DevOps Engineer · Platform Engineering · Cloud & Distributed Systems

Building reliable platforms for distributed systems.

I’m Akshat, a DevOps Engineer who works across Linux, Kubernetes, cloud platforms, and automation. My recent work has focused on large-scale AI and GPU environments, where compute, networking, storage, and observability all need to work together.

I enjoy understanding how complex systems behave and making them easier to operate.

01 / How I work

The work behind a useful AI cluster.

A résumé lists tools and titles. This is the shape of the work I enjoy: making complex systems repeatable, then staying close enough to understand them when they are not.

01

Bring-up that can be repeated

I build machine images, cluster bootstrap paths, and node lifecycle automation so a new environment does not begin as a one-off exercise.

TerraformPackerAnsiblecloud-init

02

GPU workloads that have somewhere to go

I work with Kubernetes and Slurm to connect scheduling, GPU allocation, storage, and networking—the pieces a distributed job depends on before it ever starts training.

SlurmKubernetesGPU OperatorDRA

03

Debug from the symptom back to the system

When a job is slow or a node is unhealthy, I trace the path across compute, storage, and network rather than assuming the first noisy metric is the cause.

NCCLRDMAGrafanaLinux

02 / Focus areas

The parts of the stack I spend time with.

Client work stays private. These are the recurring technical surfaces and questions that shape how I approach it.

GPU clusters under load

Testing distributed jobs with NCCL, NVBandwidth, checkpointing, and inference benchmarks to separate compute, communication, and I/O limits.

Scheduling and resource allocation

Working across Slurm and Kubernetes when placement, GPU claims, stale allocations, or multi-node behaviour need a closer look.

High-performance data paths

Connecting Lustre, EFS, Filestore, and Parallelstore with the network and workload patterns they need to serve.

Safer day-to-day operations

Host diagnostics, access controls, observability, and automated remediation that reduce guesswork during an incident.

03 / Technical expertise

Breadth, organised by what it helps me do.

Linux & systems

Ubuntu, CentOS, RHEL, host-level diagnostics, Bash, Python

Cloud

AWS, GCP, OCI, Nebius, multi-cloud environments

Containers

Kubernetes, Docker, Helm, node lifecycle, workload troubleshooting

GPU computing

NVIDIA H100, B200/B300, GB200/GB300, GPU Operator, NCCL, DRA

Scheduling

Slurm, AWS ParallelCluster, Slinky, Soperator, workload orchestration

Storage & networking

Lustre, EFS, Filestore, Parallelstore, RDMA, RoCE, InfiniBand, SR-IOV

Automation

Terraform, Terragrunt, Packer, Ansible, cloud-init, GitHub Actions, GitLab CI

Visibility & response

Prometheus, Grafana, DCGM Exporter, EFK, PagerDuty, RCA

04 / Engineering notes

Sharing what I learn.

A collection of practical write-ups on Linux, Kubernetes, GPUs, networking, and distributed systems based on real engineering work and experimentation.

RDMA fixed the network problem. It didn't fix the GPU problem

Post Summary — RDMA removes the CPU and kernel from the network path, but gradients still live in GPU memory. This note explains why GPUDirect RDMA exists, how it changes the data path, and where the real bottlenecks appear during distributed training.

GDRDMAGPUInfrastructure

Read on LinkedIn
Views
<10k+>

NVIDIA MIG - It partitions hardware guarantees

Post Summary — Enabling MIG changes how Kubernetes schedules GPU resources. This post explores why idle memory doesn't translate into schedulable capacity, how MIG profiles work, and what to expect when running mixed GPU workloads.

MIGGPUKubernetes

05 / Résumé

The short version is in my résumé.

Download résumé

06 / Contact

Get in touch.

For roles and conversations around AI systems, HPC, and DevOps.

More about meBeyond engineering

Music

I write rap music and manage my artist channel, AKSAR. It is a different kind of making: writing, recording, editing, and seeing a small idea through to a finished release.

Visit AKSAR on YouTube

Design work

I have also designed brand identities with registered trademarks, including Law Pro Classes, Mach Academy, QueensTown, Mitz, Most & More, PIE Academy, and HTC.