01
Bring-up that can be repeated
I build machine images, cluster bootstrap paths, and node lifecycle automation so a new environment does not begin as a one-off exercise.
TerraformPackerAnsiblecloud-init
DevOps Engineer · Platform Engineering · Cloud & Distributed Systems
I’m Akshat, a DevOps Engineer who works across Linux, Kubernetes, cloud platforms, and automation. My recent work has focused on large-scale AI and GPU environments, where compute, networking, storage, and observability all need to work together.
I enjoy understanding how complex systems behave and making them easier to operate.
01 / How I work
A résumé lists tools and titles. This is the shape of the work I enjoy: making complex systems repeatable, then staying close enough to understand them when they are not.
01
I build machine images, cluster bootstrap paths, and node lifecycle automation so a new environment does not begin as a one-off exercise.
TerraformPackerAnsiblecloud-init
02
I work with Kubernetes and Slurm to connect scheduling, GPU allocation, storage, and networking—the pieces a distributed job depends on before it ever starts training.
SlurmKubernetesGPU OperatorDRA
03
When a job is slow or a node is unhealthy, I trace the path across compute, storage, and network rather than assuming the first noisy metric is the cause.
NCCLRDMAGrafanaLinux
02 / Focus areas
Client work stays private. These are the recurring technical surfaces and questions that shape how I approach it.
Testing distributed jobs with NCCL, NVBandwidth, checkpointing, and inference benchmarks to separate compute, communication, and I/O limits.
Working across Slurm and Kubernetes when placement, GPU claims, stale allocations, or multi-node behaviour need a closer look.
Connecting Lustre, EFS, Filestore, and Parallelstore with the network and workload patterns they need to serve.
Host diagnostics, access controls, observability, and automated remediation that reduce guesswork during an incident.
03 / Technical expertise
Ubuntu, CentOS, RHEL, host-level diagnostics, Bash, Python
AWS, GCP, OCI, Nebius, multi-cloud environments
Kubernetes, Docker, Helm, node lifecycle, workload troubleshooting
NVIDIA H100, B200/B300, GB200/GB300, GPU Operator, NCCL, DRA
Slurm, AWS ParallelCluster, Slinky, Soperator, workload orchestration
Lustre, EFS, Filestore, Parallelstore, RDMA, RoCE, InfiniBand, SR-IOV
Terraform, Terragrunt, Packer, Ansible, cloud-init, GitHub Actions, GitLab CI
Prometheus, Grafana, DCGM Exporter, EFK, PagerDuty, RCA
04 / Engineering notes
A collection of practical write-ups on Linux, Kubernetes, GPUs, networking, and distributed systems based on real engineering work and experimentation.
Post Summary — RDMA removes the CPU and kernel from the network path, but gradients still live in GPU memory. This note explains why GPUDirect RDMA exists, how it changes the data path, and where the real bottlenecks appear during distributed training.
GDRDMAGPUInfrastructure
Post Summary — Enabling MIG changes how Kubernetes schedules GPU resources. This post explores why idle memory doesn't translate into schedulable capacity, how MIG profiles work, and what to expect when running mixed GPU workloads.
MIGGPUKubernetes
05 / Résumé
06 / Contact
For roles and conversations around AI systems, HPC, and DevOps.
I write rap music and manage my artist channel, AKSAR. It is a different kind of making: writing, recording, editing, and seeing a small idea through to a finished release.
Visit AKSAR on YouTubeI have also designed brand identities with registered trademarks, including Law Pro Classes, Mach Academy, QueensTown, Mitz, Most & More, PIE Academy, and HTC.