Available for Research Roles & Senior Positions

Nikita Mahajan

Senior Infrastructure Engineer  Β·  AI Systems & Reliability

I operate at the intersection of large-scale distributed infrastructure and AI system reliability β€” managing 1,200+ hypervisors, 24,000+ VMs, and Kubernetes clusters that keep production AI workloads running. I'm fascinated by where AI systems silently break and what it takes to build ones that don't.

CKA CKNA RHCA VCP-DCV AWS SAA Mumbai, India 🌐 Open to Remote ✈️ Open to Relocation
nikita@infra ~
❯ kubectl get nodes
NAME             STATUS   ROLES
control-plane-01  Ready    control-plane
worker-node-01    Ready    worker
worker-node-02    Ready    worker
 
❯ whoami --verbose
role: "Senior Infra Engineer"
focus: "AI Systems Reliability"
certs: ["CKA","CKNA","RHCA","VCP","AWS"]
hypervisors: 1200+
vms: 24000+
 
❯

Research Interest

I'm interested in how large-scale AI infrastructure constraints shape what AI systems can actually do in production β€” specifically, the gap between what models are theoretically capable of and what they reliably deliver under real operational conditions.

My work managing distributed compute β€” Kubernetes, OpenStack, 1,200+ hypervisors, 24,000+ VMs β€” and responding to production incidents gives me a ground-level view of where AI systems break, degrade, and fail silently.

I want to bring that operational depth into research on AI system reliability, observability, and agentic system robustness. I believe the most important work in AI right now isn't capability β€” it's making capable systems trustworthy at scale.

B.E. Electronics & Telecommunication Engineering
North Maharashtra University

8+
Years in Infrastructure
24K+
VMs Managed
1200+
Hypervisors Operated
6
Active Certifications

Professional Experience

Acqueon by Five9
Senior Infrastructure Engineer
Aug 2025 – Present
  • Design and maintain CI/CD pipelines that automate deployment of AI-powered contact centre services across cloud and on-prem environments.
  • Author AWX/Ansible playbooks to eliminate manual operational tasks β€” reducing mean time to deploy by standardising repeatable workflows across teams.
  • Manage Kubernetes clusters (K8s) supporting production AI workloads: scheduling, scaling, resource allocation, and failure recovery.
  • Lead RCA for S1/S2 incidents, documenting failure patterns, cascading dependencies, and systemic fixes β€” building an internal knowledge base of how distributed AI systems fail.
  • Maintain OpenStack environments and hypervisor infrastructure; coordinate on-call response for production outages with zero tolerance for data loss.
  • Implement observability stack using Grafana to monitor infrastructure health and correlate infrastructure events with application-level failures.
Capgemini
Consultant – Infrastructure Analyst
Jun 2023 – Aug 2025
  • Managed VMware infrastructure at scale: 72 vCenters, 1,200+ ESXi 7.0 hosts, and 24,000+ virtual machines β€” one of the largest VMware environments in the region.
  • Implemented automation using Ansible and PowerCLI; wrote Python and Bash scripts to streamline monitoring and reporting workflows.
  • Led hypervisor patching cycles with zero unplanned downtime, coordinating change windows across global teams.
  • Administered 2,000+ Linux servers (RHEL/Ubuntu): storage provisioning with LVM, user management, performance tuning, and security hardening.
  • Supported AWS and OpenStack cloud infrastructure β€” provisioning, monitoring, and orchestration of compute and network resources.
Tata Consultancy Services (TCS)
IT System Administrator
Aug 2021 – Jun 2023
  • Maintained VMware vSphere 7.0 environments with hands-on DRS, HA, vMotion, and Storage vMotion; reduced unplanned VM downtime through proactive HA tuning.
  • Administered CISCO UCS Manager (UCSM) and storage systems (3PAR, Pure Storage, NetApp); resolved storage performance bottlenecks impacting production workloads.
  • Monitored infrastructure health using Grafana and Logic Monitor; configured alerting thresholds that reduced MTTD for critical incidents.
  • Handled end-to-end incident management for S1/S2 outages: triage, escalation, resolution, and post-incident reports.
NTT India Pvt Ltd
Technical Support Engineer
Sep 2020 – Aug 2021
  • Provided L2/L3 support for data centre infrastructure including Windows/Linux servers, networking (DNS, DHCP, Firewall, MTU), and storage systems.
  • Troubleshot and resolved complex infrastructure incidents; documented resolution playbooks that reduced repeat escalations.
Ramelex Pvt Ltd
Trainee Engineer
Mar 2018 – Jan 2020
  • Gained foundational experience in server administration, network configuration, and data centre operations.

Technical Skills

βš™οΈ
Orchestration & Automation
Kubernetes (K8s) AWX / Ansible CI/CD Pipelines PowerCLI Bash Python
☁️
Cloud & Virtualisation
AWS (EC2, VPC, IAM, S3) OpenStack VMware vSphere 7.0 ESXi vCenter
πŸ“Š
Observability & Monitoring
Grafana Logic Monitor Alerting Pipelines Incident Dashboards
πŸ–₯️
Infrastructure Hardware
CISCO UCSM Dell / HP Servers 3PAR Pure Storage NetApp
🐧
Operating Systems
Red Hat Enterprise Linux Ubuntu Windows Server
πŸ”¬
Research & Reliability
RCA Documentation S1/S2 Incident Analysis Failure Pattern Recognition Post-mortem Methodology

Licences & Certifications

Certified Kubernetes Administrator (CKA)
Cloud Native Computing Foundation (CNCF)
βœ“ Active
Certified Kubernetes Network Associate (CKNA)
Cloud Native Computing Foundation (CNCF)
βœ“ Active
Red Hat Certified Architect (RHCA)
Red Hat
May 2025
Advanced Ubuntu Administration
Canonical
Apr 2026
VMware Certified Professional: Data Centre Virtualisation (VCP-DCV)
VMware
Dec 2021
AWS Certified Solutions Architect – Associate
Amazon Web Services
Jan 2023

Technical Blog

πŸ”¬
Kubernetes Reliability

Why Kubernetes Silently Fails: Patterns from Production AI Workloads

After managing K8s clusters running AI inference workloads, I've catalogued the failure modes that don't throw errors β€” they just degrade. Here's what I found and how to detect them early.

πŸ—οΈ
OpenStack Scale

OpenStack at Scale: Lessons from 1,200+ Hypervisors

Operating one of the region's largest VMware environments taught me that scale doesn't just multiply problems β€” it creates entirely new failure categories. A post-mortem of what I learned.

πŸ“‹
Incident Response RCA

My RCA Methodology for S1/S2 Incidents in Distributed Systems

A structured approach to root cause analysis that goes beyond the immediate trigger β€” tracing cascading dependencies, documenting systemic fixes, and building an institutional memory of how systems fail.

Get in Touch

βœ‰οΈ
Email
nikitamah1995@gmail.com
πŸ“
Location
Mumbai, India  Β·  Open to Remote & Relocation
πŸ’»
πŸ“¬
Want my number?
Send me a message via the form and I'll get back to you.