I'm interested in how large-scale AI infrastructure constraints shape
what AI systems can actually do in production β specifically, the gap between what models are
theoretically capable of and what they reliably deliver under real operational conditions.
My work managing distributed compute β Kubernetes, OpenStack, 1,200+ hypervisors,
24,000+ VMs β and responding to production incidents gives me a ground-level view of where
AI systems break, degrade, and fail silently.
I want to bring that operational depth into research on
AI system reliability, observability, and agentic system robustness.
I believe the most important work in AI right now isn't capability β it's making capable
systems trustworthy at scale.
B.E. Electronics & Telecommunication Engineering
North Maharashtra University