Nvidia
Engineering Manager – AI Platform & SRE
Found: Today
Site Reliability Engineering (SRE) at NVIDIA is an engineering field focused on designing, building, and operating large-scale production systems with exceptional reliability, efficiency, and availability. It combines software and systems engineering practices with expertise across distributed systems, networking, Kubernetes, public cloud, observability, capacity management, continuous delivery, and automation.
As an Engineering Manager, you will lead a team of dedicated engineers responsible for building and operating resilient AI platform capabilities at enterprise scale. You will combine people leadership with strong technical judgment, helping the team translate ambiguous business and engineering challenges into a clear strategy and executable roadmap. You will partner across Cloud, Platform, Security, and AI/ML organizations to deliver reliable systems, improve developer productivity, and advance the use of AI agents and skills in platform operations.