Network Operations Engineer, AI Networking
OpenAI · San Francisco
About this role
About the Team
OpenAI’s Infrastructure Operations team is responsible for the availability, reliability, and operational excellence of one of the world’s largest AI infrastructure networks. The team owns day-to-day operations of production AI networks across Industrial Compute's data centers, working with colocation providers, deployment teams, and hardware vendors to deliver highly available GPU infrastructure for AI training and inference workloads.
About the Role
We are seeking an Infrastructure Operations Engineer to operate and improve the large-scale Ethernet fabrics that support GPU clusters, storage systems, and management infrastructure. This role combines hands-on production operations with automation, observability, and incident response across a global AI network.
The ideal candidate has experience operating high-availability data center, cloud, AI, or HPC networks and can move comfortably from physical-layer troubleshooting to routing and fabric behavior, change execution, and root-cause analysis.…
Summary from OpenAI's official Ashby career feed — read the full description on the original posting ↗
The salary and details above come straight from OpenAI's official Ashby feed. The full job description lives on OpenAI's careers page — AI Stack Jobs links you straight to it, we never sit between you and the employer.
More AI/ML roles at OpenAI
- 2026-07-27 San Francisco
- 2026-07-27 San Francisco
- 2026-07-27 San Francisco
- 2026-07-25 San Francisco
- 2026-07-24 San Francisco
- 2026-07-24 San Francisco