NCX Senior Engineer
AI summary of the role
NVIDIA is hiring a senior engineer to lead Day 2 operational readiness for NVIDIA Cloud Partners, ensuring large-scale accelerated infrastructure runs reliably in production.
What you’ll do
- Lead NCP Day 2 operational readiness efforts with NVIDIA Cloud Partners
- Build continuous validation of GPU, CPU, storage, and network health across large-scale AI clusters
- Establish observability and telemetry across compute, GPU, InfiniBand/RoCE, storage, Kubernetes, and AI workloads
- Develop automated detection and remediation workflows to minimize workload disruption
What you’ll bring
- 8+ years in infrastructure engineering, SRE, DevOps, or similar roles
- Deep Kubernetes, containers, and cluster scheduling experience
- Strong production observability skills (metrics, logging, alerting, dashboards)
- Automation experience with Python, Go, or shell scripting
Technologies
Kubernetes · GPU · InfiniBand · RoCE · CUDA · DGX · HGX · NVLink · NVSwitch · Prometheus · Grafana · OpenTelemetry
About NVIDIA
Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.
Public
Employer postings · Data from · Sources