Senior Software Engineer, Cloud-Native Stack – CSP Engagements
AI summary of the role
NVIDIA is hiring a Senior Software Engineer to join the CSP Engagements team, focusing on the cloud-native stack for multi-rack, multi-tenant AI datacenters powered by GB200 and GB300 GPUs.
What you’ll do
- Perform deep-dive debugging of multi-rack, multi-tenant clusters: scheduler behavior, container runtime issues, device-plugin crashes, RDMA/IB fabric anomalies.
- Gather customer requirements and prototype feature extensions for Kubernetes operators, Slurm plugins, and custom micro-services that expose new GPU capabilities.
- Drive joint architecture reviews and whiteboard sessions with CSP and internal platform teams; convert findings into RFCs and upstream pull requests.
- Create reproducible testbeds (Helm/Ansible/Terraform) that mirror customer environments; automate validation and benchmark suites.
What you’ll bring
- Strong source-level expertise in Kubernetes internals (scheduler, CRI/CNI/CSI, operators) and Slurm (federation, power-save, plugins).
- Hands-on experience integrating next-gen GPUs (Blackwell/GB200/GB300) or comparable accelerators into containerized clusters.
- Proven track record debugging large-scale, cloud-native stacks across networking (RDMA/RoCE), storage, and control planes.
- Customer-facing engineering or solutions-architect background: requirements gathering, PoC ownership, roadmap influence.
Technologies
Kubernetes · Slurm · Go · Rust · C/C++ · Python · RDMA · RoCE · Helm · Ansible · Terraform · Prometheus
About NVIDIA
Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.
Public
Employer postings · Data from · Sources