Senior Solutions Architect, NVIDIA Cloud Partner Operations
AI summary of the role
Solve hard Day 2 operations problems at scale alongside partner engineers, prototyping and validating approaches under load.
What you’ll do
- Make new NVIDIA platforms and services Day 2 ready, driving adoption without degrading service.
- Improve reliability, performance, and economics using metrics like incident frequency, recovery time, utilization, and cost per token.
- Raise partner Day 2 maturity across people, process, tooling, telemetry, security, and incident response.
- Convert validated work into operating procedures, reference architectures, automation, and agentic workflows for the broader NCP ecosystem.
What you’ll bring
- BS/MS/PhD in CS, EE, CE, Physics, Math, or equivalent experience.
- 12+ years in production infrastructure, cloud engineering, solutions architecture, SRE, HPC, or similar; or 5+ years of exceptional specialist work in large-scale GPU/AI infrastructure.
- Deep hands-on expertise in at least one Day 2 stack area: DCGM, BMC/Redfish, InfiniBand/NCCL/UFM, or high-performance storage (Lustre, WEKA, VAST).
- Working experience with Kubernetes/Slurm, GPU scheduling, Prometheus/Grafana/OpenTelemetry, and automation tools (Terraform, Ansible, Argo CD).
Technologies
DCGM · BMC/Redfish · InfiniBand · NCCL · UFM · Lustre · WEKA · VAST Data · Kubernetes · Slurm · Prometheus · Grafana
About NVIDIA
Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.
Public
Employer postings · Data from · Sources