Skip to content
NEXTMOVEFDE careers · United States

Senior Solutions Architect, NVIDIA Cloud Partner Operations

AI summary of the role

Solve hard Day 2 operations problems at scale alongside partner engineers, prototyping and validating approaches under load.

What you’ll do

  • Make new NVIDIA platforms and services Day 2 ready, driving adoption without degrading service.
  • Improve reliability, performance, and economics using metrics like incident frequency, recovery time, utilization, and cost per token.
  • Raise partner Day 2 maturity across people, process, tooling, telemetry, security, and incident response.
  • Convert validated work into operating procedures, reference architectures, automation, and agentic workflows for the broader NCP ecosystem.

What you’ll bring

  • BS/MS/PhD in CS, EE, CE, Physics, Math, or equivalent experience.
  • 12+ years in production infrastructure, cloud engineering, solutions architecture, SRE, HPC, or similar; or 5+ years of exceptional specialist work in large-scale GPU/AI infrastructure.
  • Deep hands-on expertise in at least one Day 2 stack area: DCGM, BMC/Redfish, InfiniBand/NCCL/UFM, or high-performance storage (Lustre, WEKA, VAST).
  • Working experience with Kubernetes/Slurm, GPU scheduling, Prometheus/Grafana/OpenTelemetry, and automation tools (Terraform, Ansible, Argo CD).

Technologies

DCGM · BMC/Redfish · InfiniBand · NCCL · UFM · Lustre · WEKA · VAST Data · Kubernetes · Slurm · Prometheus · Grafana

About NVIDIA

Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.

Public

Employer postings · Data from · Sources