Skip to content
NEXTMOVEFDE careers · United States

Senior Solutions Architect, AI Inference

AI summary of the role

Senior Solutions Architect on NVIDIA's ISP team, driving adoption of full-stack inference technologies (Dynamo, TensorRT-LLM, vLLM) with strategic partners.

What you’ll do

  • Work with inference partners to teach NVIDIA stack value and bring back product insights
  • Build and operate inference recipes with NVIDIA Dynamo for GPU worker efficiency
  • Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang for disaggregated inference
  • Evangelize DevOps best practices for Kubernetes, compute fabrics, and observability

What you’ll bring

  • 6+ years in Solutions Architecture or similar, with 2+ years AI workloads on Kubernetes
  • Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM
  • Deep knowledge of disaggregated serving, KV cache management, speculative decoding, quantization
  • Hands-on full-stack agent design with sandboxing, memory, retrieval, evaluation

Technologies

NVIDIA Dynamo · TensorRT-LLM · vLLM · SGLang · Triton Inference Server · Kubernetes · NIXL · Grove · CUDA · speculative decoding

About NVIDIA

Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.

Public

Employer postings · Data from · Sources