Senior Solutions Architect, AI Inference
AI summary of the role
Senior Solutions Architect on NVIDIA's ISP team, driving adoption of full-stack inference technologies (Dynamo, TensorRT-LLM, vLLM) with strategic partners.
What you’ll do
- Work with inference partners to teach NVIDIA stack value and bring back product insights
- Build and operate inference recipes with NVIDIA Dynamo for GPU worker efficiency
- Accelerate inference pipelines using TensorRT-LLM, vLLM, SGLang for disaggregated inference
- Evangelize DevOps best practices for Kubernetes, compute fabrics, and observability
What you’ll bring
- 6+ years in Solutions Architecture or similar, with 2+ years AI workloads on Kubernetes
- Experience with NVIDIA Dynamo, Triton Inference Server, or TensorRT-LLM
- Deep knowledge of disaggregated serving, KV cache management, speculative decoding, quantization
- Hands-on full-stack agent design with sandboxing, memory, retrieval, evaluation
Technologies
NVIDIA Dynamo · TensorRT-LLM · vLLM · SGLang · Triton Inference Server · Kubernetes · NIXL · Grove · CUDA · speculative decoding
About NVIDIA
Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.
Public
Employer postings · Data from · Sources