Senior Staff Applied AI Inference Engineer
$250,000–$300,000San Francisco, CAOn-siteStaff
View posting at CrusoeAbout the job
This hands-on senior staff engineer owns end-to-end inference performance for large language models, optimizing serving architectures, frameworks, and GPU kernels for latency, throughput, cost, and reliability.
Summary written from the posting.
What you’ll do
- Optimize LLM serving architectures, including prefill/decode disaggregation and request routing
- Profile and improve performance across vLLM, SGLang, and CUDA kernels
- Tune production deployments for latency, throughput, cost, and reliability
- Partner with customer engineering teams to deliver monitored inference services
What you’ll bring
- Bachelor's, Master's, or Ph.D. in computer science, engineering, mathematics, or a related field
- Production software development experience in Python or C++
- Experience optimizing LLMs for high-throughput, low-latency inference
- Experience with vLLM or SGLang and kernel-level performance profiling
Technologies
vLLM · SGLang · CUDA · Python · C++ · Docker · Kubernetes
About Crusoe
Vertically integrated AI infrastructure company that sources energy, builds and manufactures AI data centers, and delivers GPU cloud and managed inference services for model builders, hyperscalers, and enterprises.
Series F · 1000–2000 people
Employer postings · How we read postings