Staff Applied AI Inference Engineer
Technologies
vLLM · SGLang · CUDA · Python · C++ · Docker · Kubernetes · GPU · LLM · inference
About Crusoe
Vertically integrated AI infrastructure company that develops energy, builds AI data centers, and runs GPU cloud and managed inference services for model builders and enterprises.
Series E
Job description
The full responsibilities and requirements are on the employer’s site.
Open application page ↗Source and classification
Production engineering · Evidence for this classification:
manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. About the Role You will spend your time making large language models run faster, cheaper, and more reliably in production. That means owning the inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when the defaults are not good enough. This is core systems and performance work on some of the most demanding models in use today. The work is applied, not academic. The optimizations you build land in real customer deployments, each with its own models, traffic patterns, latency targets, and costHow jobs are selected
Employer postings · Data from · Sources