Skip to content
NEXTMOVEFDE careers · United States

Staff Applied AI Inference Engineer

Technologies

vLLM · SGLang · CUDA · Python · C++ · Docker · Kubernetes · GPU · LLM · inference

About Crusoe

Vertically integrated AI infrastructure company that develops energy, builds AI data centers, and runs GPU cloud and managed inference services for model builders and enterprises.

Series E

Job description

The full responsibilities and requirements are on the employer’s site.

Open application page
Source and classification

Production engineering · Evidence for this classification:

manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. About the Role You will spend your time making large language models run faster, cheaper, and more reliably in production. That means owning the inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when the defaults are not good enough. This is core systems and performance work on some of the most demanding models in use today. The work is applied, not academic. The optimizations you build land in real customer deployments, each with its own models, traffic patterns, latency targets, and cost
How jobs are selected

Employer postings · Data from · Sources