Staff Applied AI Inference Engineer
AI summary of the role
Staff-level hands-on role owning the LLM inference stack end to end at a vertically integrated AI infrastructure company.
What you’ll do
- Bring current inference techniques into production and refine them.
- Design and optimize serving architectures, including prefill and decode disaggregation, request routing, and related approaches.
- Profile and tune deployments against clear targets for latency, throughput, and cost, and keep them dependable under real traffic.
- Tailor deployments to each customer's models and constraints, partnering with their engineering teams to move a workload from an early proof of concept through to a live, well-monitored production service.
What you’ll bring
- Hands-on experience shipping code in production with one or more general-purpose languages, such as Python or C++, with a strong preference for Python.
- Familiarity with methods for optimizing LLMs for high throughput / low latency inference.
- Comfort with modern LLM serving frameworks such as vLLM or SGLang, and with profiling and analyzing performance down to the kernel level.
- A firm grasp of how GPUs are built and how they behave.
Technologies
LLM inference · vLLM · SGLang · CUDA · GPU · Python · C++ · Docker · Kubernetes · prefill/decode disaggregation
About Crusoe
Vertically integrated AI infrastructure company that develops energy, builds AI data centers, and runs GPU cloud and managed inference services for model builders and enterprises.
Series E
Source and classification
Production engineering · Evidence for this classification:
manufacturing, data center construction, and cloud services. If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe. About the Role You will spend your time making large language models run faster, cheaper, and more reliably in production. That means owning the inference stack end to end: profiling where time and cost go, bringing modern optimization techniques into real deployments, and getting deep into the serving code when the defaults are not good enough. This is core systems and performance work on some of the most demanding models in use today. The work is applied, not academic. The optimizations you build land in real customer deployments, each with its own models, traffic patterns, latency targets, and costHow jobs are selected
Employer postings · Data from · Sources