Senior/Staff Machine Learning Research Engineer, General Agents, Enterprise GenAI
AI summary of the role
Design, build, and deploy production-grade AI agents for enterprise use cases on Scale's General Agents team.
What you’ll do
- Design and implement end-to-end agent systems combining LLM reasoning, tool use, memory, and control logic for recurring enterprise use cases.
- Build scalable, reliable agent architectures deployable across many customers with varying data, tools, and constraints.
- Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production.
- Productionize frontier agent techniques (planning, multi-step reasoning, multi-agent patterns) into maintainable, observable systems.
What you’ll bring
- 5+ years building and deploying ML/AI systems for real-world production use cases.
- Strong engineering fundamentals with a Bachelor's/Master's in CS, ML, AI, or equivalent practical experience.
- Deep understanding of modern LLMs, prompt/context/system-level optimization, and agentic system design.
- Proven proficiency in Python, including production-quality, testable, maintainable code.
Technologies
LLM · Python · OpenAI APIs · SFT · RLVR · LoRA · agent frameworks · tool calling · multi-agent systems
About Scale AI
Data, evals, and GenAI platform for frontier labs, enterprises, and governments — the picks-and-shovels layer for production AI.
Series G
Source and classification
Deployment team leadership · Evidence for this classification:
Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role
More from the job description
Scale AI is the data foundation for AI, helping organizations build and deploy reliable production AI applications. We partner with leading enterprises and government organizations to accelerate their AI initiatives through our data annotation platform, generative AI solutions, and enterprise AI capabilities. About the General Agents Team The General Agents team, part of Scale’s Enterprise organization, builds robust general agents for customer use cases and applications. The team sits at the intersection of frontier agent development and real-world deployment, translating state-of-the-art reasoning and agentic capabilities into reliable, production-grade systems that drive real economic value. Our agents are scalable systems built around recurring enterprise problem domains, with a strong emphasis on generalization, extensibility, and deployment across many customers. About the Role As a Senior/Staff Machine Learning Engineer (MLE) on the General Agents team, you’ll play a critical role in designing, building, and deploying production-ready AI agents that solve high-impact enterprise problems. You will work across the full agent lifecycle—from model and system design to evaluation, deployment, and iteration—bridging cutting-edge agentic techniques with the constraints and requirements of real customer environments. You will: Design and implement end-to-end agent systems [... source excerpt omitted ...] combine LLM reasoning, tool use, memory, and control logic to solve recurring enterprise use cases. Build scalable, reliable agent architectures that can be deployed across many customers with varying data, tools, and constraints. Develop evaluation frameworks, datasets, environments, and metrics to measure agent performance, reliability, and business impact in production settings. Collaborate closely with product managers, customers, data annotators, and other engineering teams to translate enterprise requirements into robust agent designs. Productionize frontier agent techniques (e.g., planning, multi-step reasoning and tool-use, multi-agent patterns) into maintainable, ob [... source excerpt omitted ...] ethods, with increasing scope and leadership at the Staff level. Ideally you’d have: 5+ years of experience building and deploying machine learning or AI systems for real-world, production use cases. Strong engineering fundamentals, supported by a Bachelor’s and/or Master’s degree in Computer Science, Machine Learning, AI, or equivalent practical experience. Deep understanding of modern LLMs, prompt-, context-, and system-level optimization, and agentic system design. Proven proficiency in Python, including writing production-quality, testable, and maintainable code. Experience building systems that integrate models with external tools, APIs, databases, and services. Ability
Employer postings · Data from · Sources