Principal Machine Learning Engineer, Conversational AI Modeling and Learning
AI summary of the role
Principal Engineer leading the engineering of Alexa+'s agentic AI platform at Amazon.
What you’ll do
- Define and drive the engineering roadmap and architecture for the agentic AI platform: evaluation, training, self-learning, and serving for LLM-based agents in production
- Architect large-scale agentic evaluation infrastructure: isolated sandboxed execution, recreatable environments, verifiable scoring, and reproducibility at hundreds of concurrent trials
- Build and scale RL and post-training systems for agentic workloads: 256K+ token contexts, multi-turn trajectory training, train/inference engine consistency, and reward attribution across long sessions
- Design the serving and inference architecture for agentic traffic (long sessions, output-generation-bound workloads, KV-cache-centric optimization), co-designing with inference-infrastructure partner teams
What you’ll bring
- 12+ years of non-internship professional software development experience
- Knowledge of object-oriented design, data structures, and algorithms
- Experience designing and building large-scale systems in a multi-tiered, distributed environment (Service Oriented Architecture)
Technologies
LLM · reinforcement learning · RL · agentic AI · KV-cache · distributed systems · Service Oriented Architecture · Alexa+ · inference · post-training
About Amazon (incl AWS)
Online retail, third-party marketplace, Prime/ads/devices, and AWS — the world's largest cloud platform powering startups and enterprises.
Public
Source and classification
Internal deployment & tooling · Evidence for this classification:
Alexa AI is building the next generation of Alexa+, Amazon's LLM-powered conversational assistant, and its future is agentic: LLM systems that reason and act over dozens of chained inferences, coupled to real environments where their actions persist. Making these agents smarter, faster, and cheaper is as much a systems problem as a modeling problem - agent performance depends on the model, the harness, the evaluation infrastructure, and the serving stack co-designed together. We are looking for a Principal Engineer to lead the engineering of this agentic platform. You will own the architecture that turns research into production capability: large-scale agentic evaluation infrastructure (sandboxed, reproducible, statistically trustworthy at high concurrency), reinforcement learning training systems for long-horizon multi-turn trajectories, self-learning pipelines that convert production
More from the job description
Alexa AI is building the next generation of Alexa+, Amazon's LLM-powered conversational assistant, and its future is agentic: LLM systems that reason and act over dozens of chained inferences, coupled to real environments where their actions persist. Making these agents smarter, faster, and cheaper is as much a systems problem as a modeling problem - agent performance depends on the model, the harness, the evaluation infrastructure, and the serving stack co-designed together. We are looking for a Principal Engineer to lead the engineering of this agentic platform. You will own the architecture that turns research into production capability: large-scale agentic evaluation infrastructure (sandboxed, reproducible, statistically trustworthy at high concurrency), reinforcement learning training systems for long-horizon multi-turn trajectories, self-learning pipelines that convert production experience into permanent model and system improvements, and the serving architecture for latency-sensitive agentic inference. You will partner closely with scientists and work backwards from committed product launches, setting the technical bar for a platform that serves every Alexa agent rather than one product at a time. The charter is the full lifecycle of a production agent: how it is measured, how it is trained, how it learns, and how it is served. You will build the harnesses and sandbox [... source excerpt omitted ...] t make their quality provable rather than asserted, the RL infrastructure that trains models on the same tasks they are measured on, and the self-improvement loop that turns every production interaction into a permanently smarter system - agents that ship better than they launched, week over week. Few places let one engineer shape the entire loop from a customer's spoken request to a model that learned from it; this role owns that loop at Alexa scale. Key job responsibilities Define and drive the engineering roadmap and architecture for the agentic AI platform: evaluation, training, self-learning, and serving for LLM-based agents in production Architect large-scale agentic evalu [... source excerpt omitted ...] anization: raise engineering standards through design reviews, operational excellence, and deep dives on the hardest cross-system problems Translate ambiguous product and science requirements into platform interfaces partner teams can build on; influence senior leadership on build-vs-adopt and ownership decisions Mentor and grow senior and principal-track engineers across multiple teams A day in the life You might spend the morning in a design review for the next generation of the evaluation platform's execution layer, midday debugging why a 100K-token training session diverges between the rollout engine and the learner, and the afternoon with the serving team deciding which KV-c
Employer postings · Data from · Sources