Skip to content
NEXTMOVEFDE careers · United States

Senior Machine Learning Engineer, Alexa-Conv A Modeling&Learning

AI summary of the role

Senior ML engineer to own core systems in Alexa's agentic AI platform, spanning evaluation infrastructure, RL training, self-learning pipelines, and inference serving.

What you’ll do

  • Design, build, and operate major components of the agentic AI platform: evaluation harnesses, sandboxed environments, RL and post-training pipelines, self-learning data pipelines, or inference serving for agentic traffic
  • Lead design work in your area: write design documents, drive them through review, and make build-versus-adopt calls within your scope
  • Own reliability and performance: instrument systems, drive down failure modes, and report platform health in metrics
  • Partner with applied scientists to turn research code into production infrastructure and expose it through usable interfaces

What you’ll bring

  • 5+ years of non-internship professional software development experience
  • 5+ years of programming with at least one software programming language
  • 5+ years of leading design or architecture (design patterns, reliability and scaling) of new and existing systems
  • Experience as a mentor, tech lead, or leading an engineering team

Technologies

LLM · reinforcement learning · agentic AI · evaluation infrastructure · inference serving · distributed systems · Python · RL training · self-learning pipelines

About Amazon (incl AWS)

Online retail, third-party marketplace, Prime/ads/devices, and AWS — the world's largest cloud platform powering startups and enterprises.

Public

Source and classification

Internal deployment & tooling · Evidence for this classification:

Alexa AI is building the next generation of Alexa+, Amazon's LLM-powered conversational assistant, and its future is agentic: LLM systems that reason and act over dozens of chained inferences, coupled to real environments where their actions persist. Making these agents smarter, faster, and cheaper is as much a systems problem as a modeling problem - agent performance depends on the model, the harness, the evaluation infrastructure, and the serving stack co-designed together. We are looking for a Senior Machine Learning Engineer to build and own core systems in this agentic platform. You will take one of its foundational areas - agentic evaluation infrastructure, reinforcement learning training systems, self-learning pipelines, or agentic inference serving - and own it end to end: the design, the implementation, the operational bar, and the interfaces that scientists and partner teams
More from the job description

Alexa AI is building the next generation of Alexa+, Amazon's LLM-powered conversational assistant, and its future is agentic: LLM systems that reason and act over dozens of chained inferences, coupled to real environments where their actions persist. Making these agents smarter, faster, and cheaper is as much a systems problem as a modeling problem - agent performance depends on the model, the harness, the evaluation infrastructure, and the serving stack co-designed together. We are looking for a Senior Machine Learning Engineer to build and own core systems in this agentic platform. You will take one of its foundational areas - agentic evaluation infrastructure, reinforcement learning training systems, self-learning pipelines, or agentic inference serving - and own it end to end: the design, the implementation, the operational bar, and the interfaces that scientists and partner teams build on. You will work directly with applied scientists, work backwards from committed product launches, and turn research prototypes into infrastructure that runs unattended at scale. The work is concrete. Agents are evaluated in sandboxed, recreatable environments at hundreds of concurrent trials, and every source of infrastructure noise you remove is a model decision the organization can trust. They are trained on long-horizon multi-turn trajectories where the rollout and learner engines hav [... source excerpt omitted ...] nder latency budgets measured in hundreds of milliseconds. And they improve week over week only if the pipeline that turns production experience into training data actually holds. You will own a piece of that loop, make it reliable, and make it fast. This is a platform role with room to grow. The systems you own serve every Alexa agent program rather than a single product, and the engineer who makes them dependable becomes the person the organization routes its hardest cross-system problems to. Key job responsibilities Design, build, and operate major components of the agentic AI platform: evaluation harnesses, sandboxed environments, RL and post-training pipelines, self-learn [... source excerpt omitted ...] , drive down failure modes that make results untrustworthy, and report platform health in metrics rather than anecdotes Partner with applied scientists to turn research code into production infrastructure and expose it through interfaces other teams can use without your involvement Mentor engineers on your team, raise the engineering bar through code and design reviews, and help set technical direction across your systems A day in the life You might spend the morning making the evaluation platform reproducible under high concurrency, tracking down why scores drift when a hundred trials share a host. Midday you pair with a scientist to get a long-context training job to converge

How jobs are selected

Employer postings · Data from · Sources