Skip to content
NEXTMOVEFDE careers · United States

Machine Learning Engineer - Agentic AI Evaluation Frameworks

AI summary of the role

Apple's Channel Sales AI Product Engineering team seeks a Machine Learning Evaluation Engineer to build and scale evaluation frameworks, datasets, and tooling for GenAI/LLM products across the Commerce domain (Store AI, Shopping AI, Conversational AI).

What you’ll do

  • Design and build automated evaluation frameworks and pipelines for LLM, GenAI, Conversational AI, and Agentic AI products.
  • Define quality metrics (accuracy, relevance, groundedness, completeness, consistency, instruction following, task completion) and evaluation methodologies.
  • Develop model-based evaluation (LLM-as-a-Judge) and Human-in-the-Loop approaches with calibration and validation.
  • Build reusable evaluation infrastructure, APIs, dashboards, and tooling; integrate into CI/CD for regression detection and quality gates.

What you’ll bring

  • 7+ years in ML Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or related field.
  • Strong Python programming and experience building production-quality software, ML systems, data pipelines, or evaluation infrastructure.
  • Experience developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or ML-driven products.
  • Experience designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies.

Technologies

Python · LLM · Generative AI · RAG · LLM-as-a-Judge · CI/CD · embeddings · agentic workflows · NLP · data pipelines

About Apple

Designs and sells iPhone, Mac, iPad, Watch, and Vision Pro plus a $100B+ Services layer (App Store, iCloud, Apple Pay, Music, TV+) atop a 2.5B device installed base.

Public

Source and classification

Internal deployment & tooling · Evidence for this classification:

Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish. The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences. In this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle. You will help define how we measure
More from the job description

Imagine what you could do here. At Apple, great ideas have a way of becoming extraordinary products, services, and customer experiences very quickly. Bring passion and dedication to your work, and there’s no telling what you could accomplish. The Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered experiences. In this role, you will develop evaluation frameworks, datasets, tooling, and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning, Software Engineering, Quality Engineering, Product, Human Interface, Data Science, and domain experts to establish rigorous evaluation practices throughout the AI product lifecycle. You will help define how we measure the quality of AI experiences across the Commerce domain, including Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI, and Platform Self-Service. This is an opportunity to work at the intersection of machine learning, software engineering, data, and product quality, helping ensure our AI experiences are accurate, relevant, grounded, reliable, and useful for users around the world. Description As a Machine Learning Evaluation Engineer, you will design and build scalable evalua [... source excerpt omitted ...] on following, and task completion. ◦ Build and maintain high-quality evaluation datasets, including golden datasets, benchmark sets, regression suites, adversarial scenarios, and production-derived test sets. ◦ Develop Auto Eval capabilities that enable teams to rapidly evaluate models, prompts, retrieval systems, agents, and end-to-end AI experiences. ◦ Design and implement model-based evaluation approaches, including LLM-as-a-Judge, while developing appropriate calibration and validation methodologies. ◦ Develop Human-in-the-Loop (HITL) evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is insufficient. ◦ Define evaluation rub [... source excerpt omitted ...] luation into development and CI/CD workflows, enabling automated regression detection, quality gates, and release-readiness assessments. ◦ Connect offline evaluation results with production signals to continuously improve evaluation coverage and product quality. ◦ Partner closely with Machine Learning, Software Engineering, Product, Quality Engineering, Human Interface, and Data Science teams throughout research, development, evaluation, launch, and continuous improvement. Minimum Qualifications Typically requires a minimum of 7 years of related experience in Machine Learning Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related techn

How jobs are selected

Employer postings · Data from · Sources