Skip to content
NEXTMOVEFDE careers · United States

Sr. Machine Learning Engineer, Speech LLM Evaluation

AI summary of the role

This role owns the data and metrics foundation for evaluating speech LLMs within Apple's Siri organization, partnering with modeling teams to design datasets and metrics that stress-test real-time speech understanding and generation models.

What you’ll do

  • Own the data and metrics foundation for evaluating speech LLMs across accuracy, robustness, and conversational quality.
  • Build and curate evaluation datasets reflecting real usage, from personalized named-entity queries to multi-turn fluid conversations.
  • Design metrics and automated judges that turn model outputs into actionable, trustworthy signal.
  • Partner with modeling, infrastructure, and product teams to ensure every new model is evaluated quickly, consistently, and with rigor before reaching customers.

What you’ll bring

  • Bachelor's degree in CS, EE, or related field, or equivalent practical experience.
  • Experience building or working with text, speech, or audio evaluation pipelines and metrics.
  • Proficiency in Python and experience building data processing pipelines at scale.
  • Experience curating or annotating datasets for ML evaluation or training.

Technologies

Python · Spark · speech LLM · ASR · TTS · LLM-as-judge · multimodal · audio evaluation · data pipelines · statistics

About Apple

Designs and sells iPhone, Mac, iPad, Watch, and Vision Pro plus a $100B+ Services layer (App Store, iCloud, Apple Pay, Music, TV+) atop a 2.5B device installed base.

Public

Source and classification

Internal deployment & tooling · Evidence for this classification:

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and real-time conversational AI efforts, partnering directly with the modeling teams. We're growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether they're ready. You'll help define how Apple measures a new class of models that listen, speak, and reason.
More from the job description

Join the team redefining what a deeply personal and integrated assistant can be. As part of the Siri organization, you will help shape one of the world's most widely used AI assistants, powered by our next-generation of Apple Intelligence, with capabilities like personal context understanding and on-screen awareness, built with privacy from the ground up. Your work will have direct, meaningful impact for users across iOS, iPadOS, macOS, watchOS, and visionOS. Our Speech Evaluation team sits at the center of Apple's ASR, TTS, and real-time conversational AI efforts, partnering directly with the modeling teams. We're growing the team to take on a role focused specifically on evaluating audio LLMs: designing the datasets that stress-test them and the metrics that decide whether they're ready. You'll help define how Apple measures a new class of models that listen, speak, and reason. This is a rare opportunity to build at the intersection of cutting-edge AI and human-centered design, shipping technology that is centered around users and their needs. Description This role owns the data and metrics foundation for evaluating speech LLMs (e.g., real-time speech understanding and generation models) across accuracy, robustness, and conversational quality. You'll build and curate evaluation datasets that reflect real usage — from personalized named-entity queries to multi-turn fluid [... source excerpt omitted ...] work closely with modeling, infrastructure, and product partners to make sure every new model is evaluated quickly, consistently, and at the right level of rigor before it reaches customers. Minimum Qualifications Bachelor's degree in Computer Science, Electrical Engineering, or a related field, or equivalent practical experience. Experience building or working with text, speech or audio evaluation pipelines and metrics. Proficiency in Python and experience building data processing pipelines at scale. Experience curating or annotating datasets for machine learning evaluation or training. Working knowledge of statistics as applied to measuring model performance and interpreting [... source excerpt omitted ...] human evaluation methods. Strong written and verbal communication skills, with the ability to explain evaluation results to both technical and non-technical audiences. Preferred Qualifications Experience evaluating audio-native or multimodal (speech-in, speech-out) large language models. Experience designing or running human evaluation studies (e.g., side-by-side comparisons, MOS ratings) at scale. Familiarity with personalization and named-entity evaluation challenges in speech systems. Experience with multilingual or international audio dataset development. Experience with distributed data processing frameworks (e.g., Spark) for large-scale audio dataset generation. Publication re

How jobs are selected

Employer postings · Data from · Sources