Skip to content
NEXTMOVEFDE careers · United States

Member of Technical Staff - Multi-Modal, Audio

AI summary of the role

Join the Audio team at Liquid AI to build frontier speech-language models that handle STT, TTS, and speech-to-speech in a single architecture.

What you’ll do

  • Build and scale data pipelines for audio model training, including preprocessing, augmentation, and quality filtering at scale
  • Design, implement, and maintain evaluation systems that measure multimodal performance across internal and public benchmarks
  • Fine-tune and adapt audio models for customer-specific use cases, owning delivery from requirements through deployment
  • Contribute production code to the core audio repository, collaborating with infrastructure and research teams

What you’ll bring

  • Strong programming fundamentals with demonstrated ability to write clean, maintainable, production-grade code
  • Experience building and shipping production ML systems beyond model training (data pipelines, evals, serving infrastructure)
  • Proficiency in PyTorch and familiarity with distributed training frameworks (DeepSpeed, FSDP, or similar)
  • Track record of collaborating effectively in shared codebases with high engineering standards

Technologies

PyTorch · DeepSpeed · FSDP · ASR · TTS · vocoders · diarization · speech-to-speech · distributed training

About Liquid AI

Efficient general-purpose AI systems optimized for on-device deployment across data centers and edge hardware, enabling low-latency, privacy-preserving enterprise AI.

Series A

Source and classification

Internal deployment & tooling · Evidence for this classification:

About Liquid AI Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there. The Opportunity Our Audio team is building frontier speech-language models that handle STT, TTS, and speech-to-speech in a single architecture. This role sits at the center of applied audio model development, working directly with the technical lead to ship production systems that run on-device under real-time constraints. You will own critical workstreams across data pipelines, evaluation systems, and customer deployments. If you want high ownership on rare
More from the job description

About Liquid AI Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there. The Opportunity Our Audio team is building frontier speech-language models that handle STT, TTS, and speech-to-speech in a single architecture. This role sits at the center of applied audio model development, working directly with the technical lead to ship production systems that run on-device under real-time constraints. You will own critical workstreams across data pipelines, evaluation systems, and customer deployments. If you want high ownership on rare technical problems in a small, elite team where your code ships, this is the role. What We're Looking For We need someone who: Builds first, theorizes later: You ship working systems, not just notebooks. Production-grade code is your default, not a stretch goal. Owns outcomes end-to-end: From data pipelines to customer deployments, you take responsibility for the full stack without waiting for someone else to handle the hard parts. Thrives under constraints: On-device, low-latency, memory-limited [... source excerpt omitted ...] tering at scale Design, implement, and maintain evaluation systems that measure multimodal performance across internal and public benchmarks Fine-tune and adapt audio models for customer-specific use cases, owning delivery from requirements through deployment Contribute production code to the core audio repository, collaborating with infrastructure and research teams Support experimentation under real hardware constraints, shifting between customer work and core development as priorities evolve Desired Experience Must-have: Strong programming fundamentals with demonstrated ability to write clean, maintainable, production-grade code Experience building and shipping product [... source excerpt omitted ...] uted GPU clusters Open-source contributions that demonstrate code quality and engineering judgment What Success Looks Like (Year One) Within 6 months, you independently deliver production-ready data pipelines or evaluation systems and own at least one customer workstream end-to-end Your PRs to the core audio repo are accepted without heavy rework, demonstrating strong judgment in system design By year end, you operate as a second pillar to the technical lead, unblocking parallel workstreams and raising overall team velocity What We Offer Rare technical problems: Work on audio-to-audio frontier systems with real ownership in a team small enough that your contributions ship di

How jobs are selected

Employer postings · Data from · Sources