Member of Technical Staff - Post Training, Applied (Audio)
AI summary of the role
Own applied post-training for Liquid's LFM2.5-Audio model, adapting it for enterprise customers with a focus on voice-driven function calling.
What you’ll do
- Act as technical owner for enterprise audio post-training engagements end-to-end.
- Design and build function calling capabilities for audio models, mapping spoken intents to structured tool calls.
- Design and execute data generation pipelines for speech-to-speech and text-to-text training.
- Run supervised fine-tuning, preference alignment, and reinforcement learning workflows on audio language models.
What you’ll bring
- Hands-on experience with post-training for language models (SFT, preference alignment, and/or RL).
- Experience with data generation and evaluation pipelines for LLM or audio model training.
- Strong intuition for data quality and evaluation design.
- Familiarity with function calling, tool use, or structured output training for language models.
Technologies
LFM2.5-Audio · speech-to-speech · ASR · TTS · supervised fine-tuning · preference alignment · reinforcement learning · function calling · structured output · on-device inference
About Liquid AI
Efficient general-purpose AI systems optimized for on-device deployment across data centers and edge hardware, enabling low-latency, privacy-preserving enterprise AI.
Series A
Source and classification
Production engineering · Evidence for this classification:
About Liquid AI Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there. The Opportunity LFM2.5-Audio is Liquid's end-to-end multimodal speech and text language model. At 1.5B parameters, it handles speech-to-speech conversation, ASR, and TTS without requiring separate components, making it uniquely suited for real-time, on-device deployment. We're now bringing this model to enterprise customers. The core challenge: teaching audio models to understand user intents and translate them into structured tool calls. Think voice-driven
More from the job description
About Liquid AI Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there. The Opportunity LFM2.5-Audio is Liquid's end-to-end multimodal speech and text language model. At 1.5B parameters, it handles speech-to-speech conversation, ASR, and TTS without requiring separate components, making it uniquely suited for real-time, on-device deployment. We're now bringing this model to enterprise customers. The core challenge: teaching audio models to understand user intents and translate them into structured tool calls. Think voice-driven function calling, where a spoken request triggers the right API, extracts the right parameters, and confirms back to the user in natural speech. This role sits at the intersection of frontier audio models and real-world deployment. You'll own the applied post-training work that adapts LFM2.5-Audio for customer use cases end-to-end, from data generation through delivery. Unlike most roles that force a trade-off between customer impact and foundational work, this one gives you both: deep ownership over [... source excerpt omitted ...] ated, and shipped, and a direct line into the evolution of Liquid's post-training and audio stacks. If you care about data quality, evaluation, and making models actually work in production, this is a chance to shape how applied audio AI is done at a foundation model company. What We’re Looking For We need someone who: Takes ownership: Owns customer post-training projects end-to-end for audio workloads, from requirements through delivery and evaluation. Thinks end-to-end: Can reason across audio data pipelines, speech-text alignment, model adaptation, and evaluation as a connected system. Is pragmatic: Optimizes for model quality and customer outcomes over publications or the [... source excerpt omitted ...] audio systems excite you. You see constraints as design parameters, not blockers. The Work Act as the technical owner for enterprise audio post-training engagements. Translate customer requirements into concrete post-training specifications and workflows for LFM2.5-Audio and future audio models. Design and build function calling capabilities for audio models: training models to map spoken user intents to structured tool calls (API invocations, parameter extraction, confirmation flows). Design and execute data generation pipelines for speech-to-speech and text-to-text training, including synthetic dialogue, function calling examples, and intent-action pairs. Run supervised
Employer postings · Data from · Sources