Member of Technical Staff - Post Training, Applied (Vision)
AI summary of the role
Own applied post-training for vision-language models end-to-end for enterprise customers while contributing to Liquid's core multimodal model development.
What you’ll do
- Act as technical owner for enterprise customer VLM post-training engagements, translating requirements into concrete specifications and workflows.
- Design and execute visual data generation, filtering, and quality assessment processes including image-text pair curation, annotation pipelines, and synthetic data generation.
- Run supervised fine-tuning, preference alignment, and reinforcement learning workflows for vision-language models.
- Design task-specific evaluations for visual understanding, grounding, OCR, document parsing, and other multimodal capabilities.
What you’ll bring
- Hands-on experience with data generation and evaluation for VLM or multimodal post-training.
- Experience training or fine-tuning vision-language models using SFT, preference alignment, and/or RL.
- Strong intuition for visual data quality, annotation design, and multimodal evaluation.
- Familiarity with vision encoders, image-text architectures, and how visual representations interact with language model backbones.
Technologies
VLM · vision-language models · SFT · preference alignment · reinforcement learning · vision encoders · image-text architectures · OCR · document parsing · synthetic data generation
About Liquid AI
Efficient general-purpose AI systems optimized for on-device deployment across data centers and edge hardware, enabling low-latency, privacy-preserving enterprise AI.
Series A
Source and classification
Production engineering · Evidence for this classification:
About Liquid AI Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there. The Opportunity This is a rare chance to sit at the intersection of frontier vision-language models and real-world deployment. You'll own applied post-training work for VLMs end-to-end for some of the world's largest enterprises, while still contributing directly to Liquid's core multimodal model development. Unlike most roles that force a trade-off between customer impact and foundational work, this role gives you both: deep ownership over how vision-language
More from the job description
About Liquid AI Spun out of MIT CSAIL, we build general-purpose AI systems that run efficiently across deployment targets, from data center accelerators to on-device hardware, ensuring low latency, minimal memory usage, privacy, and reliability. We partner with enterprises across consumer electronics, automotive, life sciences, and financial services. We are scaling rapidly and need exceptional people to help us get there. The Opportunity This is a rare chance to sit at the intersection of frontier vision-language models and real-world deployment. You'll own applied post-training work for VLMs end-to-end for some of the world's largest enterprises, while still contributing directly to Liquid's core multimodal model development. Unlike most roles that force a trade-off between customer impact and foundational work, this role gives you both: deep ownership over how vision-language models are adapted, evaluated, and shipped, and a direct line into the evolution of Liquid's multimodal post-training stack. If you care about visual understanding, data quality, evaluation, and making VLMs actually work in production, this is a chance to shape how applied multimodal AI is done at a foundation model company. What We're Looking For We need someone who: Takes ownership: Owns VLM post-training projects end-to-end, from customer requirements through delivery and evaluation. Thinks [... source excerpt omitted ...] ment, and evaluation as a single system. Is pragmatic: Optimizes for model quality and customer outcomes over publications or theory. Communicates clearly: Can translate between customer needs and internal technical teams, and push back when needed. The Work Act as the technical owner for enterprise customer VLM post-training engagements. Translate customer requirements into concrete multimodal post-training specifications and workflows. Design and execute visual data generation, filtering, and quality assessment processes, including image-text pair curation, annotation pipelines, and synthetic data generation for visual tasks. Run supervised fine-tuning, preference alignm [... source excerpt omitted ...] nding, document understanding, OCR, or video understanding tasks. Experience contributing to shared or general-purpose multimodal post-training infrastructure. Prior exposure to customer-facing or applied ML delivery environments. Familiarity with alignment or RL techniques beyond basic supervised fine-tuning in the multimodal setting. What Success Looks Like (Year One) Independently owns and delivers enterprise VLM post-training projects with minimal oversight. Is trusted by customers as the technical owner, demonstrating strong judgment and delivery quality on multimodal workloads. Has made durable contributions to Liquid's general-purpose multimodal post-training pipeli
Employer postings · Data from · Sources