Senior/Staff FDE - Synthetic Data Generation
AI summary of the role
Snorkel AI is hiring a Senior/Staff Forward Deployed Engineer to lead synthetic data generation engagements with leading AI labs and enterprises.
What you’ll do
- Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines for complex AI use cases
- Translate model objectives, failure modes, and data gaps into synthetic data strategies and experiments
- Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets
- Build automated evaluators and measurement frameworks to assess correctness, relevance, diversity, and coverage
What you’ll bring
- 5+ years of experience in ML engineering, data science, applied AI, forward deployed engineering, or similar
- Strong Python skills and experience building production data/ML systems with Docker and cloud platforms (AWS, GCP, Azure)
- Hands-on experience with LLMs and building model-based applications and data workflows with the modern GenAI/LLM stack
- Experience building synthetic data, data augmentation, or model-generated training and evaluation datasets
Technologies
Python · Docker · AWS · GCP · Azure · LLM · LLM-as-a-judge · RLVR · agentic environments · synthetic data
About Snorkel AI
Data-centric AI platform and expert-data provider that helps enterprises and frontier labs build, evaluate, and tune specialized models and agents.
Series D
Source and classification
Production engineering · Evidence for this classification:
Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives. In this role, you will lead the technical execution of complex customer engagements where synthetic data is used to improve model training, evaluation, and performance. You will translate ambiguous model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to continuously improve data quality and downstream model outcomes. You will work across the full delivery lifecycle—from technical discovery and solution design through implementation, evaluation, and production delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements. Main Responsibilities Synthetic Data Generation & Evaluation Design
More from the job description
About Snorkel At Snorkel, we believe meaningful AI doesn’t start with the model, it starts with the data. We’re on a mission to help enterprises transform expert knowledge into specialized AI at scale. The AI landscape has gone through incredible changes since 2015, when Snorkel started as a research project in the Stanford AI Lab, to the generative AI breakthroughs of today. But one thing has remained constant: the data you use to build AI is the key to achieving differentiation, high performance, and production-ready systems. We work with some of the world’s largest organizations to empower scientists, engineers, financial experts, product creators, journalists, and more to build custom AI with their data faster than ever before. Excited to help us redefine how AI is built? Apply to be the newest Snorkeler! About the Role Snorkel AI is hiring a Forward Deployed Engineer focused on Synthetic Data Generation to partner with leading AI labs and enterprises on their most critical AI initiatives. In this role, you will lead the technical execution of complex customer engagements where synthetic data is used to improve model training, evaluation, and performance. You will translate ambiguous model and data challenges into effective data strategies, build scalable generation and evaluation pipelines, and use experimentation to continuously improve data quality and downstream mo [... source excerpt omitted ...] uction delivery. You will also identify patterns across engagements and turn successful approaches into reusable capabilities, technical standards, and product improvements. Main Responsibilities Synthetic Data Generation & Evaluation Design and build scalable synthetic data generation, transformation, filtering, and evaluation pipelines for complex AI use cases Translate model objectives, failure modes, and data gaps into synthetic data strategies, experiments, and technical specifications Develop LLM- and ML-assisted workflows to generate high-quality training and evaluation datasets across targeted behaviors, domains, and edge cases Build automated evaluators, quality checks, [... source excerpt omitted ...] requirements Design and run experiments to measure the impact of synthetic data on downstream model performance and iteratively improve generation approaches Package and deliver production-grade datasets with standardized formats, quality assurance, and clear documentation Forward Deployed Engineering & Customer Partnership Lead technical workstreams from initial solution design through production delivery, navigating ambiguity and making sound technical decisions Build, refine, and iterate on solutions that address customer needs, incorporating feedback to ensure the delivered work provides tangible value Rapidly prototype and productionize solutions across models, data pipe
Employer postings · Data from · Sources