AI Engineer, Evaluation
Technologies
Python · LLM-based graders · evaluation pipelines · golden test suites · regression suites · AI-assisted test generation · prompt design · agent logic · model selection
About Distyl AI
Builds AI-native enterprise systems for Fortune 500 operators, combining forward-deployed teams with Distillery to turn domain knowledge into auditable production workflows.
Series B
Job description
The full responsibilities and requirements are on the employer’s site.
Open application page ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
Partners, Khosla Ventures, Coatue, DST Global, and the board-members of 20+ F500s. What We Are Looking For At Distyl, we build AI systems using Evaluation-Driven Development—an approach where evaluation is not an afterthought, but the primary mechanism for iterating, improving, and trusting AI behavior in production. AI Evaluation Engineers focus on designing and implementing the evaluation systems that drive this process. They are hands-on engineers who write production Python code, build evaluation pipelines, and use structured signals to guide system design, prompt iteration, and deployment decisions for real customer-facing AI systems. This role is for engineers who believe that AI systems only improve when measurement is tightly coupled to development—and who want to apply that philosophy directly to systems that matter. Key Responsibilities Design and implement evaluationHow jobs are selected
Employer postings · Data from · Sources