Skip to content
NEXTMOVEFDE careers · United States

Staff Applied Scientist - Agentic Interfaces

AI summary of the role

Define and build the evaluation and measurement systems for Datadog's AI agent integrations, covering answer quality, tool-selection accuracy, retrieval relevance, latency, token cost, and end-to-end agent success.

No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.

What you’ll do

  • Own the evaluation strategy for Datadog's AI agent integrations, defining offline and online metrics across quality, cost, single-turn, and trajectory-level dimensions.
  • Build eval datasets, golden traces, and regression harnesses that catch quality changes before they hit customers, making them reusable across teams.
  • Drive measurable improvements to retrieval relevance, tool-selection accuracy, and context efficiency in partnership with AI engineers.
  • Run applied research on agent–data interaction: tool selection under large catalogs, multi-turn evaluation, grounding and hallucination control, cost/quality tradeoffs.

What you’ll bring

  • BS/MS/PhD in a scientific field or equivalent experience.
  • 10+ years of relevant engineering or applied science experience, including time as a technical lead.
  • Proven track record of leading ML or GenAI initiatives in a product-driven environment from research through production.
  • Significant experience with evaluation, experimentation, or measurement of ML systems at scale.

Technologies

MCP Server · Bits SRE · Bits Assistant · Bits Dev Agent · Claude Code · Cursor · Copilot · GenAI · ML evaluation · retrieval

About Datadog

Cloud observability SaaS unifying metrics, traces, logs, and security signals for engineering and SRE teams running cloud-native stacks.

Public

Source and classification

Internal deployment & tooling · Evidence for this classification:

Team description At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the platform that connects these agents to Datadog: the MCP Server, the tools and retrieval surfaces agents call into, and — critically — the evaluation systems that tell us whether an agent's experience on Datadog data is actually getting better over time. This role is about that last piece. We're hiring a Staff Applied Scientist to define what "good" means for an Agentic interface at Datadog and to build the measurement systems that make it true. "Good" isn't one number — it spans answer quality, tool-selection accuracy, retrieval relevance, latency, token cost, and end-to-end agent success on
More from the job description

Team description At Datadog, AI agents are becoming first-class consumers of observability, security, and software delivery data — from third-party coding agents like Claude Code, Cursor, and Copilot, to our own Bits SRE, Bits Assistant, and Bits Dev Agent. The Agentic Interfaces team owns the platform that connects these agents to Datadog: the MCP Server, the tools and retrieval surfaces agents call into, and — critically — the evaluation systems that tell us whether an agent's experience on Datadog data is actually getting better over time. This role is about that last piece. We're hiring a Staff Applied Scientist to define what "good" means for an Agentic interface at Datadog and to build the measurement systems that make it true. "Good" isn't one number — it spans answer quality, tool-selection accuracy, retrieval relevance, latency, token cost, and end-to-end agent success on real customer workflows. You'll design the evals, build the datasets, define the metrics, and partner with the AI engineers on the team to land the platform that lets every product group at Datadog ship integrations that are demonstrably better release over release. The space is full of open research questions. How do you evaluate an agent end-to-end when the trajectory is non-deterministic? How do you score tool selection when the tool catalog has hundreds of entries and grows weekly? How do you b [... source excerpt omitted ...] f those are the problems you want to spend your time on, come build this with us. Datadog values people from all walks of life. We understand not everyone will meet all the above qualifications on day one. That's okay. If you’re passionate about technology and want to grow your skills, we encourage you to apply. What You’ll Do: Own the evaluation strategy for Datadog's AI agent integrations. Define the metrics — offline and online, quality and cost, single-turn and trajectory-level — that the team and the broader organization optimize against. Build the eval datasets, golden traces, and regression harnesses that catch quality changes before they hit customers, and make those assets [... source excerpt omitted ...] g or applied science experience, including time as a technical lead. Proven track record of leading ML or GenAI initiatives in a product-driven environment, from research through production. Significant experience with evaluation, experimentation, or measurement of ML systems at scale. You bring a strong product mindset and are comfortable driving initiatives across cross-functional teams. You thrive in ambiguity and can make sound technical calls when the path isn’t yet defined. Benefits and Growth: New hire stock equity (RSUs) and employee stock purchase plan (ESPP) Continuous professional development, product training, and career pathing An inclusive company culture, giv

How jobs are selected

Employer postings · Data from · Sources