Software Engineer 3
Job description
The full responsibilities and requirements are on the employer’s site.
Read the job description ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
The Agent Research and Tooling team, part of MongoDB's AI Builder Experience organization, owns the platform layer around agents: how teams author, distribute, evaluate, monitor, and improve agent skills and agent behavior. We are hiring a software engineer to build and maintain the tooling, evaluation systems, and quality gates behind MongoDB's agent skills. This is a software engineering role at the intersection of developer tooling, applied AI, and software quality. You will take loosely defined agent and tooling problems, break them into workable plans, and ship durable internal systems: command-line tools, reusable libraries, evaluation harnesses, and CI workflows. This role is open to remote work in the US or can be based out of any of our US offices. What you'll do Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them Design
More from the job description
The Agent Research and Tooling team, part of MongoDB's AI Builder Experience organization, owns the platform layer around agents: how teams author, distribute, evaluate, monitor, and improve agent skills and agent behavior. We are hiring a software engineer to build and maintain the tooling, evaluation systems, and quality gates behind MongoDB's agent skills. This is a software engineering role at the intersection of developer tooling, applied AI, and software quality. You will take loosely defined agent and tooling problems, break them into workable plans, and ship durable internal systems: command-line tools, reusable libraries, evaluation harnesses, and CI workflows. This role is open to remote work in the US or can be based out of any of our US offices. What you'll do Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression Create CLIs, libraries, and MCP integrations that other repositories adopt and that r [... source excerpt omitted ...] m into reusable improvements Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams Examples of the problems you'll solve How can tests verify an agent tool's structured result when item order may vary, but counts, required fields, and values must remain correct How can a CI gate flag unsafe instructions in an agent skill without treating every neutral mention as an incident or letting cautionary wording hide a real instruction How can an evaluation suite show whether a skill improves answers over a baseline and give authors enough signal to improve it What we're looking for 2+ years of experience build [... source excerpt omitted ...] Python, JavaScript/TypeScript, Java, or C# Experience with GitHub Actions security, secret handling, static rule engines, or policy enforcement Experience moving prototypes into production What success looks like In your first year, you will: Ship tooling that makes agent skills or developer workflows easier to test, review, and adopt Improve the quality and interpretability of evaluations, not just their count Convert recurring manual work and fragile scripts into documented, reusable automation Make security, correctness, and operational trade-offs explicit in the designs you ship Earn adoption from partner teams through clear interfaces and reliable CI Own projects ind
Employer postings · Data from · Sources