Skip to content
NEXTMOVEFDE careers · United States

Senior AI/ML Engineer — LLM & Agent Stack (Customer Facing)

AI summary of the role

TrueFoundry is building an AI Gateway control plane for production agentic AI systems, and this Senior AI/ML Engineer will design and own core components for enterprise customers.

What you’ll do

  • Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads.
  • Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows.
  • Build and improve tracing, benchmarking and observability for LLMs and agents — token/cost accounting, latency p95, throughput, and correctness checks.
  • Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement.

What you’ll bring

  • 5–10 years of software engineering with substantial experience building distributed systems, infra, or ML platforms.
  • Deep practical experience integrating and deploying LLMs in production (RAG, retrieval, embeddings pipelines).
  • Hands-on experience with agent orchestration frameworks (LangGraph / LangChain or custom agent runtimes) and stateful workflow design.
  • Strong systems knowledge: Kubernetes, container orchestration, service meshes, and performance tuning.

Technologies

LLM · LangGraph · LangChain · Kubernetes · RAG · MCP · vector databases · observability · agent orchestration · OpenAI · Anthropic · Google

Source and classification

Production engineering · Evidence for this classification:

yet. You can't just duct-tape together a few API calls and call it production-ready. You need a control plane that handles: Intelligent routing with observability, cost policies, and fallback logic Centralized tool and MCP server management with security and lifecycle controls Agent orchestration with governance and guardrails A unified compute layer to run self-hosted models, custom tools, and agents AI Gateway is the control plane: five composable components (Prompts, LLM Gateway, MCP Gateway, Guardrails, Agent Gateway) that handle routing, orchestration, and governance. We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform. Role summary You’ll design and own core components that enable enterprise customers to run production agentic AI safely and efficiently on
More from the job description

About TrueFoundry Every production AI system, whether it's powering customer support, writing code, analyzing financial data, or diagnosing medical conditions, needs the same foundational infrastructure.A way to route between models. A way to manage tools and integrate them securely. A way to orchestrate agents and enforce governance. A unified compute layer to run it all. That infrastructure layer is being built right now. We are looking for a Senior AI/ML Engineer — LLM & Agent Stack (Customer-Facing) to join the team. The Problem We're Solving Companies are moving beyond simple chatbots to production agentic systems. These systems route between OpenAI, Anthropic, Google, and self-hosted models. They integrate dozens of tools via protocols like MCP. They orchestrate multi-agent workflows where agents coordinate with other agents. The infrastructure to support this doesn't exist yet. You can't just duct-tape together a few API calls and call it production-ready. You need a control plane that handles: Intelligent routing with observability, cost policies, and fallback logic Centralized tool and MCP server management with security and lifecycle controls Agent orchestration with governance and guardrails A unified compute layer to run self-hosted models, custom tools, and agents AI Gateway is the control plane: five composable components (Prompts, LLM Gateway, MCP Gat [... source excerpt omitted ...] that handle routing, orchestration, and governance. We're Series A, backed by Intel Capital and Sequoia. Companies like CVS, Mastercard, Siemens, Paytm, Synopsys, and Zscaler run production AI workloads on our platform. Role summary You’ll design and own core components that enable enterprise customers to run production agentic AI safely and efficiently on TrueFoundry. This includes building robust orchestration for multi-step agents (graph/stateful workflows), model/routing logic, observability and policy enforcement (cost, data residency, rate limiting), and integrating upstream tooling like LangGraph, LangChain, vector stores, and specialized LLM runtimes. This is a customer- [... source excerpt omitted ...] code in their environments. What you’ll do Architect and implement scalable agent orchestration patterns (graph-based executors, state management, multi-agent coordination) for production workloads. Own critical integrations: model adapters, LLM gateway hooks, vector DBs, tools & external APIs, and the platform’s LLMops flows. Build and improve tracing, benchmarking and observability for LLMs and agents — token/cost accounting, latency p95, throughput, and correctness checks. Drive design for safety/guardrails: moderation hooks, human-in-the-loop checkpoints, replayable audit trails and policy enforcement. Mentor junior engineers, run design reviews, and improve engineering

How jobs are selected

Employer postings · Data from · Sources