Software Engineer, Infrastructure
AI summary of the role
Senior Infrastructure Engineer owning the Kubernetes-based, container-native infrastructure for an AI agent platform serving regulated financial institutions.
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
What you’ll do
- Own and evolve Kubernetes infrastructure including cluster management, service mesh, and container security policies.
- Design and implement progressive delivery pipelines with canary deployments, automated rollbacks, and deployment health validation.
- Build and maintain observability infrastructure in Datadog (dashboards, monitors, SLOs, distributed tracing).
- Drive incident response for high-severity outages and model capacity needs for low-latency AI inference.
What you’ll bring
- 8+ years in infrastructure engineering or DevOps at high-growth or hyperscale companies.
- Experience with Docker and Kubernetes including production cluster management, Helm, and service mesh technologies.
- Proven track record architecting and operating AWS (preferred), GCP, or Azure at enterprise scale.
- Experience with observability platforms, preferably Datadog (metrics, logs, APM, distributed tracing).
Technologies
Docker · Kubernetes · Helm · service mesh · AWS · Datadog · Terraform · Kustomize · Python · GitOps
About Bretton AI
Builds AI agents that automate risk, compliance, and financial-crime investigations for financial-services companies (sells to banks/fintechs like Robinhood, Mercury).
Series A · 20–50 people
Source and classification
Internal deployment & tooling · Evidence for this classification:
About Bretton AI Bretton runs AI-native operations for the financial back office. OCC, Fed and FDIC regulated banks use Bretton AI to run compliance, risk, fraud, and operations on one platform, on their own data and policies, with every output traced, cited, and audit-ready. We’ve raised over $95M from Greylock, Y Combinator, Thomson Reuters Ventures and other top tier investors. We’re based in downtown San Francisco and our team comes from world-class organizations like Google, Netflix, Stripe, Plaid, Brex, and more. The Role As a Senior Infrastructure Engineer, you will own the foundation that enables us to deploy secure, compliant AI systems at major financial institutions fighting financial crime at a massive scale. Our infrastructure is built on a modern, container-native architecture, leveraging Docker and Kubernetes to deliver consistent, auditable deployments across diverse
More from the job description
About Bretton AI Bretton runs AI-native operations for the financial back office. OCC, Fed and FDIC regulated banks use Bretton AI to run compliance, risk, fraud, and operations on one platform, on their own data and policies, with every output traced, cited, and audit-ready. We’ve raised over $95M from Greylock, Y Combinator, Thomson Reuters Ventures and other top tier investors. We’re based in downtown San Francisco and our team comes from world-class organizations like Google, Netflix, Stripe, Plaid, Brex, and more. The Role As a Senior Infrastructure Engineer, you will own the foundation that enables us to deploy secure, compliant AI systems at major financial institutions fighting financial crime at a massive scale. Our infrastructure is built on a modern, container-native architecture, leveraging Docker and Kubernetes to deliver consistent, auditable deployments across diverse customer environments. You will work directly with our largest customers—institutions serving over a billion people—to architect, automate, and harden our on-premises and cloud environments to meet the strictest regulatory and performance requirements, including SOC 2 compliance. Your work will be informed by real customer needs and will ship to everyone, so you must build enterprise-grade systems, work effectively with engineering and customer teams, understand financial services compliance, a [... source excerpt omitted ...] ode for VPCs, IAM policies, Kubernetes manifests, and private cloud deployments. Maintain and improve the infrastructure controls that support our SOC 2 compliance posture. Lead customer engagements for enterprise rollouts and mentor mid-level engineers on infrastructure best practices. What We’re Looking For Must-Haves: 8+ years in infrastructure engineering or DevOps at high-growth or hyperscale companies. Experience with Docker and Kubernetes, including production cluster management, Helm, and service mesh technologies. A proven track record of architecting and operating AWS (preferred), GCP, or Azure at an enterprise scale. Experience with observability platforms, pre [... source excerpt omitted ...] with TypeScript. Direct involvement in SOC 2 or other compliance audit preparation or remediation. Direct experience with private-cloud or on-premises deployments for regulated customers. Previous experience at startups scaling infrastructure from the early stages to the enterprise level. A background in fintech or building systems for highly regulated industries. Experience with AI/ML infrastructure and model deployment at scale. Why You’ll Love Working Here Build for Scale: You thrive at the intersection of technical leadership and customer impact, building systems that enable rapid development while maintaining the highest standards of security, compliance, and reliabi
Employer postings · Data from · Sources