Skip to content
NEXTMOVEFDE careers · United States

Cloud Infrastructure Engineer

AI summary of the role

Build and maintain scalable infrastructure using Terraform, Kubernetes, and CI/CD for Braintrust's internal AWS environment and customer self-hosting across AWS, Azure, GCP.

What you’ll do

  • Build/maintain Terraform modules for internal and customer infra
  • Support customers in Slack for self-hosting and troubleshooting
  • Own CI/CD pipeline improvements for faster, safer releases
  • Scale observability with logs, metrics, dashboards, alerts

What you’ll bring

  • 5+ years DevOps/SRE/Infrastructure Engineering
  • Deep Terraform + AWS experience
  • Strong Kubernetes skills
  • Proficient in Python/Typescript/Go

Technologies

Terraform · Kubernetes · AWS · Azure · GCP · CI/CD · observability

About Braintrust

AI observability platform for teams shipping LLM apps and agents, combining tracing, evals, prompt iteration, and model gateway infrastructure in one workflow.

Series B · 100–200 people

Source and classification

Implementation & delivery · Evidence for this classification:

you’ll contribute across our internal AWS environment and help customers deploy our stack in AWS, Azure, and GCP. What you’ll do Build and maintain Terraform modules for both internal infrastructure and customer deployments Work directly with customers in Slack to support self-hosting and troubleshoot infrastructure issues. Build tools to make it easier for them to support themselves. Own and improve our CI/CD pipeline: reduce build times, improve failure visibility, and enable safer, faster releases Centralize and scale observability - including logs, metrics, dashboards, and alerts Partner with engineering teams to build and evolve a secure, developer-friendly infrastructure platform Support multi-cloud deployment patterns (AWS primarily, with Azure and GCP support for enterprise customers) Implement tools and automation to improve deployment, rollback, and infrastructure
More from the job description

About the company Braintrust is the agent observability platform. By actively applying intelligence to agent traces and automatically surfacing the most critical patterns, Braintrust gives teams the visibility to understand how agents behave in production and the tools to improve them. Teams at Notion, Stripe, Box, OpenAI, and Cloudflare use Braintrust to trace their agents, find the issues in their observability data, and run evals that tell them how to improve. About the role We’re looking for a Cloud Infrastructure Engineer to help us build reliable, scalable infrastructure and give developers a world-class platform to ship code with speed and confidence. You’ll lead efforts across Terraform, Kubernetes, CI/CD, observability, and support, and play a key role in how we scale Braintrust both internally and for customers self-hosting our platform. This is a high-impact role where you’ll contribute across our internal AWS environment and help customers deploy our stack in AWS, Azure, and GCP. What you’ll do Build and maintain Terraform modules for both internal infrastructure and customer deployments Work directly with customers in Slack to support self-hosting and troubleshoot infrastructure issues. Build tools to make it easier for them to support themselves. Own and improve our CI/CD pipeline: reduce build times, improve failure visibility, and enable safer, faster releases Centralize and scale observability - including logs, metrics, dashboards, and alerts Partner with engineering teams to build and evolve a secure, developer-friendly infrastructure platform Support multi-cloud deployment patterns (AWS primarily, with Azure and GCP support for enterprise customers) Implement tools and automation to improve deployment, rollback, and infrastructure reliability Ideal candidate credentials 5+ years of experience in DevOps, SRE, or Infrastructure Engineering roles Deep experience with Terraform and at least one major cloud provider (AWS strongly preferred) Strong Kubernetes skills: deploying, debugging, and scaling real workloads Proficient in scripting or programming (Python, Typescript, or Go) Experience supporting production systems and responding to incidents Comfortable working directly with customers in a support or deployment context Bonus: experience with multi-cloud environments or self-hosted enterprise software Benefits include Medical, dental, and vision insurance Daily lunch, snacks, and beverages Flexible time off Competitive salary and equity Wifi & cellphone stipend Equal opportunity Braintrust is an equal opportunity employer. All applicants will be considered for employment without attention to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran or disability status.

How jobs are selected

Employer postings · Data from · Sources