Skip to content
NEXTMOVEFDE careers · United States

Staff Site Reliability Engineer

AI summary of the role

Staff SRE at a remote-first insurtech defining reliability standards, SLOs, and observability across a growing engineering org.

What you’ll do

  • Define reliability standards, patterns, and reference implementations for SLOs, error budgets, instrumentation, alerting, and incident practice.
  • Drive end-to-end SLOs for claims journeys and SLA compliance reporting.
  • Lead consolidation onto OpenTelemetry as the single instrumentation standard.
  • Build reporting that surfaces risk and coverage gaps across the platform.

What you’ll bring

  • 10+ years of site reliability, production, or platform engineering experience.
  • Deep expertise with OpenTelemetry, AWS, Kubernetes, PostgreSQL, and modern incident tooling.
  • Proven experience designing and landing SLOs and error budgets that teams actually use.
  • Strong ability to lead through influence and mentor engineers.

Technologies

OpenTelemetry · AWS · Kubernetes · PostgreSQL · SLOs · error budgets · incident management · Claude · Codex · Cursor

Source and classification

Internal deployment & tooling · Evidence for this classification:

rewarding. As a Staff Site Reliability Engineer, you'll define the standards, patterns, and platforms that let every engineering team at Assured run their services reliably — and partner directly with the teams adopting them. Some of the Problems You'll Solve: 🎯 Define what reliability means across a growing engineering organization. Set the standards, patterns, and reference implementations teams adopt for SLOs, error budgets, instrumentation, alerting, and incident practice — and make them easy enough to adopt that teams actually do. 📊 Measure reliability the way customers experience it. Move us beyond component-level availability targets to end-to-end SLOs for the claims journeys insurers and claimants depend on, spanning many services and teams, and extend that into how we measure and report SLA compliance. 🔭 Unify a fragmented observability picture. Help drive our
More from the job description

Assured is on a mission to modernize insurance. Claims processing (i.e. should we pay this claim?), while often overlooked, is the foundation of the entire industry. It’s currently highly manual, involving phone calls, faxes, and gut instinct, costing tens of billions of dollars a year. We can do better. At Assured, we provide large insurers with the software solutions they need to win in a modern, technology-driven world. From self-service claim-filing software to backend fraud detection, we’re the engine that powers claims processing for some of the largest insurers in the world. The challenges we face are deep and diverse, from creating digital experiences that provide comfort and clarity to claimants at their most stressed and vulnerable to orchestrating large-scale ML-driven decision-making on billions of dollars of claims payments, life at Assured is dynamic, collaborative, and rewarding. As a Staff Site Reliability Engineer, you'll define the standards, patterns, and platforms that let every engineering team at Assured run their services reliably — and partner directly with the teams adopting them. Some of the Problems You'll Solve: 🎯 Define what reliability means across a growing engineering organization. Set the standards, patterns, and reference implementations teams adopt for SLOs, error budgets, instrumentation, alerting, and incident practice — and make them [... source excerpt omitted ...] t, respond to, and learn from failure — incident tooling and automation, post-incident review practice, and making sure action items get closed rather than quietly aging out. How You'll Make an Impact: 🤝 Help other teams run their own systems well. Embed with product teams for a period at a time: specify what reliability looks like for their most critical paths, help them build it, then hand it over with them as the durable owner. 🔍 Find the risk before it finds us. Surface coverage gaps, weak signals, and single points of failure across the platform, and make the case for fixing them before they become incidents. 📟 Support engineering when things go wrong. Share an inte [... source excerpt omitted ...] re effectively. Use tools such as Claude, Codex, Cursor, and similar platforms to support tooling development, incident analysis, debugging, documentation, and operational work. You'll Probably Thrive Here If You: 📐 Have deep site reliability and systems expertise. You bring 10+ years of site reliability, production, or platform engineering experience, ideally within SaaS platforms or high-scale distributed systems environments. 🛠 Work across a modern reliability and observability stack. You're comfortable with OpenTelemetry, metrics, traces and logs, AWS, Kubernetes, PostgreSQL, and modern incident tooling. Experience with every tool isn't required — we value strong fund

How jobs are selected

Employer postings · Data from · Sources