Senior Site Reliability Engineer, Colorado Springs
AI summary of the role
Onebrief is hiring a Senior Site Reliability Engineer to join their Infrastructure & Security team, working on-site at customer locations in Colorado Springs.
What you’ll do
- Design, implement, and manage monitoring, logging, and alerting stack (Prometheus, Loki, Alloy, Grafana).
- Define, measure, and own alerting feeding into SLIs and SLOs.
- Act as incident responder and lead blameless post-mortems / After Action Reviews.
- Design, build, and manage secure Kubernetes clusters and cloud/on-prem environments using Terraform and Ansible.
What you’ll bring
- Active Top Secret clearance.
- 5+ years in Platform, DevOps, or Site Reliability Engineering with infrastructure and operations focus.
- Proficiency with Terraform, Ansible, Kubernetes, and CI/CD pipelines (GitLab CI/CD, Jenkins, GitHub Actions).
- Scripting proficiency in Python, Go, or Bash.
Technologies
Terraform · Ansible · Kubernetes · GitLab CI/CD · Jenkins · GitHub Actions · Python · Go · Bash · AWS GovCloud · Prometheus · Grafana
About One Brief
AI-powered command operating system for military staffs, unifying planning, collaboration, decision intelligence, and wargaming across classified and coalition networks.
Series D · 200–500 people
Employer postings · Data from · Sources