Distributed Software Engineer
AI summary of the role
Cerebras is hiring a Distributed Software Engineer for its Cluster engineering team, which owns the software that turns thousands of wafers, servers, and switches into a reliable, observable cloud.
What you’ll do
- Build declarative, CRD-driven automation for bare-metal networking, OS, and application software across clusters of thousands of nodes.
- Develop Kubernetes operators for scheduling large inference workloads with resource locks, priority queues, and health-aware placement.
- Implement gRPC control-plane services, authorization, admission webhooks, and quota policy for a multi-tenant fleet.
- Design metrics and log pipelines with exporters for wafer-scale systems, servers, and network fabric on Prometheus and Grafana.
What you’ll bring
- 5+ years building and operating production distributed systems or infrastructure software.
- Production-quality Go and Python.
- Deep Kubernetes expertise: controllers/operators, CRDs, reconciliation, informer caches, admission webhooks, RBAC.
- Strong debugging skills across distributed systems, Linux, and networking.
Technologies
Go · Python · Kubernetes · CRD · gRPC · Prometheus · Grafana · Redfish · IPMI · gNMI · sFlow · eBPF
About Cerebras Systems
Builds wafer-scale AI processors, CS-3 systems, and cloud inference services that deliver ultra-fast training and inference without conventional multi-GPU orchestration overhead.
Growth
Source and classification
Internal deployment & tooling · Evidence for this classification:
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. The Role The Cluster engineering team owns the software that turns thousands of wafers, servers, and switches into a cloud that stays up, stays busy, and stays debuggable. We stand clusters up
More from the job description
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. The Role The Cluster engineering team owns the software that turns thousands of wafers, servers, and switches into a cloud that stays up, stays busy, and stays debuggable. We stand clusters up from bare metal, schedule training and inference workloads across the fleet, keep it healthy, and make it observable to users, operators, and increasingly to AI agents. The stack is Go and Python on Kubernetes, running both on-premise deployments and our own cloud. Responsibilities · Declarative, CRD-driven automation of bare-metal networking, OS, and application software across clusters of Cerebras systems, servers, and switches, built to reconcile thousands of nodes · Push-button cluster instal [... source excerpt omitted ...] alerting · Failure detection, HA control planes, and automated recovery, plus the CLIs, APIs, and MCP gateway that expose the fleet to users, operators, and AI agents Skills and Qualifications · 5+ years building and operating production distributed systems or infrastructure software · Production-quality Go and Python · Real Kubernetes depth: you have written or debugged controllers and operators, and you understand CRDs, reconciliation semantics, informer caches, admission webhooks, and RBAC · Strong debugging skills across distributed systems, Linux, and networking · Prometheus and Grafana as a practitioner: PromQL, exporter design, cardinality discipline, useful alerts · Str
Employer postings · Data from · Sources