Skip to content
NEXTMOVEFDE careers · United States

Engineering Manager, Kernel Reliability

AI summary of the role

Lead the on-field Kernel Reliability team at Cerebras, owning the technical vision and roadmap for kernel-centric reliability of internal and customer-facing systems.

What you’ll do

  • Provide hands-on technical leadership, owning the technical vision and roadmap for kernel-centric reliability of internal and customer-facing systems
  • Assist System and Cluster Operations teams on reducing system and service downtime after failure by providing tooling and manual intervention for failure analysis and diagnostic
  • Work with the Debug Team to enhance debug tools with the goal of speeding up failure analysis
  • Collaborate with SW teams to improve the software stack, including Kernels, to improve on-field debugging and failure analysis

What you’ll bring

  • 6+ years in software engineering, with 3+ years leading teams in SW/HW reliability, debug, diagnostic, failure analysis or related fields
  • Expertise in parallel and distributed programming (message passing, multicore, GPU, embedded, etc.)
  • Expertise in debug and diagnostic tool development or expert usage (debuggers, core dump handling, code sanitizers, etc.)
  • Experience debugging distributed and parallel applications (deadlocks, livelocks, race conditions, etc.)

Technologies

wafer-scale architecture · kernel reliability · debuggers · core dump handling · code sanitizers · parallel programming · distributed programming · GPU · ASIC · HW architecture

About Cerebras Systems

Builds wafer-scale AI processors, CS-3 systems, and cloud inference services that deliver ultra-fast training and inference without conventional multi-GPU orchestration overhead.

Growth

Source and classification

Deployment team leadership · Evidence for this classification:

GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. The Role We're looking for a deeply technical, hands-on engineering leader for our on-field Kernel Reliability team. You will lead a high performing team to tackle a critical challenge: improving the reliability of our advanced compute clusters and the underlying inference, training, and internal production services. In this role, you'll set the technical vision while staying close to the code and designing solutions that will scale to our exponentially growing system production and software service offerings. If you have proven expertise in software or hardware reliability, diagnostic tool building, or failure analysis and debugging, we want to hear
More from the job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly run large-scale ML applications, without the hassle of managing hundreds of GPUs or TPUs. Cerebras' current customers include top model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. The Role We're looking for a deeply technical, hands-on engineering leader for our on-field Kernel Reliability team. You will lead a high performing team to tackle a critical challenge: improving the reliability of our advanced compute clusters and the underl [... source excerpt omitted ...] rnal production services. In this role, you'll set the technical vision while staying close to the code and designing solutions that will scale to our exponentially growing system production and software service offerings. If you have proven expertise in software or hardware reliability, diagnostic tool building, or failure analysis and debugging, we want to hear from you. Responsibilities Provide hands-on technical leadership, owning the technical vision and roadmap for the kernel-centric reliability of our internal and customer-facing systems Assist System and Cluster Operations teams on reducing system and service downtime after failure by providing tooling and manual interve [... source excerpt omitted ...] res with reliability and ease of debug in mind Lead, mentor, and grow a high-caliber team of engineers, fostering a culture of technical excellence and rapid execution. Skills & Qualifications 6+ years in software engineering, with 3+ years leading teams in SW/HW reliability, debug, diagnostic, failure analysis or related fields Expertise in parallel and distributed programming (message passing, multicore, GPU, embeded, etc.), debug and diagnostic tool development or expert usage (debuggers, core dump handling, code sanitizers, etc.), experience debugging distributed and parallel applications (deadlocks, livelocks, race conditions, etc.), deep understanding of computer architecture

How jobs are selected

Employer postings · Data from · Sources