Engineering Manager, Data Infrastructure
AI summary of the role
Lead a team of Software Engineers and SREs responsible for the infrastructure powering CoreWeave's data platform, including compute engines, orchestration frameworks, and storage layers.
What you’ll do
- Lead a team of Software Engineers and Site Reliability Engineers responsible for the infrastructure that powers CoreWeave's data platform
- Own the reliability, scalability, and performance of core systems such as compute engines, orchestration frameworks, and storage layers
- Partner closely with Data Engineering teams and cross-functional groups including Production Engineering, Developer Experience, Security Engineering, and IT Operations
- Balance people leadership with deep technical ownership, stepping in hands-on when needed to support critical initiatives
What you’ll bring
- 7+ years of experience in software engineering, infrastructure engineering, or data platform engineering roles
- 2+ years of experience managing engineering teams, including hiring, coaching, performance management, and career development
- Strong hands-on experience operating and scaling data platform infrastructure (e.g., Spark, Airflow, Iceberg, StarRocks) in production environments
- Deep expertise in Kubernetes and containerized software development, including cluster design, operations, and scaling in production environments
Technologies
Spark · Airflow · Iceberg · StarRocks · Kubernetes · Python · Java · Go · Rust
Source and classification
Internal deployment & tooling · Evidence for this classification:
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com. What You’ll Do: The Platform & Infrastructure Engineering team in the Data Infrastructure organization is responsible for the performance, reliability, scalability, and security of the company’s data platform. The team builds and operates the foundational systems that power ingestion, transformation, analytics, and AI workloads at scale. This includes ownership of
More from the job description
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at www.coreweave.com. What You’ll Do: The Platform & Infrastructure Engineering team in the Data Infrastructure organization is responsible for the performance, reliability, scalability, and security of the company’s data platform. The team builds and operates the foundational systems that power ingestion, transformation, analytics, and AI workloads at scale. This includes ownership of the underlying infrastructure for orchestration, compute, and storage systems that enable data engineering teams to build and deliver data products. We operate with production-grade discipline, supporting mission-critical services with stringent uptime requirements and a focus on automation, observability, and resilience. About the role: As an Engineering Manager, you will lead a team of Software Engineers and Site Reliability Engineers responsible for the infrastructure that powers CoreWeave’ [... source excerpt omitted ...] ore systems such as compute engines, orchestration frameworks, and storage layers. You’ll partner closely with Data Engineering teams, as well as cross-functional groups including Production Engineering, Developer Experience, Security Engineering, and IT Operations to ensure the platform is robust, secure, and easy to operate. This role balances people leadership with deep technical ownership, including stepping in hands-on when needed to support critical initiatives. Who You Are: 7+ years of experience in software engineering, infrastructure engineering, or data platform engineering roles 2+ years of experience managing engineering teams, including hiring, coaching, performance [... source excerpt omitted ...] (e.g., OKRs) and holding teams accountable to outcomes Strong hands-on experience operating and scaling data platform infrastructure (e.g., Spark, Airflow, Iceberg, StarRocks) in production environments Deep expertise in Kubernetes and containerized software development, including cluster design, operations, and scaling in production environments Experience building and operating distributed systems with high availability and performance requirements, including SLOs and incident management Strong understanding of data platform architecture (compute, orchestration, storage) and experience driving reliability, performance, and cost optimization at the platform level Ability to c
Employer postings · Data from · Sources