Skip to content
NEXTMOVEFDE careers · United States

Sr. Manager, Core Infrastructure Engineering

AI summary of the role

Lead a team of core infrastructure engineers providing white-glove engineering and operational support for OCI's largest GPU/AI/ML customers.

What you’ll do

  • Lead, mentor, and develop a team of Core Infrastructure Engineers supporting OCI's largest GPU/AI/ML customers.
  • Drive design, development, testing, and deployment of automated GPU cluster deployment tools (like AWS Parallel Cluster, Azure Cycle Cloud) with Slurm and/or OKE.
  • Act as technical liaison between customers, core engineering teams, and support; build collaborative relationships with OCI Services and sales.
  • Optimize infrastructure performance via tuning, resource utilization, caching, and data pre-processing; troubleshoot scalability and reliability issues.

What you’ll bring

  • 7+ years of senior software engineering leadership or related experience.
  • Proven experience building and managing distributed/cloud software engineering solutions.
  • Experience with Ansible, Terraform, Python, containerization (Docker, Kubernetes), and orchestration tools.
  • Strong Linux skills (Oracle Linux/RHEL/CentOS, Ubuntu, Debian) including system administration and shell scripting.

Technologies

GPU clusters · Slurm · Oracle Kubernetes Engine (OKE) · Ansible · Terraform · Python · Docker · Kubernetes · Linux · HPC

About Oracle

Oracle sells database software, enterprise applications (ERP/HCM/CX), and Oracle Cloud Infrastructure to large enterprises and governments.

Public · 5000+ people

Source and classification

Production engineering · Evidence for this classification:

Oracle Cloud Infrastructure (OCI) is building some of the world's largest and most advanced GPU clusters to power the next generation of AI. The Strategic Customer Engineering (SCE) Core Infrastructure team—also known as the AI/ML Forward Deployed Infrastructure Engineering team—provides white-glove engineering and operational support to OCI's most strategic AI/ML infrastructure customers. As a trusted partner to our customers, we play a critical role in designing, deploying, operating, and optimizing the infrastructure that powers some of the largest and most demanding GPU and AI/ML environments in the world. Our team works closely with customers and internal engineering organizations to ensure exceptional reliability, performance, and scalability for mission-critical AI workloads. We are seeking an experienced Core Infrastructure Engineering Leader to lead a high-performing
How jobs are selected

Employer postings · Data from · Sources