AI Infrastructure Engineer
AI summary of the role
This is a pre-sales technical deployment role for vCluster's GPU infrastructure platform, working directly with neocloud and AI Factory customers from bare metal to production.
What you’ll do
- Lead end-to-end technical deployments for GPU neocloud and AI Factory customers from bare metal to validated vCluster environment
- Configure and troubleshoot bare metal GPU node infrastructure including CNI, GPU Operator, distributed storage, and RDMA/InfiniBand
- Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s
- Document reusable playbooks and deployment architectures to scale customer onboarding
What you’ll bring
- 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or high-complexity environments
- Practical knowledge of NVIDIA GPU Operators, CUDA tooling, and systems-level configuration for GPU nodes
- Deep understanding of CNI plugins, overlay networks, load balancing, and connectivity diagnosis
- Experience with persistent volume configuration, CSI drivers, and distributed systems like Ceph, Rook, Weka, or Longhorn
Technologies
Kubernetes · NVIDIA GPU Operators · CUDA · CNI · RDMA · InfiniBand · Ceph · Rook · Weka · Longhorn · Bash · Python
About vcluster
Kubernetes virtualization platform that runs isolated virtual clusters inside one real cluster, used by enterprises and AI neoclouds for multi-tenancy and cost reduction.
Series A · 50–100 people
Source and classification
Implementation & delivery · Evidence for this classification:
As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer. vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen. As an AI Infrastructure Engineer, your role will include: Lead Technical Deployments: Drive end-to-end
More from the job description
As vCluster’s AI Infrastructure Specialist, you will work directly with customers at the earliest and most critical stage of their journey: from bare metal GPU nodes through to a production-ready deployment. This is not a traditional professional services role; you operate pre-sale as part of a proof of value engagement scoped to reach production. You will be one of the first team members a neocloud or AI Factory engages with at a technical depth, and the playbooks you develop will scale the motion for the next hire and customer. vCluster is gaining rapid traction with GPU AI Clouds and enterprises building AI Factories: organizations that need to offer Kubernetes as a managed service on bare metal GPU infrastructure, and need to do it fast. This role exists to make that happen. As an AI Infrastructure Engineer, your role will include: Lead Technical Deployments: Drive end-to-end technical deployments for GPU neocloud and AI Factory customers, from initial bare metal configuration to a validated vCluster environment. Infrastructure Optimization: Configure and troubleshoot bare metal GPU node infrastructure, including CNI configuration, GPU Operator setup, distributed storage backends, and RDMA/InfiniBand. Validation: Deploy and validate Kubernetes and vCluster to provide GPU-powered managed K8s. Knowledge Transfer: Work alongside customer teams to build self-sufficiency, [... source excerpt omitted ...] ng they can operate and grow the platform independently. Scaling through Documentation: Document reusable playbooks and deployment architectures so your learnings become the next customer's head start. Feedback Loop: Collaborate with Engineering and Product to surface recurring infrastructure challenges, acting as a direct feedback loop from the field into the roadmap. Strategic Partnering: Join Sales in the pre-sales process where deep infrastructure work is required to achieve a meaningful proof of value. This role could be a fit for you if you bring: Production K8s Mastery: 5+ years of experience deploying and operating Kubernetes in production, ideally on bare metal or i [... source excerpt omitted ...] d we have a remote-first work culture. We are the leading platform for operating GPU infrastructure, enabling AI Cloud providers to deliver a hyperscaler-like experience to their customers and AI factories that need to build that same experience for their internal teams. Our platform delivers the full operational stack operators need to run their GPU data centers — managed Kubernetes, fast isolated tenant provisioning, and automated node provisioning and lifecycle management — enabling them to accelerate time to value, reduce operational burden, and maximize the ROI of every GPU. We're the company behind vCluster, an open-source technology for virtualizing Kubernetes (10k+ GitH
Employer postings · Data from · Sources