Forward Deployed Infrastructure Engineer - Eastern US
Technologies
GPU · NCCL · container · AWS · Lambda · CoreWeave · Runpod · ML model benchmarks · inference · training
About Hyperbolic Labs
Marketplace-style AI cloud for startups and researchers to rent GPUs, run serverless inference, and reserve dedicated capacity without hyperscaler friction.
Series A · 10–50 people
Job description
The full responsibilities and requirements are on the employer’s site.
Open application page ↗Source and classification
Implementation & delivery · Evidence for this classification:
misbehave the problem is rarely simple. You are the engineer embedded with those customers: you stand their cluster up, you hand it over, and you stay with it. You do not own tickets. You own environments. Technical Support Engineers own the ticket lifecycle and pull you in when an issue needs real depth: multinode collective performance, hardware faults, fabric problems, or a provider who needs to be told what is wrong with their hardware. One thing we will be straight about, because it shapes the job. We aggregate capacity from suppliers rather than owning most of the hardware ourselves. That means a real part of this role is technical liaison work: proving where a fault actually lives, taking it to the provider with evidence, and coordinating the fix on the customer's behalf. The engineers who enjoy this role are the ones who find that interesting rather than frustrating. Who You
More from the job description
Who We Are Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality. NOTE: This role is focused on EASTERN TIME ZONE About the Role Our reserved customers run large multinode GPU clusters, and when those clusters misbehave the problem is rarely simple. You are the engineer embedded with those customers: you stand their cluster up, you hand it over, and you stay with it. You do not own tickets. You own environments. Technical Support Engineers own the ticket lifecycle and pull you in when an issue needs real depth: multinode collective performance, hardware faults, fabric problems, or a provider who needs to be told what is wrong with their hardware. One thing we will be straight about, because it shapes t [... source excerpt omitted ...] lves. That means a real part of this role is technical liaison work: proving where a fault actually lives, taking it to the provider with evidence, and coordinating the fix on the customer's behalf. The engineers who enjoy this role are the ones who find that interesting rather than frustrating. Who You Are Cluster stand-up and handoff. Build, validate, and benchmark new customer clusters, then hand them over with documentation the customer's own engineers can work from. Deep escalations. Multinode and NCCL performance debugging, GPU and hardware faults (XID and ECC errors, lspci, dmesg), driver and fabric issues, container and scheduler problems. Provider escalation and coor [... source excerpt omitted ...] ribution. You bring the evidence that makes the provider act, and you keep the customer informed while it happens. Embedded ownership of named accounts. You are the engineer your customers know by name. You learn their workload, not just their infrastructure, and you tell them what to change before they hit the wall. Proactive monitoring. Own the monitoring and alerting we put in front of customer clusters (Grafana, Prometheus) so we find faults before the customer reports them. Tooling and pushing work down. Automate the repeat work and turn your own escalations into runbooks the L1 tier can run. Anything you fix three times should stop reaching you. Deep Linux experience an
Employer postings · Data from · Sources