Skip to content
NEXTMOVEFDE careers · United States

Customer Success Engineer (CSE), GPU Cluster

AI summary of the role

This role is the named technical owner for a strategic customer of Together AI's GPU cloud, responsible for end-to-end technical relationship and operational health of large-scale GPU deployments across compute, networking, storage, and facilities.

No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.

What you’ll do

  • Serve as named technical point of contact for a dedicated strategic customer, owning end-to-end technical relationship across compute, networking, storage, and facilities
  • Drive structured engagement through status reporting, technical steering meetings, QBRs, and EBRs
  • Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains
  • Own RMA coordination and hardware lifecycle management for large-scale GPU deployments

What you’ll bring

  • 5+ years in customer-facing technical role, with 2+ years in dedicated TAM or solutions architecture for large-scale AI or HPC infrastructure
  • Deep expertise in GPU infrastructure: health diagnostics, RMA workflows, hardware acceptance testing
  • Hands-on experience with large-scale Ethernet and InfiniBand fabric architecture
  • Working knowledge of enterprise storage systems (high-density NVMe, parallel file systems, metadata infrastructure)

Technologies

GPU · InfiniBand · Ethernet · NVMe · parallel file systems · Prometheus · Grafana · Python · Bash · RMA

About Together AI

GPU cloud and inference platform optimized for open-source LLMs, serving 450K+ developers and enterprises with serverless APIs, dedicated clusters, and fine-tuning.

Series B

Source and classification

Customer adoption & accounts · Evidence for this classification:

About the role As a Customer Success Engineer at Together AI, you will serve as the named technical owner for one of our most strategic customer relationships. You will be the primary technical point of contact across all infrastructure domains — compute, networking, storage, and facilities — ensuring flawless delivery and operational health of large-scale GPU deployments. This role sits at the intersection of deep infrastructure expertise and high-stakes customer partnership, making you a critical driver of both customer success and company growth. Responsibilities Serve as the named technical point of contact for a dedicated strategic customer, owning the end-to-end technical relationship across compute, networking, storage, and facilities Drive structured engagement through regular cadences — status reporting, technical steering meetings, quarterly business reviews (QBRs), and
More from the job description

About the role As a Customer Success Engineer at Together AI, you will serve as the named technical owner for one of our most strategic customer relationships. You will be the primary technical point of contact across all infrastructure domains — compute, networking, storage, and facilities — ensuring flawless delivery and operational health of large-scale GPU deployments. This role sits at the intersection of deep infrastructure expertise and high-stakes customer partnership, making you a critical driver of both customer success and company growth. Responsibilities Serve as the named technical point of contact for a dedicated strategic customer, owning the end-to-end technical relationship across compute, networking, storage, and facilities Drive structured engagement through regular cadences — status reporting, technical steering meetings, quarterly business reviews (QBRs), and executive business reviews (EBRs) — spanning both operational and strategic levels Translate customer operational feedback into actionable input for Engineering, Product, and Infrastructure roadmaps Lead issue lifecycle management, escalation, and RCA authorship across all infrastructure domains in partnership with Support, SRE, DC Ops, and Engineering teams Own end-to-end RMA coordination and hardware lifecycle management, including acceptance testing, spare inventory management, and hardware h [... source excerpt omitted ...] ompute, high-speed fabric, and large-scale storage systems — advising on configuration, operational best practices, and incident resolution Own the observability strategy for the customer estate, including alert policy definition, dashboard development, and proactive health management across all infrastructure layers Coordinate DC operations and facilities events in partnership with internal teams and hosting providers, ensuring SLA compliance and cluster availability Act as project manager for all capacity expansions, owning the full node deployment lifecycle from freight receipt through production acceptance Qualifications 5+ years in a customer-facing technical role, with [... source excerpt omitted ...] Experience with DC operations, facilities coordination, and hosting provider SLA management Strong ownership mindset for incident management, RCA authorship, and executive-level customer communication Proficiency in infrastructure monitoring and observability tooling (Prometheus, Grafana, or equivalent) Proven ability to manage multiple concurrent workstreams with hyperscaler-level rigor and communication standards Proficiency in Python, Bash, or infrastructure automation tools preferred About Together AI Together AI, the AI Native Cloud, is purpose-built for AI engineers. AI application developers get high-performance inference that scales reliably, fine-tuning and reinfo

How jobs are selected

Employer postings · Data from · Sources