Senior Solutions Engineer
AI summary of the role
Senior Solutions Engineer acting as the elite escalation point between the Global Operations Center and Core Engineering for an AMD-exclusive AI cloud provider.
What you’ll do
- Resolve complex escalations as final authority on issues exceeding GOC scope using code-level debugging and architectural investigation.
- Partner with customer technical leads to diagnose production issues and ensure rapid resolution.
- Develop diagnostic scripts and workarounds to maintain operations while long-term patches are developed.
- Own end-to-end P1 resolution and deliver clear post-incident analysis with TAMs.
What you’ll bring
- 5–9 years in Infrastructure Engineering, Platform Engineering, or SRE with focus on HPC or large-scale AI stacks.
- Deep Kubernetes expertise in cluster administration and scheduler internals.
- Proficient in orchestrating GPU workloads and diagnosing training job failures using ROCm or CUDA.
- Skilled in RDMA/RoCEv2, SRIOV, and BGP with ability to interpret switch telemetry.
Technologies
Kubernetes · ROCm · CUDA · RDMA · RoCEv2 · SRIOV · BGP · Python · Ansible · GPU
About TensorWave
AMD-exclusive AI cloud provider offering bare-metal and reserved inference infrastructure for training and serving large models without NVIDIA/CUDA lock-in.
Series A
Employer postings · Data from · Sources