Senior AI Infrastructure Engineer
Technologies
PyTorch Distributed · Ray Train · NCCL · InfiniBand · Kubernetes · Terraform · Helm · TensorRT · Triton Inference Server · MLFlow · Argo Workflows · LangGraph
About Gatik AI
Builds and operates driverless middle-mile freight networks for retailers and CPG shippers using fixed-route Level 4 medium-duty trucks.
Series C
Job description
The full responsibilities and requirements are on the employer’s site.
Read the job description ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
Carrier™ serves as an all-encompassing solution that integrates advanced software and hardware powering the fleet, facilitating effortless integration into customers' logistics operations. About the role We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that enables distributed training, experiment tracking, and seamless model deployment. You will bridge the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient This role is onsite 5 days a week at our Santa Clara, CA office! What you'll do Distributed Training & ML Systems Support Scale Research Workloads: Enable researchers to scale complex models
More from the job description
Who we are Gatik, the leader in autonomous middle-mile logistics, is revolutionizing the B2B supply chain with its autonomous transportation-as-a-service (ATaaS) solution and prioritizing safe, consistent deliveries while streamlining freight movement by reducing congestion. The company focuses on short-haul, B2B logistics for Fortune 500 retailers and in 2021 launched the world’s first fully driverless commercial transportation service with Walmart. Gatik's Class 3-7 autonomous trucks are commercially deployed across major markets, including Texas, Arkansas, and Ontario, Canada, driving innovation in freight transportation. The company's proprietary Level 4 autonomous technology, Gatik Carrier™, is custom-built to transport freight safely and efficiently between pick-up and drop-off locations on the middle mile. With robust capabilities in both highway and urban environments, Gatik Carrier™ serves as an all-encompassing solution that integrates advanced software and hardware powering the fleet, facilitating effortless integration into customers' logistics operations. About the role We are seeking a Senior AI Infrastructure Engineer to design, build, and scale the high-performance AI platform powering our autonomous driving models. While researchers focus on developing perception, planning, and world models, you will be responsible for the underlying infrastructure that ena [... source excerpt omitted ...] the gap between research and production, ensuring our AI stack is scalable, resilient, and highly efficient This role is onsite 5 days a week at our Santa Clara, CA office! What you'll do Distributed Training & ML Systems Support Scale Research Workloads: Enable researchers to scale complex models (VLA, World Models) across multi-node setups using PyTorch Distributed, and Ray Train. Performance Optimization: Architect and optimize multi-GPU setups, ensuring efficient model parallelism and data parallelism techniques across H100/A100 clusters. Networking & Hardware Tuning: Optimize low-level communication (e.g., NCCL tuning, InfiniBand, or RoCE v2) to minimize latency for [... source excerpt omitted ...] Helm. Data & CI/CD: Proficiency in Apache Airflow, Kafka, Spark, and GitOps automation. Core Skills: Proficiency in Python and Bash; experience with Go or Rust is a plus Bonus Qualifications Advanced AI Protocols: Familiarity with the Model Context Protocol (MCP) to standardize how AI agents interact with internal databases and orchestration APIs. Hybrid & Physical AI: Experience in hybrid cloud and on-prem GPU cluster management for Physical AI workloads (e.g., 3DGS, World Models). Agentic Observability: Experience utilizing LLMs for semantic monitoring and log analysis to detect complex distributed system failures that traditional threshold-based alerts miss. Salary Ranges -
Employer postings · Data from · Sources