Skip to content
NEXTMOVEFDE careers · United States

Software Engineer – ML Platform

Technologies

Kubernetes · Argo Workflows · Ray · Python · Go · MLflow · GPU scheduling · distributed training

About Avride

Builds robotaxis and sidewalk delivery robots on a shared autonomous driving stack, commercialized through partners such as Uber, Grubhub, and Rakuten.

Series C

Job description

The full responsibilities and requirements are on the employer’s site.

Read the job description
Source and classification

Internal deployment & tooling · Evidence for this classification:

About the team The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use — scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle. About the role As an ML Platform Engineer at Avride, you'll own critical pieces of the ML stack: workflow orchestration, distributed execution, resource governance, performance.You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience.
More from the job description

About the team The ML Platform team at Avride builds the infrastructure that powers large-scale ML training and data processing for autonomous driving. We sit between Cloud Platform and ML engineers, turning low-level compute, storage, and networking primitives into an ML platform that teams actually use — scalable orchestration, distributed compute, and production-grade tooling for the full model lifecycle. About the role As an ML Platform Engineer at Avride, you'll own critical pieces of the ML stack: workflow orchestration, distributed execution, resource governance, performance.You will shape how ML teams across the company run experiments and train models at scale. You will build the abstractions and services that make training workloads reliable, cost-efficient, and fast, helping ML teams run at scale on Kubernetes with strong reliability and excellent developer experience. What you will do Build and scale our ML compute platform on Kubernetes, using Argo Workflows for training, evaluation, and data processing orchestration Design and implement core platform capabilities, including a Ray-based internal SDK for distributed execution, and multi-tenant resource governance — scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and IO Improve end-to-end training throughput and platform efficiency by optimizing data access patterns, caching, an [... source excerpt omitted ...] recurring problems into platform-level fixes Evaluate, integrate and extend open-source tooling (Argo Workflows, Ray, Kubernetes ecosystem) to meet evolving platform needs What you will need Strong proficiency in Python or Go; C++ is a plus Track record of designing and building scalable, maintainable systems and services Experience operating production services end-to-end: APIs, reliability practices, observability Deep knowledge of Kubernetes: how scheduling, resource management, controllers, and pod lifecycle actually behave under pressure Solid Linux and systems debugging skills: performance investigation, networking, storage/IO Ability to troubleshoot complex produc

How jobs are selected

Employer postings · Data from · Sources