Lead Machine Learning Engineer
AI summary of the role
Lead the design and scaling of distributed training systems for petabyte-scale multimodal robotics data (video, point clouds) at a public autonomous sidewalk delivery company.
What you’ll do
- Design and maintain training systems for petabyte-scale multimodal datasets (video, point cloud) across large GPU clusters.
- Identify and resolve bottlenecks in data loading, preprocessing, model computation, and inter-node communication to maximize GPU utilization.
- Develop and refine neural network architectures for autonomy tasks handling high-dimensional sequential sensor data.
- Configure, monitor, and maintain large-scale distributed training jobs across multiple machines and GPUs ensuring stability and fault tolerance.
What you’ll bring
- Master’s or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or closely related field.
- Minimum 5 years professional experience developing, training, and deploying ML models in production.
- Hands-on experience training ML models across multiple GPUs or compute nodes with distributed training frameworks.
- Strong programming skills in Python for ML models, data pipelines, and training workflows.
Technologies
distributed training · GPU clusters · Python · multimodal data · video · point cloud · LiDAR · neural network architectures · loss functions · data pipelines
About Serve Robotics
Builds and operates AI-powered sidewalk delivery robots and supporting autonomy software for restaurant, retail, and now hospital logistics workflows.
Public
Source and classification
Internal deployment & tooling · Evidence for this classification:
diverse, and driven. We believe that the best way to solve complicated dynamic problems is collaboratively and respectfully. This role develops and scales large-scale machine learning training systems for multimodal robotics data, enabling the creation of high-performance autonomy models. By optimizing distributed training pipelines, neural network architectures, and data processing workflows, the position improves training efficiency, accelerates model iteration, and maximizes GPU utilization. The role collaborates closely with ML researchers and infrastructure teams, influencing the design, deployment, and performance of end-to-end autonomy models and the large-scale data pipelines that support them. Responsibilities Design and maintain training systems that can process and learn from petabyte-scale multimodal datasets (e.g., video and point cloud data). This includes ensuring data
More from the job description
At Serve Robotics, we’re reimagining how things move in cities. Our personable sidewalk robot is our vision for the future. It’s designed to take deliveries away from congested streets, make deliveries available to more people, and benefit local businesses. The Serve fleet has been delighting merchants, customers, and pedestrians along the way in Los Angeles, Miami, Dallas, Atlanta and Chicago while doing commercial deliveries. We’re looking for talented individuals who will grow robotic deliveries from surprising novelty to efficient ubiquity. Who We Are We are tech industry veterans in software, hardware, and design who are pooling our skills to build the future we want to live in. We are solving real-world problems leveraging robotics, machine learning and computer vision, among other disciplines, with a mindful eye towards the end-to-end user experience. Our team is agile, diverse, and driven. We believe that the best way to solve complicated dynamic problems is collaboratively and respectfully. This role develops and scales large-scale machine learning training systems for multimodal robotics data, enabling the creation of high-performance autonomy models. By optimizing distributed training pipelines, neural network architectures, and data processing workflows, the position improves training efficiency, accelerates model iteration, and maximizes GPU utilization. The ro [... source excerpt omitted ...] h ML researchers and infrastructure teams, influencing the design, deployment, and performance of end-to-end autonomy models and the large-scale data pipelines that support them. Responsibilities Design and maintain training systems that can process and learn from petabyte-scale multimodal datasets (e.g., video and point cloud data). This includes ensuring data is efficiently loaded, distributed, and processed across large GPU clusters. Identify and resolve bottlenecks in the training pipeline, including data loading, preprocessing, model computation, and inter-node communication, to maximize GPU utilization and reduce training time. Work with the ML team to develop and refine neu [... source excerpt omitted ...] datasets so that they are suitable for model training. Work closely with ML scientists and other engineers to integrate new models, experiments, and training approaches into the production training pipeline. Analyze training metrics, model outputs, and experiment logs to assess model performance and guide improvements in architecture, data usage, or training strategies. Develop tools and workflows that allow teams to run experiments, track results, and iterate quickly on new model ideas or training approaches. Qualifications Master’s or PhD in Computer Science, Robotics, Electrical Engineering, Machine Learning, or a closely related technical discipline. Minimum of 5 years o
Employer postings · Data from · Sources