Skip to content
NEXTMOVEFDE careers · United States

Senior Data Engineer

AI summary of the role

TRI's Automated Driving Advanced Development division is building a fully end-to-end learned driving stack in collaboration with Woven by Toyota.

What you’ll do

  • Design and implement scalable, production-grade pipelines for data ingestion, transformation, storage, and retrieval from vehicle fleets and simulation environments.
  • Build internal tools and services for data labeling, curation, indexing, and cataloging across large and diverse datasets.
  • Collaborate with ML researchers, autonomy engineers, and data scientists to design schemas and APIs for model training, evaluation, and debugging.
  • Develop and maintain feature stores, metadata systems, and versioning infrastructure for structured and unstructured data.

What you’ll bring

  • 8+ years of experience building data-intensive software systems, ideally in robotics, autonomous driving, or large-scale ML environments.
  • Proficient in Python, SQL, and familiar with C++.
  • Experience designing ETL pipelines using modern frameworks (e.g., Apache Spark, Flyte, Union).
  • Strong knowledge of cloud-native architectures, including AWS services (e.g., S3, or equivalents (Google Cloud platform).

Technologies

Python · SQL · C++ · Apache Spark · Flyte · Union · AWS · S3 · protobuf · ROS2bag · MCAP · Kubeflow

About Toyota Research Institute

Toyota's advanced R&D lab building AI, robotics, driving, and materials breakthroughs that feed directly into Toyota's vehicles, factories, and mobility platforms.

Established

Source and classification

Internal deployment & tooling · Evidence for this classification:

Behavior Models. We are looking for a Senior Data Engineer to design and build the foundational data infrastructure and tools that power our autonomy research and development workflows. This includes large-scale ingestion pipelines, structured feature stores, labeling infrastructure, scene search and data discovery tools, and performance diagnostics for machine learning and simulation workflows. Responsibilities Design and implement scalable, production-grade pipelines for data ingestion, transformation, storage, and retrieval from vehicle fleets and simulation environments. Build internal tools and services for data labeling, curation, indexing, and cataloging across large and diverse datasets. Collaborate with ML researchers, autonomy engineers, and data scientists to design schemas and APIs that power model training, evaluation, and debugging. Develop and maintain feature
More from the job description

At Toyota Research Institute (TRI), we’re on a mission to improve the quality of human life. We’re developing new tools and capabilities to amplify the human experience. To lead this transformative shift in mobility, we’ve built a world-class team advancing the state of the art in AI, robotics, driving, and material sciences. The Automated Driving Advanced Development division at TRI will focus on enabling innovation and transformation at Toyota by building a bridge between TRI research and Toyota products, services, and needs. We achieve this through partnership, collaboration, and shared commitment. This new division is leading a new cross-organizational project between TRI and Woven by Toyota to conduct research and develop a fully end-to-end learned driving stack. This cross-org collaborative project is harmonious with TRI’s robotics divisions' efforts in Diffusion Policy and Large Behavior Models. We are looking for a Senior Data Engineer to design and build the foundational data infrastructure and tools that power our autonomy research and development workflows. This includes large-scale ingestion pipelines, structured feature stores, labeling infrastructure, scene search and data discovery tools, and performance diagnostics for machine learning and simulation workflows. Responsibilities Design and implement scalable, production-grade pipelines for data ingestion, tra [... source excerpt omitted ...] and consistency across environments. Partner with simulation and cloud platform teams to automate workflows for closed-loop testing, scenario mining, and performance analytics. Qualifications Bachelor’s or Master’s degree in Computer Science, Data Engineering, or a related field. 8+ years of experience building data-intensive software systems, ideally in robotics, autonomous driving, or large-scale ML environments. Proficient in Python, SQL, and familiar with C++. Experience designing ETL pipelines using modern frameworks (e.g., Apache Spark, Flyte, Union). Strong knowledge of cloud-native architectures, including AWS services (e.g., S3, or equivalents (Google Cloud platform) [... source excerpt omitted ...] a quality, observability, and lineage in high-volume systems. Track record of building reliable and performant infrastructure that supports both ad-hoc exploration and repeatable production workflows. Bonus Qualifications Experience in AD/ADAS, robotics, or autonomous systems — especially handling perception or planning datasets. Familiarity with ML pipeline orchestration frameworks (e.g. Kubeflow, SageMaker, etc). Experience working with temporal or spatial data, including geospatial indexing and time-series alignment. Exposure to synthetic data generation, simulation logging, or scenario replay pipelines. Strong software engineering fundamentals, CI/CD, testing, code revie

How jobs are selected

Employer postings · Data from · Sources