Skip to content
NEXTMOVEFDE careers · United States

Senior Machine Learning Engineer - AI Foundation

AI summary of the role

This role focuses on building state-of-the-art ML infrastructure for training very large foundation models and accelerating model training/inference for autonomous driving.

What you’ll do

  • Optimize transformer-based LLMs for low-latency and high-throughput inference.
  • Optimize kernels and model graphs using CUDA, Triton, and custom fused operators.
  • Implement and benchmark quantization, knowledge distillation, pruning, and KV-cache optimization.
  • Deploy optimized models across GPUs, CPUs, and edge accelerators.

What you’ll bring

  • Master in CS/CE/EE or equivalent with 3+ years industry experience.
  • Good knowledge of PyTorch.
  • Knowledge of transformer architecture and ways to accelerate training/inference.

Technologies

PyTorch · CUDA · Triton · TensorRT · Torchscript · transformer · LLM · GPU · NPU · DSP

About XPENG Motors

Chinese smart-EV maker building in-house autonomous driving, robotaxi, eVTOL flying cars, and the IRON humanoid robot on a unified physical-AI foundation.

Public · 5000+ people

Source and classification

Internal deployment & tooling · Evidence for this classification:

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity. We are looking for a full-time Machine Learning Engineer - AI Foundation, with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training very large foundation model and accelerating model training/inference. Our mission is to solve the autonomous driving problem. You will work with a team of talented software engineers, machine learning engineers and research scientists to push the boundary of state-of-art
More from the job description

XPENG is a leading smart technology company at the forefront of innovation, integrating advanced AI and autonomous driving technologies into its vehicles, including electric vehicles (EVs), electric vertical take-off and landing (eVTOL) aircraft, and robotics. With a strong focus on intelligent mobility, XPENG is dedicated to reshaping the future of transportation through cutting-edge R&D in AI, machine learning, and smart connectivity. We are looking for a full-time Machine Learning Engineer - AI Foundation, with deep knowledge and strong enthusiasm towards establishing a state-of-art ML infrastructure for training very large foundation model and accelerating model training/inference. Our mission is to solve the autonomous driving problem. You will work with a team of talented software engineers, machine learning engineers and research scientists to push the boundary of state-of-art machine learning models which will enable the next-generation E2E solution of autonomous driving. Job Responsibilities: Optimize transformer-based LLMs for low-latency and high-throughput inference. Optimize kernels and model graphs using tools like CUDA, Triton, and custom fused operators. Implement and benchmark (Quantization, Knowledge distillation, structured and unstructured pruning, KV-cache optimization, etc.). Deploy optimized models across GPUs, CPUs, and edge acceleators. Contribu [... source excerpt omitted ...] f industry experience. Good knowledge of PyTorch. Knowledge of transformer architecture and ways to accelerate the training and inference of transformer models. Preferred Skill Requirements: Previous experience in the autonomous driving industry. Knowledge of Torchscript and Nvidia TensorRT. Strong programming skills in Python and C++ Familiarity with GPU CPU, NPU, DSP architecture. Deep understanding of memory bandwidth, compute bottlenecks, and hardware-aware model optimization Being efficiently in solving complex problems collaboratively on larger teams What do we provide: A fun, supportive and engaging environment. Infrastructures and computational resources to suppor

How jobs are selected

Employer postings · Data from · Sources