Agentic AI/ML Engineer, Multimodal
AI summary of the role
This role is for an AI/ML engineer on the Field-insight Foundation Model (FiFM) team, focusing on transforming multimodal data from autonomous robots into actionable insights.
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
What you’ll do
- Train and fine-tune million- to billion-parameter multimodal models for computer vision, video understanding, and vision-language integration.
- Build and optimize Multi-VectorRAG pipelines with vector DBs and knowledge graphs.
- Curate datasets and develop tools for model interpretability and scalable evaluation pipelines.
- Fine-tune and optimize open-source VLMs and multimodal embedding models for efficiency and robustness.
What you’ll bring
- Master’s/Ph.D. in Computer Science, AI/ML, Robotics, or equivalent industry experience.
- 2+ years of industry experience or relevant publications in CV/ML/AI.
- Strong expertise in computer vision, video understanding, temporal modeling, and VLMs.
- Proficiency in Python and PyTorch with production-level coding skills.
Technologies
PyTorch · HuggingFace · DeepSpeed · vLLM · FSDP · LoRA/QLoRA · LangChain · LangGraph · LlamaIndex · OpenSearch · FAISS · Pinecone
About Field AI
Field Foundation Models — embodiment-agnostic robot brains for quadrupeds, humanoids, wheeled and tracked robots in construction, energy, mining, and logistics.
Series D
Source and classification
Internal deployment & tooling · Evidence for this classification:
unlocking the full potential of embodied intelligence. We go beyond typical data-driven approaches or pure transformer-based architectures, and are charting a new course, with already-globally-deployed solutions delivering real-world results and rapidly improving models through real-field applications. Learn more at https://fieldai.com. About the Job Our Field Foundation Model (FFM) powers a global fleet of autonomous robots that capture massive streams of multimodal data across diverse, dynamic environments every day. As part of the Insight Team our mission is to transform this raw, multimodal data into actionable insights that empower our customers and engineers to deliver value. Field-insight Foundation Model (FiFM) is at the core of how we transform multimodal data from autonomous robots into actionable insights. As an AI/ML Engineer on the FiFM team, you will drive research and
More from the job description
FieldAI’s Irvine team is where embodied AI meets real robots, real sensors, and real field deployments. Based in the heart of Southern California’s robotics ecosystem, we build risk-aware, reliable, field-ready AI systems that solve the hardest problems in robotics and unlock the full potential of embodied intelligence. If you want your work to ship, get tested on hardware, and improve through real deployments, Irvine is the place. We go beyond typical data-driven approaches or pure transformer-only architectures, combining rigorous engineering with learning systems proven in globally deployed solutions that deliver results today and get better every time our robots run in the field. Who are We? Field AI is transforming how robots interact with the real world. We are building risk-aware, reliable, and field-ready AI systems that address the most complex challenges in robotics, unlocking the full potential of embodied intelligence. We go beyond typical data-driven approaches or pure transformer-based architectures, and are charting a new course, with already-globally-deployed solutions delivering real-world results and rapidly improving models through real-field applications. Learn more at https://fieldai.com. About the Job Our Field Foundation Model (FFM) powers a global fleet of autonomous robots that capture massive streams of multimodal data across diverse, dynamic envir [... source excerpt omitted ...] e. Field-insight Foundation Model (FiFM) is at the core of how we transform multimodal data from autonomous robots into actionable insights. As an AI/ML Engineer on the FiFM team, you will drive research and model development for one of Field AI’s most ambitious initiatives. Your work will span computer vision, vision-language models (VLMs), multimodal scene understanding, and long-memory video analysis and search, with a strong emphasis on agentic AI (tool use, memory, multimodal retrieval-augmented generation).This is a full-cycle ML role: you’ll curate datasets, fine-tune and evaluate models, optimize inference, and deploy them into production. It’s a blend of applied research [... source excerpt omitted ...] ry experience or relevant publications in CV/ML/AI. Strong expertise in computer vision, video understanding, temporal modeling, and VLMs. Proficiency in Python and PyTorch with production-level coding skills. Experience building pipelines for large-scale video/image datasets. Familiarity with AWS or other cloud platforms for ML training and deployment. Understanding of MLOps best practices (CI/CD, experiment tracking). Hands-on experience fine-tuning open-source multimodal models using HuggingFace, DeepSpeed, vLLM, FSDP, LoRA/QLoRA. Knowledge of precision tradeoffs (FP16, bfloat16, quantization) and multi-GPU optimization. Ability to design scalable evaluation pipelines fo
Employer postings · Data from · Sources