Research Engineer, Multimodal Data
Technologies
VLM · VQA · embedding models · convolutional perception models · GPU · Spark · Ray · Daft · object stores · distributed data processing
About Eventual Computing
Builds Daft and Daft Cloud so AI teams can process multimodal data with Python-native, fault-tolerant pipelines instead of stitching together Spark-era infrastructure.
Series A
Job description
The full responsibilities and requirements are on the employer’s site.
Open application page ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
About Eventual From humanoid robots to autonomous vehicles, every Physical AI model is trained on petabytes of video, lidar, radar, and sensor data. Today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not video corpora. And understanding that video still means paying a person to watch it, ten dollars an hour of footage at the low end. So teams check a sample and hope it represents the rest. The footage grows every year; the budget to look at it doesn't. Eventual was founded in 2022 to close that gap. Our open-source engine, Daft, is purpose-built for multimodal AI: 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at companies like Mobileye, TogetherAI. On top of it we're building the infrastructure that finds any situation you can describe across a fleet's entire video history, and turns it into a training set or an alert
More from the job description
About Eventual From humanoid robots to autonomous vehicles, every Physical AI model is trained on petabytes of video, lidar, radar, and sensor data. Today's data platforms (Databricks, Snowflake) were built for spreadsheet-like analytics, not video corpora. And understanding that video still means paying a person to watch it, ten dollars an hour of footage at the low end. So teams check a sample and hope it represents the rest. The footage grows every year; the budget to look at it doesn't. Eventual was founded in 2022 to close that gap. Our open-source engine, Daft, is purpose-built for multimodal AI: 2 PB/day at Amazon, 60-100 PB at another FAANG company, and in production at companies like Mobileye, TogetherAI. On top of it we're building the infrastructure that finds any situation you can describe across a fleet's entire video history, and turns it into a training set or an alert someone can still act on. We fine-tune and run the vision models ourselves, which makes indexing every hour cheaper than annotating a sample. We're building this with the top Physical AI labs and GPU cloud providers. We've raised $30M from investors like Felicis, CRV, Y Combinator, and angels from the co-founders of Databricks and Perplexity. Our team comes from AWS, Lyft, and Tesla. We powered the last generation of Physical AI in self-driving; now we're doing it for the next. Join our small ( [... source excerpt omitted ...] s object stores with no way to find what they need without weeks of human annotation. Eventual runs vision/language models and pipelines over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask "left-arm grasp failures on deformable objects" and get a curated dataset in minutes. You'll define the roadmap for our visual understanding capabilities, train and select the models that make corpus-scale annotation tractable at single-digit cents per hour of video, and build the rich datasets that go on to train customer models. This is a applied research role — meaning you'll read papers a [... source excerpt omitted ...] on and your work has a real impact on our customers’ models and robots. Key Responsibilities Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale. Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and CNNs against customer datasets and benchmarks. Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so "annotate every clip in a 10K-hour corpus" stays economical. Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version the outputs. Partner wit
Employer postings · Data from · Sources