Eventual logo
Eventual

5 open roles

Research Engineer, Multimodal Data

$150k to $300k

San FranciscoFull-timePosted Apr 29, 2026

Against the San Francisco typical range

Job description

ABOUT EVENTUAL Robots and world models learn about the physical world from video: millions of hours of people cooking, building, carrying and fixing things. That footage is piling up faster than anyone can review it, and much of it is mislabeled, out of sync or not useful for training. Most teams still choose what to train on by having people watch a small sample.

When a model learns from bad examples, it learns the wrong behavior. Eventual is building data curation for Physical AI. We process video and sensor data at petabyte scale and extract signals from it: how hands and objects interact, what happens over time (a failed grasp, a retry, a recovery) and whether a clip is reliable enough to learn from.

Leading Physical AI labs and robotics teams use these signals to decide what goes into their next training run. The work spans large-scale data systems, computer vision and ML research, and it ships directly into how frontier models get trained. We’re a small team from AWS, Lyft and Tesla, backed by $30M from Felicis, CRV, Y Combinator and the co-founders of Databricks and Perplexity.

We helped power the last generation of Physical AI in self-driving, and now we’re building what the next generation trains on. Join our small (but powerful!) team, 4 days/week in our SF Mission District office.

YOUR ROLE

As a Research Engineer on the Visual Understanding team, you'll own the layer that makes petabytes of video queryable by content. Physical AI teams have video, lidar, radar, and sim outputs scattered across object stores with no way to find what they need without weeks of human annotation. Eventual runs vision/language models and pipelines over every clip in a corpus along axes the customer cares about (gripper type, failure mode, object class, scene, motion density), so a researcher can ask "left-arm grasp failures on deformable objects" and get a curated dataset in minutes.

You'll define the roadmap for our visual understanding capabilities, train and select the models that make corpus-scale annotation tractable at single-digit cents per hour of video, and build the rich datasets that go on to train customer models. This is a applied research role — meaning you'll read papers and run experiments, but you ship to production and your work has a real impact on our customers’ models and robots.

KEY RESPONSIBILITIES

  • Own the visual understanding roadmap end-to-end: from picking the model family for a customer's taxonomy to landing it in production inference at corpus scale.
  • Train, fine-tune, and evaluate VLMs, VQA models, embedding models, and CNNs against customer datasets and benchmarks.
  • Drive down per-clip annotation cost — model selection, distillation, batching, decode pipelining — so "annotate every clip in a 10K-hour corpus" stays economical.
  • Build the rich, queryable datasets that customers train on: design taxonomies with researchers, instrument quality, version the outputs.
  • Partner with the dataloading and storage teams so visual understanding outputs flow into the index and on to the GPU without re-engineering.
  • Work directly with researchers at our partner labs — your shortest feedback loop is their next training iteration. WHAT WE LOOK FOR
  • Strong familiarity with modern vision and multimodal models — convolution nets, VLMs, VQA, embeddings — and a sense for the SOTA that's actually deployable today vs. on a leaderboard.
  • Experience running these models at scale on real video and sensor data, ideally for perception tasks (detection, tracking, segmentation, retrieval, captioning).
  • Background from a perception team at a self-driving, robotics, or visual-data company — or equivalent depth from a research lab.
  • Comfortable with cloud infrastructure and large-scale data processing — you don't need to be a distributed-systems engineer, but you've shipped jobs that run on thousands of GPU-hours of video.

NICE TO HAVE

  • Experience training and evaluating vision or multimodal models (not just calling APIs).
  • ML/AI research background — papers, citations, or a research org on your resume.
  • Worked on embeddings, retrieval, or content-aware search at scale.
  • Experience designing labeling taxonomies or running annotation programs. PERKS & BENEFITS
  • In-person, tight-knit team — 4 days/week in our SF Mission office.
  • Competitive comp and meaningful startup equity.
  • Catered lunches and dinners for SF employees.
  • Commuter benefit.
  • Team-building events and poker nights.
  • Health, vision, and dental coverage.
  • Flexible PTO.
  • Latest Apple equipment. - 401(k) plan with match. If you're excited about being on the team that turns petabytes of raw video into the training data for the next generation of Physical AI, we'd love to talk.

Description copied from Eventual's careers page. Read the full posting before you apply.

More jobs at Eventual

See all openings at Eventual

More jobs in San Francisco

See all jobs in San Francisco