Skip to content
NEXTMOVEFDE careers · United States

Research Engineer, Applied AI

AI summary of the role

Research Engineer on a new Applied AI team at HeyMilo, designing and building RL environments with verifiable rewards for real-world hiring use cases.

No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.

What you’ll do

  • Design and build RL environments for real-world use cases: simulators, tool interfaces, episodic task generation with isolation and reproducibility
  • Design verifiable reward functions that score intermediate actions and end states, robust to reward hacking
  • Build and maintain evaluation harnesses for frontier and open-weight models with clean scoring and cost tracking
  • Run post-training experiments (e.g., RLVR-style fine-tuning) to validate learnable signals

What you’ll bring

  • Master's or PhD in AI, ML, CS, or related field
  • Solid grounding in reinforcement learning and LLM post-training (reward design, policy optimization, evaluation methodology)
  • Strong software engineering skills in Python
  • Hands-on experience with LLMs: evaluations, agentic loops, tool calling, fine-tuning

Technologies

reinforcement learning · LLM post-training · Python · Docker · Linux · cloud · RLVR · fine-tuning · agentic loops · tool calling

About HeyMilo

Agentic recruiting software that sources, screens, interviews, and scores candidates for staffing firms, BPOs, and enterprise hiring teams inside their ATS workflows.

Seed

Source and classification

Internal deployment & tooling · Evidence for this classification:

HeyMilo is building AI interviewers that automate and improve hiring through conversational AI. We work closely with companies to bring AI into real hiring workflows. We also run an applied AI team that studies where today's models succeed and fail across industries. The Role We're hiring a Research Engineer to join our Applied AI team in San Francisco. You'll design and build reinforcement learning environments with verifiable rewards for specific real-world use cases: the simulators, reward functions, and evaluation harnesses that let us measure and improve how models perform on real work. The work is equal parts ML research and systems engineering. You'll work directly with the founders and our research advisors, taking a use case from problem definition to a reproducible environment that models can be evaluated and trained against. What you'll do Design and build RL
More from the job description

HeyMilo is building AI interviewers that automate and improve hiring through conversational AI. We work closely with companies to bring AI into real hiring workflows. We also run an applied AI team that studies where today's models succeed and fail across industries. The Role We're hiring a Research Engineer to join our Applied AI team in San Francisco. You'll design and build reinforcement learning environments with verifiable rewards for specific real-world use cases: the simulators, reward functions, and evaluation harnesses that let us measure and improve how models perform on real work. The work is equal parts ML research and systems engineering. You'll work directly with the founders and our research advisors, taking a use case from problem definition to a reproducible environment that models can be evaluated and trained against. What you'll do Design and build RL environments for specific real-world use cases: realistic simulators, tool interfaces, and episodic task generation with proper isolation and reproducibility Design verifiable reward functions that score correct intermediate actions as well as end states, and hold up against reward hacking Build and maintain evaluation harnesses that run task suites across frontier and open-weight models, with clean scoring and cost tracking Run post-training experiments (e.g. RLVR-style fine-tuning of open models) to validate that your environments produce a learnable signal Package environments and results for reproducibility, and contribute to research write-ups and published evaluations Help define which use cases we pursue next, informed by where models are weakest What we're looking for Master's or PhD in AI, Machine Learning, Computer Science, or a closely related field Solid grounding in reinforcement learning and LLM post-training (reward design, policy optimization, evaluation methodology) Strong software engineering skills in Python, with code that others can run and build on Hands-on experience with LLMs: running evaluations, building agentic loops, tool calling, fine-tuning Comfortable with containers and infrastructure (Docker, Linux, cloud environments) for reproducible experiment setups Ability to operate in ambiguity and move quickly; comfortable owning a problem end to end Based in (or willing to relocate to) the San Francisco Bay Area Bonus Published research or open-source contributions in ML, RL, or evaluation Experience with RL/eval frameworks and simulated or sandboxed environments Experience training or fine-tuning open-weight models at any scale Why join Ground-floor role on a new applied AI team with real influence over how we build Work directly with the founders and experienced research advisors, with your name on published work High visibility, fast-paced, execution-driven environment Competitive pay, equity, and benefits

How jobs are selected

Employer postings · Data from · Sources