Staff Applied AI Engineer, Product & Agent Performance
AI summary of the role
Staff-level individual contributor owning product-layer decisions for agentic AI in healthcare: prompting, retrieval, memory, evaluation, and escalation.
What you’ll do
- Design and iterate on agent behavior across live, long-horizon, multi-turn workflows
- Architect retrieval, context, memory, and state handling to keep agents grounded in real data
- Build and run production-grounded evaluations, rubrics, and regression checks
- Design escalation paths and human-in-the-loop logic for safety under edge cases
What you’ll bring
- 8+ years production software engineering, including 3+ years owning ML/LLM/agentic systems in production
- Experience in healthcare, finance, or another regulated industry
- Hands-on RAG architecture, production evaluation frameworks, and human-in-the-loop logic
- Working familiarity with AWS Bedrock and SageMaker
Technologies
LLM · RAG · AWS Bedrock · SageMaker · agentic systems · prompt engineering · evaluation frameworks · human-in-the-loop · model cards
Source and classification
Internal deployment & tooling · Evidence for this classification:
healthcare deliver better outcomes for every person, every community, and every generation. Why This Role Is Important to Arcadia Arcadia’s data and analytics platform is used by hundreds of health systems, ACOs, payers, and life sciences organizations, touching tens of millions of patient lives. This role owns how our agentic capabilities perform at that same scale: accurate, transparent about their own confidence, and safe for the clinicians, care teams, and patients who depend on them. As a staff-level individual contributor, you will own the product-layer decisions that shape agent behavior, including prompting, retrieval and context, memory and state, evaluation, and escalation, while partnering with Product and Engineering on the systems that support them. Your work will help Arcadia make evidence-based launch decisions and scale responsible AI that is steerable, trustworthy,
More from the job description
Arcadia is the most trusted healthcare platform powering outcomes. We transform complex healthcare data into trusted intelligence, helping providers, payers, and life sciences organizations act with clarity, make confident decisions, and achieve measurable clinical, operational, and financial outcomes. Built on a comprehensive data foundation spanning tens of millions of patient lives, Arcadia combines advanced analytics and responsible AI to surface meaningful insights, coordinate action, and improve performance at scale. Our approach to AI and automation is governed and transparent — designed to strengthen human expertise, not replace judgment or obscure responsibility. Hundreds of organizations rely on Arcadia to improve cost, quality, and outcomes. Backed by Nordic Capital, we continue to invest in our platform, AI capabilities, and people as we pursue our purpose: helping healthcare deliver better outcomes for every person, every community, and every generation. Why This Role Is Important to Arcadia Arcadia’s data and analytics platform is used by hundreds of health systems, ACOs, payers, and life sciences organizations, touching tens of millions of patient lives. This role owns how our agentic capabilities perform at that same scale: accurate, transparent about their own confidence, and safe for the clinicians, care teams, and patients who depend on them. As a staff- [... source excerpt omitted ...] ed launch decisions and scale responsible AI that is steerable, trustworthy, and ready for real healthcare workflows. What Success Looks Like In 3 months You have established a production-grounded baseline for priority agentic workflows, with documented failure modes, severity-weighted evaluation rubrics, and a clear measurement plan You have mapped the current retrieval, context, memory, and escalation patterns and identified the highest-value opportunities to improve reliability, calibration, and cost You have earned trust across Product and Engineering by turning production evidence into clear, actionable recommendations In 6 months Production-representative evaluation su [... source excerpt omitted ...] e-case conditions, with decision criteria and ownership boundaries clearly documented In 12 months Arcadia has a repeatable product-layer AI performance practice that moves from production failure to diagnosis, experiment, evaluation, and release decision High-severity regressions are caught earlier, and agent behavior is more transparent, calibrated, and trustworthy at scale Model cards, intended-use guidance, limitations, and performance documentation are current and useful to product and customer-facing teams What You'll Be Doing Design and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks Design retrieval and context
Employer postings · Data from · Sources