Staff Software Engineer, Data Warehouse
AI summary of the role
Staff-level IC role owning Commure's data warehouse platform end-to-end, from CDC pipelines and data lake to query and transformation layers.
What you’ll do
- Own the data warehouse platform end-to-end: CDC pipelines, data lake, query layer, transformation layer, and analytics tooling.
- Design and operate CDC pipelines with Debezium and streaming backbones like Kafka or Redpanda.
- Architect the data lake on object storage using open table formats (Iceberg, Delta Lake, Hudi) with Parquet.
- Run and scale StarRocks as the query and serving layer, including schema design and performance tuning.
What you’ll bring
- 6+ years of software engineering experience, with significant time building or operating data platforms at scale.
- Experience with CDC (Debezium or equivalent), streaming (Kafka, Redpanda), data lake formats (Iceberg, Delta, Hudi), and an MPP/lakehouse engine (StarRocks, ClickHouse, Trino, Snowflake, Databricks).
- Fluent in SQL, schema design, query optimization, and cost/latency trade-offs on large datasets.
- Experience running production data infrastructure including orchestration, observability, on-call, and data quality.
Technologies
Debezium · StarRocks · dbt · Kafka · Redpanda · Iceberg · Delta Lake · Hudi · Parquet · Airflow · Dagster · ClickHouse
About Commure
Builds an AI-native platform for health systems that unifies ambient documentation, agents, patient engagement, and revenue-cycle automation across deeply integrated EHR workflows.
Private Late · 1000–2000 people
Source and classification
Internal deployment & tooling · Evidence for this classification:
Awards for “Overall NLP Company of the Year.” Our team works directly alongside clinicians, not through layers of process, which means the gap between what you build and its impact on patient care is immediate. We move fast, deploy daily, and take full ownership from early thinking to production. If you're energized by hard problems, high stakes, and a team that holds itself to a high bar, you'll find your people here. The future of healthcare is being built right now. Come deliver this transformation. About the Role We're hiring a Staff Software Engineer to own Commure's data warehouse platform end-to-end. You'll design, build, and operate every layer of the stack: CDC pipelines Streaming transport Schema governance and data contracts Query and serving layer Analytics platform The core stack today: Debezium for CDC, StarRocks as our MPP query and serving engine, and dbt for
More from the job description
At Commure, we're building the AI Operating System for healthcare, the foundation that defines how care is delivered, documented, and financed. Our platform spans the full care journey: Ambient AI and Dictation eliminating documentation burden at the point of care, intelligent Agents automating patient and revenue workflows, and autonomous RCM processing billions in claims, all on a single AI-native platform integrated with 60+ EHRs. Healthcare carries a $1 trillion administrative burden and we're at the center of transforming it. Today, 500,000+ clinicians across 500+ healthcare organizations nationwide trust Commure to handle $25B+ in annual claims and support over 200 million patient interactions. Our latest $70M raise at a $7B valuation reflects the confidence the market has placed in this mission. We've also been named to the Fortune Future 50 list and the 2026 AI Breakthrough Awards for “Overall NLP Company of the Year.” Our team works directly alongside clinicians, not through layers of process, which means the gap between what you build and its impact on patient care is immediate. We move fast, deploy daily, and take full ownership from early thinking to production. If you're energized by hard problems, high stakes, and a team that holds itself to a high bar, you'll find your people here. The future of healthcare is being built right now. Come deliver this transforma [... source excerpt omitted ...] tform The core stack today: Debezium for CDC, StarRocks as our MPP query and serving engine, and dbt for transformation and modeling. This is a hands-on IC role with broad scope. You'll make architectural calls, write the code that matters most, and set the patterns other teams build on. What You'll Do Own the data warehouse platform end-to-end: CDC pipelines, data lake, query layer, transformation layer, and the analytics-facing tooling that sits on top. Design and operate CDC pipelines with Debezium (and Kafka, Redpanda, or an equivalent streaming backbone) that move data from operational databases into the warehouse with low latency and high fidelity. Architect the data [... source excerpt omitted ...] , Trino, Snowflake, Databricks), and dbt. Fluent in SQL, schema design, query optimization, and reasoning about cost and latency trade-offs on large datasets. Experience running production data infrastructure (orchestration, observability, on-call, data quality, and incident response). Preferred Direct experience with Debezium, StarRocks, and dbt in production. Experience building semantic layers (dbt Semantic Layer, Cube) or data catalogs / lineage (DataHub, OpenMetadata, Amundsen). Experience with HIPAA-regulated data (PHI handling, de-identification, and access governance). Experience powering AI/ML workloads: feature stores, training-set curation, embedding pipelines, or
Employer postings · Data from · Sources