Skip to content
NEXTMOVEFDE careers · United States

Data Engineer, Platform

AI summary of the role

Basis Research Institute seeks a Data Engineer for its Platform team to build trustworthy ML data pipelines with provenance and quality gates, curate documented datasets, and coordinate cross-project data initiatives.

What you’ll do

  • Design and build data pipelines for training and evaluation across Basis research projects and platform offerings, ensuring reliability, performance, and scalability.
  • Implement data quality frameworks including validation rules, quality gates, anomaly detection, and monitoring.
  • Develop and maintain feature stores or equivalent systems to prevent train-serve skew.
  • Ensure data provenance and lineage tracking for reproducible experiments and debugging.

What you’ll bring

  • Demonstrated significant achievements in data engineering for ML/AI systems (e.g., building data pipelines at scale, developing feature stores, creating data quality frameworks).
  • Strong proficiency in SQL (expert level), Python, distributed computing frameworks (Spark, Dask), and workflow orchestration (Airflow, Dagster, Prefect).
  • Experience with cloud data platforms (Snowflake, BigQuery, Redshift), object storage (S3), and streaming systems (Kafka, Kinesis, Flink).
  • Understanding of ML data requirements including feature engineering, data versioning, and experiment reproducibility.

Technologies

SQL · Python · Spark · Dask · Airflow · Dagster · Prefect · Snowflake · BigQuery · Redshift · S3 · Kafka

About Basis Research Institute

501(c)(3) applied-AI research institute building causal and probabilistic reasoning systems, and applying them to intractable scientific and societal problems.

Bootstrapped · 10–50 people

Source and classification

Internal deployment & tooling · Evidence for this classification:

About Basis Basis is a nonprofit applied AI research organization with two mutually reinforcing goals. The first is to understand and build intelligence. This means to establish the mathematical principles of what it means to reason, to learn, to make decisions, to understand, and to explain; and to construct software that implements these principles. The second is to advance society’s ability to solve intractable problems. This means expanding the scale, complexity, and breadth of problems that we can solve today, and even more importantly, accelerating our ability to solve problems in the future. To achieve these goals, we’re building both a new technological foundation that draws inspiration from how humans reason, and a new kind of collaborative organization that puts human values first. About the Role Data Engineers on the Platform team at Basis build trustworthy data
More from the job description

About Basis Basis is a nonprofit applied AI research organization with two mutually reinforcing goals. The first is to understand and build intelligence. This means to establish the mathematical principles of what it means to reason, to learn, to make decisions, to understand, and to explain; and to construct software that implements these principles. The second is to advance society’s ability to solve intractable problems. This means expanding the scale, complexity, and breadth of problems that we can solve today, and even more importantly, accelerating our ability to solve problems in the future. To achieve these goals, we’re building both a new technological foundation that draws inspiration from how humans reason, and a new kind of collaborative organization that puts human values first. About the Role Data Engineers on the Platform team at Basis build trustworthy data pipelines with comprehensive provenance and quality gates, curate documented datasets for training and evaluation, and ensure data infrastructure scales reliably. You will work on both platform-specific data needs and cross-project data coordination, preventing duplicate work and facilitating shared datasets. We are looking for people who are technically excellent and treat data quality as a first-class concern. The ideal Data Engineer has experience with ML data pipelines, understands the full lifecyc [... source excerpt omitted ...] valuation, and brings rigor to data provenance, lineage tracking, and quality assurance. You combine software engineering discipline with deep understanding of data systems and ML requirements. This role is embedded across Platform and Research teams, working on infrastructure that supports both commercial offerings and internal research. You will help Basis scale data operations to support medium-scale models, ensure data governance as we serve external customers, and build systems that researchers can trust for reproducible experiments. We seek individuals who aspire to do rigorous, high-quality, robust data engineering, but are not afraid to iterate, learn from real usage, and e [... source excerpt omitted ...] houses (Snowflake, BigQuery, Redshift), data lakes, object storage (S3), and streaming systems (Kafka, Kinesis, Flink) for both batch and real-time processing. Understand ML data requirements including feature engineering, training/validation/test splits, data versioning, experiment reproducibility, and the specific data needs of different model types and training procedures. Be skilled at data quality and governance including implementing validation frameworks, anomaly detection, data lineage tracking, metadata management, and ensuring compliance with privacy and security policies. Have knowledge of data modeling principles for both relational and NoSQL systems, understanding of

How jobs are selected

Employer postings · Data from · Sources