Member of Technical Staff (Software Engineer, Data Platform)
AI summary of the role
Senior/Staff engineer on the Data Platform team owning the end-to-end data lifecycle at Perplexity, from ingestion to serving.
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
What you’ll do
- Design and operate large-scale batch and streaming data pipelines powering product features, AI training, analytics, and experimentation.
- Build event-driven and streaming systems (Kafka, Kinesis, PubSub) for real-time ingestion and transformation.
- Lead architecture of data orchestration using Airflow or Dagster, owning scheduling, SLAs, and observability.
- Set guarantees for data correctness, freshness, lineage, and recoverability at scale.
What you’ll bring
- 5+ years (Senior) or 8+ years (Staff) of software engineering experience.
- Strong experience building production data infrastructure systems.
- Hands-on experience with batch and/or streaming data processing at scale.
- Deep familiarity with data orchestration systems (Airflow, Dagster, or similar).
Technologies
Databricks · Snowflake · Spark · Kafka · Flink · Airflow · Dagster · dbt · Iceberg · Delta Lake · ClickHouse · Python
About Perplexity
Conversational AI answer engine and agentic browser (Comet) that researches, shops, and executes tasks for consumers and enterprises.
Series E
Source and classification
Internal deployment & tooling · Evidence for this classification:
About the Role The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake. The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse. In this senior/staff role, you will shape architecture, set standards, and drive the long-term technical direction of Perplexity’s data ecosystem. Key Responsibilities Design and operate large-scale batch and streaming data pipelines that directly power Perplexity product features, AI training and evaluation
More from the job description
About the Role The Data Platform team owns the end-to-end data lifecycle at Perplexity, from ingestion through processing, storage, and serving, powering product features, analytics, experimentation, AI workloads, and the company’s data lake. The team defines the architecture for batch and streaming systems, the orchestration and observability stack, and a self-serve data platform, while thoughtfully combining platforms such as Databricks and Snowflake with open-source technologies including Spark, Kafka, Flink, Airflow, Dagster, dbt, Iceberg, Delta Lake, and ClickHouse. In this senior/staff role, you will shape architecture, set standards, and drive the long-term technical direction of Perplexity’s data ecosystem. Key Responsibilities Design and operate large-scale batch and streaming data pipelines that directly power Perplexity product features, AI training and evaluation workflows, analytics, and experimentation. Build event-driven and streaming systems (Kafka, Kinesis, PubSub, or similar) for real-time ingestion, transformation, and delivery, alongside batch frameworks for backfills, aggregations, and offline computation. Lead the architecture of data orchestration using tools like Airflow or Dagster, owning scheduling, dependency management, retries, SLAs, and end-to-end observability for critical data flows. Set and enforce guarantees for data correctness, freshn [... source excerpt omitted ...] rplexity’s roadmap. Mentor engineers, review designs, and raise the technical bar for data infrastructure through thoughtful feedback, documentation, and hands-on collaboration. Qualifications 5+ years (Senior) or 8+ years (Staff) of software engineering experience. Strong experience building production data infrastructure systems. Hands-on experience with batch and/or streaming data processing at scale. Deep familiarity with data orchestration systems (Airflow, Dagster, or similar). Proficiency in Python and at least one additional backend language (Go, TypeScript, etc.). Strong systems thinking around reliability, latency, cost, and complexity tradeoffs. Experience supportin
Employer postings · Data from · Sources