Lead Data Engineer
AI summary of the role
Lead data engineer role at Empower, a US retirement and wealth management firm, focused on architecting and optimizing enterprise-scale streaming data pipelines and platforms.
What you’ll do
- Lead design and implementation of advanced data pipelines, ingestion frameworks, and real-time streaming solutions.
- Build and maintain scalable Kafka producers/consumers using Python, Java, Go, or Scala for sub-second latency.
- Configure and manage CDC solutions (Debezium, AWS DMS, Qlik Replicate) across relational and NoSQL sources.
- Develop stream-processing logic using Flink, Spark Streaming, Kafka Streams, or dbt-mesh.
What you’ll bring
- 6-8+ years of data engineering experience with enterprise scale platforms.
- 6-8+ years hands-on with Apache Kafka, Confluent Cloud, AWS Kinesis, or Apache Pulsar.
- Advanced proficiency in Python, PySpark, and SQL.
- Deep expertise with CDC tools such as Debezium and Kafka Connect.
Technologies
Apache Kafka · Confluent Cloud · AWS Kinesis · Apache Pulsar · Debezium · AWS DMS · Qlik Replicate · Apache Flink · Spark Streaming · Kafka Streams · dbt-mesh · Confluent Schema Registry
About Empower
US retirement recordkeeping giant (19.5M participants, $2T AUA) expanding into personal wealth advisory, equity compensation, and private-market DC investing.
Acquired · 5000+ people
Source and classification
Internal deployment & tooling · Evidence for this classification:
design, development, and optimization of scalable data pipelines, real-time ingestion frameworks, and data platforms. This role will lead complex data initiatives, including event-driven and streaming solutions, while ensuring performance, scalability, reliability, security, data quality, and governance align with enterprise standards. Operating with significant independence, the Lead Data Engineer will guide technical decision-making, mentor engineers, collaborate with architects and stakeholders, and serve as an escalation point for complex data engineering challenges. What you will do: Lead the design and implementation of advanced data pipelines, ingestion frameworks, and real-time streaming solutions. Build, tune, and maintain scalable Kafka producers and consumers using Python, Java, Go, or Scala to support sub-second latency and high availability across distributed services.
More from the job description
Our vision for the future is based on the idea that transforming financial lives starts by giving our people the freedom to transform their own. We have a flexible work environment, and fluid career paths. We not only encourage but celebrate internal mobility. We also recognize the importance of purpose, well-being, and work-life balance. Within Empower and our communities, we work hard to create a welcoming and inclusive environment, and our associates dedicate thousands of hours to volunteering for causes that matter most to them. Chart your own path and grow your career while helping more customers achieve financial freedom. Empower Yourself. Applicants must be authorized to work for any employer in the U.S. We are unable to sponsor or take over sponsorship of an employment visa at this time, including CPT/OPT. The Lead Data Engineer will provide technical leadership in the design, development, and optimization of scalable data pipelines, real-time ingestion frameworks, and data platforms. This role will lead complex data initiatives, including event-driven and streaming solutions, while ensuring performance, scalability, reliability, security, data quality, and governance align with enterprise standards. Operating with significant independence, the Lead Data Engineer will guide technical decision-making, mentor engineers, collaborate with architects and stakeholders, and [... source excerpt omitted ...] onsumer group offsets, partition rebalancing, and dead-letter queues. Guide the technical execution of complex data initiatives across teams and translate highly complex business requirements into scalable data solutions. Ensure data platforms are optimized for performance, scalability, reliability, and high availability. Enforce data governance, security, quality, data contract, and schema management standards. Mentor engineers, provide technical leadership on projects, and lead code reviews to ensure adherence to engineering standards. Collaborate with architects and stakeholders to align solutions with enterprise direction and contribute to architectural and design decisions. [... source excerpt omitted ...] nce with Apache Kafka, Confluent Cloud, AWS Kinesis, or Apache Pulsar. Advanced proficiency in Python, PySpark, and SQL. Proficiency in Python, Java, Scala, or Go, combined with production experience using Flink, Kafka Streams, or Spark Structured Streaming. Strong experience with AWS and modern data platforms. Deep expertise with CDC tools such as Debezium and Kafka Connect and log-based replication mechanisms such as Oracle Redo, PostgreSQL WAL, and MySQL binlog. Hands-on experience sinking streaming events into modern lakehouse formats such as Apache Iceberg and Delta Lake and cloud warehouses such as Snowflake or BigQuery. Solid foundation in Docker, Kubernetes, Terraform
Employer postings · Data from · Sources