GCP Python Data Engineer
Job description
The full responsibilities and requirements are on the employer’s site.
Read the job description ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
in the United States without current or future sponsorship. No visa sponsorship, transfers, or C2C arrangements are available. Key Responsibilities Data Engineering & Pipeline Development Design, develop, and optimize ETL/ELT pipelines for structured and unstructured data. Build scalable batch and streaming data processing solutions using GCP technologies. Develop event-driven data processing solutions leveraging Pub/Sub and Cloud Functions. Create and maintain data ingestion frameworks for enterprise data platforms. Data Storage & Analytics Design and optimize data lake, lakehouse, and data warehouse solutions. Build efficient data models supporting analytics, reporting, and AI/ML workloads. Optimize performance, scalability, and cost efficiency of data pipelines and queries. Development & Automation Develop robust Python-based solutions and frameworks. Automate
More from the job description
Choosing Capgemini means choosing a company where you will be empowered to shape your career in the way you’d like, where you’ll be supported and inspired by a collaborative community of colleagues around the world, and where you’ll be able to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more sustainable, more inclusive world. About the Role We are seeking a highly skilled GCP Python Data Engineer to design, build, and optimize scalable cloud-based data solutions on Google Cloud Platform (GCP). The ideal candidate will possess strong Python development skills, hands-on experience with modern data engineering technologies, and expertise in building both batch and real-time data pipelines supporting analytics, AI/ML, and enterprise reporting initiatives. Work Authorization: Candidates must be authorized to work in the United States without current or future sponsorship. No visa sponsorship, transfers, or C2C arrangements are available. Key Responsibilities Data Engineering & Pipeline Development Design, develop, and optimize ETL/ELT pipelines for structured and unstructured data. Build scalable batch and streaming data processing solutions using GCP technologies. Develop event-driven data processing solutions leveraging Pub/Sub and Cloud Functions. Create and maintain data ingestion frameworks fo [... source excerpt omitted ...] nitoring, and observability. Collaboration & Support Partner with business stakeholders, analytics teams, data scientists, and engineers to deliver data solutions. Troubleshoot production issues and perform root cause analysis. Continuously improve reliability, scalability, security, and operational excellence. Required Qualifications 5+ years of data engineering, software engineering, or related experience. 2+ years of hands-on Google Cloud Platform (GCP) experience. 2+ years of professional Python development experience. Experience developing batch and real-time data pipelines. Advanced SQL development and query optimization skills. Experience with data warehousing, ET [... source excerpt omitted ...] Data Engineering Python SQL Bash/Shell Scripting ETL/ELT Data Warehousing Data Lake / Lakehouse Architectures Batch Processing Real-Time Streaming Architectures Preferred Qualifications Experience with Vertex AI and Google AI services. Experience building AI/ML data platforms and pipelines. Knowledge of Large Language Models (LLMs) and GenAI concepts. Experience with Gemini models, Agentic AI frameworks, Prompt Engineering, RAG architectures, and Vector Search. Experience with Dataproc, Spark, or PySpark. Familiarity with event-driven architectures. Experience with Terraform or Infrastructure as Code. Understanding of cloud cost optimization and FinOps practices. Fina
Employer postings · Data from · Sources