Principal Data Engineer, LLM/AI Platforms (Remote)
AI summary of the role
Lead the design and build of exabyte-scale data infrastructure powering AI-driven security products at CrowdStrike.
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
What you’ll do
- Architect and optimize data platforms for LLMs, RAG, and AI agentic systems at Exabyte scale.
- Drive adoption of agentic workflows and agent harnessing for autonomous security features.
- Write production-ready code with focus on performance, maintainability, and testing rigor.
- Establish MLOps/DataOps best practices for LLMs including monitoring and zero-touch recovery.
What you’ll bring
- 10+ years in Data Engineering/Platform Engineering with 3+ years architecting AI/ML platforms at massive scale.
- Hands-on experience in LLM engineering (fine-tuning, prompt engineering, deployment), RAG, and agentic workflows.
- Proven track record designing large-scale distributed systems (sharding, partitioning, concurrency).
- Expert-level proficiency in Python or JVM technologies.
Technologies
LLM · RAG · LangChain · LlamaIndex · Spark · Dask · Flink · Kafka · Pulsar · Snowflake · BigQuery · Airflow
About CrowdStrike
AI-native Falcon platform protects endpoints, cloud workloads, identities, data, and enterprise AI for organizations consolidating security on one sensor and data graph.
Public · 5000+ people
Source and classification
Internal deployment & tooling · Evidence for this classification:
is looking for a Principal Data Engineer with deep expertise in Large Language Models (LLMs) and AI platforms to join our growing Data Science Platform Engineering Team. You will be a key leader, responsible for designing, building, and deploying cutting-edge data infrastructure that powers our next generation of AI-driven security products. This role requires significant hands-on experience in LLM integration, agentic workflows, and agent harnessing to deliver high-impact, scalable solutions. You will champion engineering excellence, focusing on shipping fast, writing elegant, high-quality code, and actively mentoring and strengthening the team's technical knowledge and capabilities. The scale of our systems and data are approaching Exabytes in size. Experience with extremely large-scale systems, including DevSecOps patterns, practices, and standards are important for this work. What
More from the job description
As a global leader in cybersecurity, CrowdStrike protects the people, processes and technologies that drive modern organizations. Since 2011, our mission hasn’t changed — we’re here to stop breaches, and we’ve redefined modern security with the world’s most advanced AI-native platform. Our customers span all industries, and they count on CrowdStrike to keep their businesses running, their communities safe and their lives moving forward. We’re also a mission-driven company. We cultivate a culture that gives every CrowdStriker both the flexibility and autonomy to own their careers. We’re always looking to add talented CrowdStrikers to the team who have limitless passion, a relentless focus on innovation and a fanatical commitment to our customers, our community and each other. Ready to join a mission that matters? The future of cybersecurity starts with you. About the Role: CrowdStrike is looking for a Principal Data Engineer with deep expertise in Large Language Models (LLMs) and AI platforms to join our growing Data Science Platform Engineering Team. You will be a key leader, responsible for designing, building, and deploying cutting-edge data infrastructure that powers our next generation of AI-driven security products. This role requires significant hands-on experience in LLM integration, agentic workflows, and agent harnessing to deliver high-impact, scalable solutions. Yo [... source excerpt omitted ...] and data are approaching Exabytes in size. Experience with extremely large-scale systems, including DevSecOps patterns, practices, and standards are important for this work. What You'll Do: Architect, implement, and optimize data platforms and pipelines specifically designed to support LLMs, Retrieval-Augmented Generation (RAG), and sophisticated AI agentic systems at Exabyte scale. Drive the adoption and deployment of agentic workflows and agent harnessing techniques to create autonomous, data-driven security features. Design and implement highly scalable, fault-tolerant, and cost-effective data solutions, emphasizing rapid iteration and high-quality deployment. Write ele [... source excerpt omitted ...] ge AI platform technologies. Collaborate across the organization with Data Scientists, Product Managers, and other engineering teams to transform research prototypes into robust, production-grade services. Own the end-to-end lifecycle of critical data services: development, testing, deployment, and monitoring. Tech Stack (Expertise in several key areas is expected): MLOps Tools (MLflow, Sagemaker, Vertex AI) Experience with common agentic workflow frameworks (e.g., LangChain, LlamaIndex). Expert-level proficiency in a high-level coding language (Python, or JVM technologies). Deep experience with distributed data processing frameworks (e.g., Spark, Dask, Flink). Strong exper
Employer postings · Data from · Sources