Staff Machine Learning Platform Engineer
AI summary of the role
Design, improve, and operate a scalable ML platform at a wholesale marketplace.
What you’ll do
- Design and operate ML infrastructure including workspaces, clusters, jobs, and workflows
- Productionize ML workloads using Spark, Delta Lake, MLflow, and Databricks Workflows
- Implement Unity Catalog for data governance, lineage, access control, and secure multi-tenant usage
- Build CI/CD pipelines for ML using Terraform and Git-based workflows (e.g., GitHub Actions)
What you’ll bring
- 8+ years of experience building production ML or data platforms
- Strong hands-on expertise with Databricks, Spark, Delta Lake, and MLflow
- Proficiency in Python, SQL, and distributed systems concepts
- Experience with cloud platforms and infrastructure-as-code
Technologies
Databricks · Spark · Delta Lake · MLflow · Unity Catalog · Terraform · GitHub Actions · Python · SQL · AWS
About Faire
Wholesale marketplace connecting independent retailers and brands with discovery, ordering, payments, logistics, and inventory tools in one platform.
Series G · 1000–2000 people
Source and classification
Internal deployment & tooling · Evidence for this classification:
design, improve, and operate a scalable ML platform to accelerate model training, deployment, and governance. You are the technical bridge between data science and production engineering. You’ll be joining a small but deeply critical team that scales Faire’s ability to support tens of thousands of local businesses in a constantly narrowing retail landscape. What You Will Do Design and operate ML infrastructure, including workspaces, clusters, jobs, and workflows Productionize ML workloads using Spark, Delta Lake, MLflow, and Databricks Workflows Teach data scientists how to utilize our ML platform to advance development from notebook to production for our most critical models Implement Unity Catalog for data governance, lineage, access control, and secure multi-tenant usage Build CI/CD pipelines for ML using Terraform and Git-based workflows (e.g., GitHub Actions) Optimize
More from the job description
About Faire Faire is a technology wholesale platform built on the belief that the future is local. Independent retailers around the globe collectively represent a multi-hundred-billion-dollar wholesale market that has historically been fragmented and offline. At Faire, we're using the power of tech, data, and machine learning to connect this thriving community of entrepreneurs across the globe. Picture your favorite boutique in town — we help them discover the best products from around the world to sell in their stores. With the right tools and insights, we believe that we can level the playing field so businesses can grow and local communities can thrive. We’re looking for smart, resourceful and passionate people to join us as we power the shop local movement. If you believe in community, come join ours. About this role As a Staff Machine Learning Platform Engineer, you will help design, improve, and operate a scalable ML platform to accelerate model training, deployment, and governance. You are the technical bridge between data science and production engineering. You’ll be joining a small but deeply critical team that scales Faire’s ability to support tens of thousands of local businesses in a constantly narrowing retail landscape. What You Will Do Design and operate ML infrastructure, including workspaces, clusters, jobs, and workflows Productionize ML workloads using [... source excerpt omitted ...] lish observability for data quality, model performance, and platform health Build and maintain ML Platform technical documentation What it takes 8+ years of experience building production ML or data platforms A degree (preferably graduate level) in Computer Science, Engineering, Statistics, or a related technical field Strong hands-on expertise with Databricks, Spark, Delta Lake, and MLflow. Proficiency in Python, SQL, and distributed systems concepts Experience with cloud platforms and infrastructure-as-code Solid understanding of MLOps best practices: CI/CD, monitoring, reproducibility, and security Experience supporting multiple ML teams in a shared platform environment [... source excerpt omitted ...] Workplace and Information Technology positions may require onsite attendance 5 days per week as will be indicated in the job posting. Why you’ll love working at Faire Move fast: You'll own meaningful problems that serve customers around the globe with the agency to move fast and see your results clearly. Equipped to scale: We invest in what matters, including the latest enterprise AI tools, to help you work smarter and get more out of every day. Best in class: Our team is full of sharp, kind, and generous colleagues who care about their craft and about helping you grow in yours. Real rewards. Competitive pay, equity, and comprehensive benefits designed to support your life
Employer postings · Data from · Sources