Staff Software Engineer, ML Infrastructure
AI summary of the role
Voxel is hiring a Staff Software Engineer to own ML infrastructure for its computer vision platform, which processes thousands of cameras across Fortune 500 customers in manufacturing, logistics, retail, and pharma.
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
What you’ll do
- Set technical direction for ML infrastructure: build vs. buy decisions and architecture as the team scales
- Architect and build training infrastructure for concurrent experiments on PyTorch and AWS
- Own train-to-deploy handoff: export models to TensorRT/ONNX, quantify accuracy/latency tradeoffs
- Select and roll out experiment tracking stack (Weights & Biases, MLflow, ClearML, etc.)
What you’ll bring
- 7+ years building large-scale software systems, with 3+ years in ML infrastructure or large-scale data infrastructure
- Track record of owning architecture decisions (tool selection, framework choices, build-vs-buy)
- Deep fluency in PyTorch and modern ML training stack
- Strong Python skills with performant, maintainable production code
Technologies
PyTorch · AWS · TensorRT · ONNX · Weights & Biases · MLflow · ClearML · Ray · Sematic · Flyte · Metaflow · Prefect
About Voxel
Computer-vision platform that converts existing security cameras into real-time hazard detection for Fortune 500 warehouses, factories, and logistics sites to prevent workplace incidents.
Series B · 100–200 people
Source and classification
Internal deployment & tooling · Evidence for this classification:
time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team. We're hiring a Staff Software Engineer to own ML Infrastructure at Voxel. Our applied ML team is shipping vision models into production every week, across thousands of cameras at Fortune 500 customers, and the infrastructure underneath determines how fast we can move. You'll set the technical direction for how we train, track, and ship vision models, build the foundational systems that the applied ML team relies on, and shape the architectural decisions that will define our ML stack for the next several years. This is a hands-on role. You'll write code, make architecture
More from the job description
Who We Are Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs. About the Role Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team. We're hiring a Staff Software Engineer to own ML Infrastructure at Voxel. Our applied ML team is shipping vision models into production every week, across thousands of cameras at Fo [... source excerpt omitted ...] foundational systems that the applied ML team relies on, and shape the architectural decisions that will define our ML stack for the next several years. This is a hands-on role. You'll write code, make architecture calls, and own outcomes end to end. You'll partner closely with applied CV engineers, the ML Data team, and the Platform team, and you'll be the technical voice in the room when ML infrastructure tradeoffs come up. What You'll Do Set the technical direction for ML infrastructure at Voxel: what we build, what we buy, and how the pieces fit together as the team and model portfolio scale Architect and build the training infrastructure that lets the applied ML team [... source excerpt omitted ...] h, AWS) Own the train-to-deploy handoff: export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment Pick and roll out the experiment tracking and lifecycle stack (Weights & Biases, MLflow, ClearML, or similar) so researchers can run, compare, and reproduce experiments efficiently Establish DevOps-for-ML best practices (IaC, CI/CD, observability, cost monitoring) so researchers can iterate quickly and safely Mentor engineers across Vision & AI on ML infrastructure best practices, raising the bar for how the org thinks about training, evaluation, and deployment Anticipate where t
Employer postings · Data from · Sources