Senior Software Engineer, ML Infrastructure
AI summary of the role
Voxel is hiring a Senior Software Engineer to own the ML Infrastructure that powers training and shipping of vision models for computer vision applications in operations, risk, and safety.
What you’ll do
- Build and maintain training infrastructure for concurrent model training and rapid iteration.
- Own the train-to-deploy handoff, exporting models to optimized inference formats (TensorRT, ONNX) and quantifying accuracy/latency impact.
- Establish ML experiment tracking and lifecycle management using tools like Weights & Biases, MLflow, or ClearML.
- Establish DevOps-for-ML best practices on AWS (IaC, CI/CD, observability, cost monitoring).
What you’ll bring
- 4+ years building and shipping large scale software solutions.
- Hands-on experience building ML training pipelines in PyTorch.
- Hands-on experience with ML experiment tracking and lifecycle tools (Weights & Biases, MLflow, ClearML, or similar).
- Experience with AWS (S3, EC2, EKS, or similar) for ML workloads.
Technologies
PyTorch · TensorRT · ONNX · Weights & Biases · MLflow · ClearML · AWS · S3 · EC2 · EKS
About Voxel
Computer-vision platform that converts existing security cameras into real-time hazard detection for Fortune 500 warehouses, factories, and logistics sites to prevent workplace incidents.
Series B · 100–200 people
Source and classification
Internal deployment & tooling · Evidence for this classification:
time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team. We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train multiple models concurrently, manage experiments and ship optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers. What You'll Do Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new
More from the job description
Who We Are Voxel is building the future of Computer Vision and Machine Learning for operations, risk, and safety. We use computer vision and AI to enable existing security cameras to automatically detect hazards and high-risk activities, keep people safe and drive operational efficiencies. Our technology addresses the key cost drivers for workers’ compensation, general liability, and property damage, which cost US employers over $500 billion annually. Our customers include Fortune 500 companies across grocery, retail, manufacturing, food and beverage, logistics, and pharmaceutical distribution. We’ve passed $10M ARR with strong expansion revenue. Based in SF, backed by industry-leading VCs. About the Role Voxel’s perception system is the technical core of everything we ship. Our models detect human activity, equipment interactions, environmental hazards, and operational state in real time across thousands of cameras in manufacturing, logistics, retail, and pharmaceutical environments. Safety was our wedge; it proved our platform works. Now customers are pulling us into operations: equipment utilization, workflow compliance, process efficiency. Every new use case runs through the perception team. We're hiring a strong software engineer to own the ML Infrastructure that powers how Voxel trains and ships vision models. You’ll build systems that let our applied ML team train mu [... source excerpt omitted ...] hip optimized models to production. You'll set technical direction, write code, make architecture calls, and partner closely with applied CV, ML Data and Platform engineers. What You'll Do Build and maintain training infrastructure that lets the applied ML team train multiple models concurrently, manage experiments, and iterate quickly on new architectures. Own the train-to-deploy handoff - export trained models to optimized inference formats (TensorRT, ONNX), quantify accuracy and latency impact, and partner with Platform on production deployment. Establish ML experiment tracking and lifecycle management - pick the right tools (Weights & Biases, MLflow, ClearML, or similar [... source excerpt omitted ...] tools (Weights & Biases, MLflow, ClearML, or similar). Experience with AWS (S3, EC2, EKS, or similar) for ML workloads. Strong Python. Write performant code that scales well in production environments. Track record of owning infrastructure end-to-end: scoping, building, shipping, and improving systems that internal teams depend on. Bias toward shipping. You'd rather ship something good this week than something perfect next quarter. Strong communication skills. Nice to Have Experience with modern ML orchestration tools (Ray, Sematic, Flyte, Metaflow, Prefect, or similar) Familiarity with GPU performance profiling and optimization (Nsight, PyTorch profiler, or similar) Back
Employer postings · Data from · Sources