AI Platform Engineer, Training and Inference
No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.
Technologies
Ray · KubeRay · GKE · vLLM · SGLang · NVIDIA Triton · PyTorch · Flyte · MLflow · Qdrant
About Saviynt
AI-powered identity security cloud governing human, non-human, and AI-agent access to apps and data for Fortune 500 enterprises and governments.
Series B · 1000–2000 people
Job description
The full responsibilities and requirements are on the employer’s site.
Open application page ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com. The AI Platform team is building the compute layer that trains, evaluates, and serves every AI model at Saviynt. We need an ML Platform Engineer to own distributed training on Ray + H100s, the multi-engine LLM inference mesh (vLLM,
More from the job description
AI Platform Engineer – Training & Inference Saviynt's AI-powered identity platform manages and governs human and non-human access to all of an organization's applications, data, and business processes. Customers trust Saviynt to safeguard their digital assets, drive operational efficiency, and reduce compliance costs. Built for the AI age, Saviynt is today helping organizations safely accelerate their deployment and usage of AI. Saviynt is recognized as the leader in identity security, with solutions that protect and empower the world's leading brands, Fortune 500 companies and government institutions. For more information, please visit www.saviynt.com. The AI Platform team is building the compute layer that trains, evaluates, and serves every AI model at Saviynt. We need an ML Platform Engineer to own distributed training on Ray + H100s, the multi-engine LLM inference mesh (vLLM, SGLang, NVIDIA Triton), and the full model promotion lifecycle — from shadow mode through canary rollout to GA. The AI Platform team's mission is to build a secure, scalable, product-agnostic AI foundation that enables Saviynt's identity products to deliver measurable AI-powered outcomes. Training & Inference is the engine — it turns data into deployed models that make Saviynt's products smarter. What You Will Be Doing • Own the Ray ecosystem end-to-end: manage KubeRay on GKE, tune Ray Core Task/ [... source excerpt omitted ...] ector similarity search, context assembly, and prompt construction before LLM inference What You Bring • Experience in ML engineering with time in an ML platform or MLOps role • Production Ray depth: Ray Train, Serve, Core, and Data — debugged real production failures including NCCL timeouts, Plasma OOM, and Serve autoscaling lag • LLM serving engines: hands-on with vLLM, SGLang, or NVIDIA Triton — PagedAttention, prefix caching, and continuous batching tuned for latency/throughput targets • Distributed training: DDP, FSDP, NCCL collectives, gradient checkpointing, and mixed precision (BF16/FP8) • RL working knowledge: PPO, policy gradient, or RLHF — able to translate an algorith
Employer postings · Data from · Sources