Cloud Platform Architect
Technologies
AWS · GCP · Azure · Kubernetes · Docker · Terraform · Ansible · Prometheus · Grafana · Datadog · Python · Go
About SambaNova Systems
Full-stack AI platform from chip to cloud for enterprise inference, built on proprietary RDU silicon and optimized for on-prem or cloud deployment with customer data ownership.
Series E · 500–1000 people
Job description
The full responsibilities and requirements are on the employer’s site.
Read the job description ↗Source and classification
Internal deployment & tooling · Evidence for this classification:
service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America. About the role The Cloud Operations team is seeking an experienced engineering leader to scale the platform our internal and external customers use to access SambaNova RDUs. Responsibilities In this role, you'll architecting our next-generation system from the ground up, running the Kubernetes infrastructure that powers some of the most advanced AI workloads in the industry, and bridging multi-cloud and on-prem environments in ways no generic SaaS company can offer. Your work will directly impact the productivity of every engineer at SambaNova and by extension, the speed at which we ship the future of AI computing. Architect, build, and maintain our next-generation
More from the job description
The era of pervasive AI has arrived. In this era, organizations will use generative AI to unlock hidden value in their data, accelerate processes, reduce costs, drive efficiency and innovation to fundamentally transform their businesses and operations at scale. SambaNova Suite™ is the first full-stack, generative AI platform, from chip to model, optimized for enterprise and government organizations. Powered by the intelligent SN40L chip, the SambaNova Suite is a fully integrated platform, delivered on-premises or in the cloud, combined with state-of-the-art open-source models that can be easily and securely fine-tuned using customer data for greater accuracy. Once adapted with customer data, customers retain model ownership in perpetuity, so they can turn generative AI into one of their most valuable assets. About the team The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America. About the role The Cloud Operations team is seeking an experienced engineering leader to scale the platform our internal and external customers use to access SambaNova RDUs. Responsibilities In this role, you'll architecting our next-generation system from the ground up, running the [... source excerpt omitted ...] vice meshes) that seamlessly connect our multi-cloud and hybrid environments Collaborate with AI and software engineering teams to understand their needs, provide golden paths to production, and build internal tools that accelerate their development cycles Implement best practices for observability (monitoring, logging, tracing) to ensure system reliability and performance, and participate in on-call rotation Required Qualifications 7+ years of experience in DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure roles Proficiency in at least one programming language (e.g., Python, Go, Rust) Expertise with Kubernetes (EKS, GKE, or self-managed) in production envir [... source excerpt omitted ...] rstanding of the core services (compute, storage, networking, IAM) Networking fundamentals (TCP/IP, DNS, HTTP, load balancing) and security best practices in the cloud Preferred Qualifications Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure Experience managing infrastructure for data-intensive or ML/AI workloads Knowledge of building and maintaining CI/CD pipelines (e.g., GitLab CI, Jenkins, ArgoCD) Experience with service mesh technologies (e.g., Istio, Linkerd) Contributions to open-source projects or a public portfolio of code (GitHub) Base Salary Range: Base Pay Range $245,000—$325,000 USD Submission Guidelines Please note that
Employer postings · Data from · Sources