Skip to content
NEXTMOVEFDE careers · United States

Solutions Architect - AI Inference Specialist

AI summary of the role

FriendliAI seeks a Solutions Architect to help enterprises deploy, scale, and operate generative/agentic AI workloads on its inference platform.

No longer in the current catalog. Last included 2026-09-09. Check the employer’s posting for availability.

What you’ll do

  • Design and implement large-scale deployment architectures for LLM and multimodal inference
  • Deploy and manage containerized workloads across Kubernetes clusters
  • Diagnose production issues and implement temporary fixes as needed
  • Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows

What you’ll bring

  • 3+ years of experience in cloud infrastructure, DevOps, or reliability engineering
  • Proficiency with Kubernetes, Docker, Terraform, and Helm
  • Experience with GPU-based computing and generative AI model serving workloads
  • Experience operating workloads on AWS, GCP, or OCI

Technologies

Kubernetes · Docker · Terraform · Helm · AWS · GCP · OCI · Triton · vLLM · TensorRT · DeepSpeed-Inference · Prometheus

About FriendliAI

Generative AI inference platform selling fast, cost-optimized LLM/agent serving (managed endpoints + on-prem engine) to enterprises.

Seed · 50–200 people

Source and classification

Implementation & delivery · Evidence for this classification:

About the job FriendliAI is seeking a Solution Architect to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container. Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product. You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps,
More from the job description

About the job FriendliAI is seeking a Solution Architect to assist enterprises in deploying, scaling, and operating generative and agentic AI workloads on FriendliAI infrastructure. You will work directly with customers to solve and implement production-grade applications using our products, such as Serverless Endpoints, Dedicated Endpoints, or Container. Friendli Container is our service that allows customers to download our inference engine as Docker images and deploy it in their chosen environment, such as private clouds or on-premises. Our Friendli Container can be adopted directly to AWS EKS clusters using our EKS add-on product. You will work directly on our customers’ projects, collaborating with their engineering teams to solve AI inference challenges like scaling, orchestration, and monitoring. This is a hands-on, customer-embedded role. If you have worked in DevOps, platform engineering, or SRE for AI applications, this is your ideal position. Key Responsibilities Design and implement large-scale deployment architectures for LLM and multimodal inference Deploy and manage containerized workloads across Kubernetes clusters Diagnose production issues, such as performance bottlenecks, and implement temporary fixes as needed Collaborate with customers’ DevOps teams to integrate FriendliAI’s infrastructure into their CI/CD workflows Develop scripts, Helm charts, an [... source excerpt omitted ...] y repeated deployments Contribute field insights to shape our platform reliability, observability, and scaling strategies Lead workshops, technical sessions, or webinars to help customers master infrastructure best practices. Qualifications 3+ years of experience in cloud infrastructure, DevOps, or reliability engineering Bachelor’s or Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or equivalent Proficiency with Kubernetes, Docker, Terraform, and Helm Strong foundation in distributed systems, networking, and performance tuning Experience with GPU-based computing and generative AI model serving workloads Strong technical background in b

How jobs are selected

Employer postings · Data from · Sources