Fellow Software Engineer — AI Performance & Reliability
AI summary of the role
Principal/Fellow-level software engineer on AMD's AI Infrastructure team, responsible for improving performance, efficiency, and reliability of AI workloads (training and inference) across large language, diffusion, and recommendation models.
What you’ll do
- Profile and optimize AI model training and inference workloads to improve throughput, latency, memory efficiency, scalability, and reliability.
- Identify bottlenecks across models, frameworks, compilers, runtimes, operating systems, and hardware.
- Develop performance tooling, benchmarks, automation, and observability systems.
- Collaborate with customers to understand requirements, reproduce issues, and translate feedback into product improvements.
What you’ll bring
- Strong software engineering skills and experience building production-quality systems.
- Experience with AI infrastructure for model training, inference, or both.
- Demonstrated experience profiling and optimizing ML models or AI workloads.
- Strong foundations in computer architecture and systems performance concepts.
Technologies
PyTorch · TensorFlow · JAX · ROCm · HIP · CUDA · Triton · XLA · MLIR · NCCL
Source and classification
Implementation & delivery · Evidence for this classification:
reliability of AI workloads across both model training and inference. Our team supports a broad range of machine learning systems, including large language models, diffusion models, and recommendation models. You will collaborate closely with customers and internal engineering teams to understand performance bottlenecks, optimize workloads, and ensure that models run reliably at scale. This role is a strong fit for an engineer who enjoys working across the AI software and hardware stack, solving technically challenging performance problems, and partnering directly with customers to make them successful. You will help customers achieve meaningful improvements in model performance and system reliability. You will identify difficult bottlenecks, develop reusable solutions, and help shape the infrastructure and product capabilities needed to run demanding AI workloads efficiently atHow jobs are selected
Employer postings · Data from · Sources