Skip to content
NEXTMOVEFDE careers · United States

Principal Software Engineer, AI Networking

AI summary of the role

NVIDIA seeks a Principal Software Engineer to lead the technical strategy for AI networking deployments at hyperscaler and AI Factory customers.

What you’ll do

  • Lead technical strategy for AI Factory networking deployments at strategic customers, including architecture reviews, risk assessments, and multi-phase execution plans.
  • Serve as principal-level technical authority for embedded networking products (BlueField, ConnectX) and ecosystem (DOCA, RDMA, RoCE, InfiniBand).
  • Lead deep technical engagements with hyperscalers and AI Factory customers covering design-in, coding, bring-up, performance tuning, failure analysis, and production hardening.
  • Partner with internal engineering, product, and architecture teams to transform customer needs into product features, reference architectures, and tooling.

What you’ll bring

  • BS/MS/PhD in Computer Science, Computer Engineering, Electrical Engineering, or equivalent experience.
  • 15+ years of relevant industry experience including technical leadership across complex systems.
  • Deep knowledge of networking protocols and distributed systems with strong understanding of RoCE/InfiniBand, L1–L4 fundamentals, and performance/latency tradeoffs.
  • Proven low-level software expertise with proficiency in C/C++ and comfort debugging across firmware, driver, and user space.

Technologies

BlueField · ConnectX · DOCA · RDMA · RoCE · InfiniBand · C/C++ · DPDK · NCCL · CUDA

About NVIDIA

Designs and manufactures GPUs and system-on-chips powering data centers, AI workloads, gaming, autonomous vehicles, and HPC. The foundational hardware for modern deep learning.

Public

Source and classification

Implementation & delivery · Evidence for this classification:

networking systems. You will apply your deep expertise to manage complex customer engagements and help develop our product and architecture direction. This role offers an outstanding opportunity to influence NVIDIA's networking technologies and make a significant impact on the industry! What you'll be doing: Lead the technical strategy for AI Factory networking deployments at strategic customers, including conducting architecture reviews, risk assessments, and crafting multi-phase execution plans. Serve as the principal-level technical authority for embedded networking products like BlueField and ConnectX. This role also covers the surrounding technology ecosystem, including DOCA, RDMA, RoCE, and Infiniband. Lead deep technical engagements with hyperscalers and AI Factory customers, involving design-in, coding, bring-up, performance tuning, failure analysis, and production
More from the job description

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It’s a unique legacy of innovation that’s fueled by great technology—and amazing people. Today, we’re tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what’s never been done before takes vision, innovation, and the world’s best talent. As an NVIDIAN, you’ll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world. Join NVIDIA, where the future is defined by our innovative advances in AI, computer graphics, and accelerated computing. As a Principal Software Engineer, you will lead the transformation of AI networking systems. You will apply your deep expertise to manage complex customer engagements and help develop our product and architecture direction. This role offers an outstanding opportunity to influence NVIDIA's networking technologies and make a significant impact on the industry! What you'll be doing: Lead the technical strategy for AI Factory networking deployments at strategic customers, including conducting architecture reviews, risk assessments, and crafting multi-phase execution plans. Ser [... source excerpt omitted ...] ConnectX. This role also covers the surrounding technology ecosystem, including DOCA, RDMA, RoCE, and Infiniband. Lead deep technical engagements with hyperscalers and AI Factory customers, involving design-in, coding, bring-up, performance tuning, failure analysis, and production hardening. Partner with internal engineering, product, and architecture teams to transform customer needs into product features, reference architectures, tooling, and guidelines. Drive performance, reliability, and debuggability improvements across customer stacks and translate findings into actionable product, firmware, and software roadmap items. What we need to see: BS/MS/PhD in Computer Science [... source excerpt omitted ...] rops, retransmissions, congestion, QoS, ordering, and buffer management. Excellent interpersonal skills, with the ability to clearly explain complex topics to engineers, PMs, and customer collaborators, and align cross-organizational teams toward a decision. Ways To Stand Out from the crowd: Prior experience in customer-facing technical leadership at hyperscalers/CSPs/AI factories (or similarly complex production environments). Hands-on expertise with DPDK, DOCA, RDMA verbs, NCCL, CUDA-aware networking, congestion control, and performance tuning at scale. Experience building internal tools, telemetry, and automation that improve triage speed and operational excellence. Demo

How jobs are selected

Employer postings · Data from · Sources