Skip to main content
Loading
loadingbar
Loading, Please wait..!!

GPU Network Engineer

  • Job type Posted on: Jul 22, 2026
  • Experience level Blue Signal Search
  • Employment type Santa Clara, California
  • Employment type Onsite
  • Salary Full-time

Point Apply Here APPLY LATER

Curious about compensation?

Explore the historical salary trends, average pay, and estimated compensation for GPU Network Engineer roles in California.

View Salary Guide →

Job Title :

GPU Network Engineer

Job Type :

Full-time

Job Location :

Santa Clara California United States

Remote :

No

Jobcon Logo Job Description :

An innovative technology organization at the forefront of AI infrastructure is seeking a GPU Network Engineer to help design and operate the high-performance networking backbone powering advanced GPU computing environments. This opportunity is ideal for an engineer who enjoys solving complex networking challenges, optimizing ultra low latency communication, and building highly scalable infrastructure supporting large AI training and inference workloads. You will work alongside experienced infrastructure and compute professionals while influencing the architecture of next generation GPU clusters. What You Will Do Design, deploy, and optimize high performance InfiniBand (NDR and XDR) and RoCEv2 fabrics supporting large scale GPU and AI workloads. Build scalable CLOS and ECMP network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments. Implement and support EVPN, VXLAN, Layer 2, and Layer 3 networking for modern data center infrastructure. Configure and maintain Arista, Juniper, and NVIDIA networking platforms while ensuring highly available, resilient network operations. Automate network provisioning, configuration, and lifecycle management using Netconf, Ansible, Terraform, and NetBox. Monitor, troubleshoot, and optimize network performance through packet analysis, telemetry, and root cause analysis. Partner with infrastructure, compute, and platform engineering teams to support AI training, inference, and high performance computing initiatives. Create and maintain technical documentation, operational procedures, and implementation standards while participating in change management and infrastructure improvements. Required Qualifications Minimum of 5 years of experience supporting enterprise data center or HPC networking environments with direct GPU cluster networking experience. Extensive hands on experience with InfiniBand (NDR or XDR) and RoCEv2 networking technologies. Strong expertise designing and supporting east west networking for GPU or AI infrastructure. Experience with CLOS architectures, ECMP routing, EVPN, and VXLAN. Hands on administration of Arista EOS, Juniper Junos, and NVIDIA or Mellanox networking platforms. Experience using NetBox, Netconf, Ansible, and Terraform to automate network deployment and management. Strong troubleshooting skills, packet analysis experience, and network performance optimization capabilities. Excellent communication and technical documentation skills. Preferred Qualifications Experience optimizing fabrics supporting large scale AI training workloads. Knowledge of adaptive routing, congestion management, RDMA optimization, and lossless networking. Python scripting for infrastructure automation. Industry certifications such as CCNP, CCIE, JNCIP, JNCIE, or equivalent practical experience. #J-18808-Ljbffr

Jobcon Logo Position Details

Posted:

Jul 22, 2026

Reference Number:

14660_C34858E1CE378FDB35BBF7806ECDF22D

Employment:

Full-time

Salary:

Not Available

City:

Santa Clara

Job Origin:

APPCAST_CPC

Share this job:

  • linkedin

Jobcon Logo
A job sourcing event
In Dallas Fort Worth
Aug 19, 2017 9am-6pm
All job seekers welcome!

GPU Network Engineer    Apply

Click on the below icons to share this job to Linkedin, Twitter!

An innovative technology organization at the forefront of AI infrastructure is seeking a GPU Network Engineer to help design and operate the high-performance networking backbone powering advanced GPU computing environments. This opportunity is ideal for an engineer who enjoys solving complex networking challenges, optimizing ultra low latency communication, and building highly scalable infrastructure supporting large AI training and inference workloads. You will work alongside experienced infrastructure and compute professionals while influencing the architecture of next generation GPU clusters. What You Will Do Design, deploy, and optimize high performance InfiniBand (NDR and XDR) and RoCEv2 fabrics supporting large scale GPU and AI workloads. Build scalable CLOS and ECMP network architectures that deliver low latency, high bandwidth east west traffic across distributed GPU environments. Implement and support EVPN, VXLAN, Layer 2, and Layer 3 networking for modern data center infrastructure. Configure and maintain Arista, Juniper, and NVIDIA networking platforms while ensuring highly available, resilient network operations. Automate network provisioning, configuration, and lifecycle management using Netconf, Ansible, Terraform, and NetBox. Monitor, troubleshoot, and optimize network performance through packet analysis, telemetry, and root cause analysis. Partner with infrastructure, compute, and platform engineering teams to support AI training, inference, and high performance computing initiatives. Create and maintain technical documentation, operational procedures, and implementation standards while participating in change management and infrastructure improvements. Required Qualifications Minimum of 5 years of experience supporting enterprise data center or HPC networking environments with direct GPU cluster networking experience. Extensive hands on experience with InfiniBand (NDR or XDR) and RoCEv2 networking technologies. Strong expertise designing and supporting east west networking for GPU or AI infrastructure. Experience with CLOS architectures, ECMP routing, EVPN, and VXLAN. Hands on administration of Arista EOS, Juniper Junos, and NVIDIA or Mellanox networking platforms. Experience using NetBox, Netconf, Ansible, and Terraform to automate network deployment and management. Strong troubleshooting skills, packet analysis experience, and network performance optimization capabilities. Excellent communication and technical documentation skills. Preferred Qualifications Experience optimizing fabrics supporting large scale AI training workloads. Knowledge of adaptive routing, congestion management, RDMA optimization, and lossless networking. Python scripting for infrastructure automation. Industry certifications such as CCNP, CCIE, JNCIP, JNCIE, or equivalent practical experience. #J-18808-Ljbffr

Loading
Please wait..!!