Network Engineer (AI GPU Cluster Operations)
aquila hashBuffalo, NY
5 days ago
Occupations
Network and Computer Systems AdministratorsComputer Network Support SpecialistsComputer Network ArchitectsIndustries
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesComputer Systems Design ServicesComputer Facilities Management ServicesAbout the role
Position Overview The AI GPU Cluster Operations – Network Engineer is responsible for the operations, maintenance and troubleshooting of high-performance networking infrastructure supporting large-scale AI GPU computing clusters. The role focuses on Infini Band, RoCE and data center IP networks, while also requiring a strong understanding of GPU servers, Linux and the overall AI cluster architecture. The engineer will monitor network and cluster health, troubleshoot connectivity and performance issues, maintain high-speed network infrastructure, support production incidents, implement approved network changes and work with internal engineering teams and OEM vendors to drive technical issues through resolution. Key Responsibilities• Operate and maintain high-performance GPU cluster networks, including Infini Band and RoCE, as well as standard data center IP networks.• Manage and troubleshoot NVIDIA/Mellanox Infini Band switches, high-speed Ethernet switches, NICs/HCAs, optical transceivers, cables and related network infrastructure.• Monitor network health and identify link degradation, link flapping, congestion, packet loss, bandwidth or latency issues.• Perform Infini Band health checks and diagnostics using tools such as ibstat, ibdiagnet and other NVIDIA/Mellanox diagnostic utilities.• Support Infini Band Subnet Management, partition configuration, routing, link-health monitoring and performance-counter analysis.• Troubleshoot RoCE environments, including connectivity, congestion and lossless Ethernet-related issues.• Support data center IP network operations, including BGP, OSPF, VLAN, MLAG and TCP/IP.• Troubleshoot GPU node connectivity and determine whether infrastructure issues originate from the server, NIC/HCA, optics, cabling, switch or network fabric.• Work with GPU Cluster Operations Engineers to troubleshoot GPU-to-GPU and node-to-node communication issues, including network-related NCCL failures and performance degradation.• Monitor network and infrastructure health using Prometheus, SNMP, Grafana and related monitoring platforms.• Develop and maintain network monitoring, alerting and diagnostic capabilities.• Perform approved network configuration changes, maintenance and upgrades according to established change-management procedures and Standard Operating Procedures (SOPs).• Support network capacity planning and expansion of GPU cluster infrastructure.• Use Python, Shell, Ansible and/or Terraform to automate network configuration, health checks, monitoring and repetitive operational tasks.• Participate in infrastructure incident response, troubleshooting, post-incident reviews and Root Cause Analysis (RCA).• Maintain network diagrams, configuration documentation, SOPs, troubleshooting guides and operational records.• Coordinate with NVIDIA/Mellanox, server OEMs, network vendors and other technical support teams to resolve hardware and network issues.• Support hardware replacement and RMA activities for switches, NICs/HCAs, optical transceivers and other network components. Qualifications• Degree or relevant educational background in Computer Science, Information Technology, Networking, Telecommunications, Engineering or a related field.• 3+ years of hands-on network operations, data center networking or infrastructure experience.• Hands-on experience operating Infini Band and/or RoCE networks in GPU, AI, HPC or large-scale data center environments.• Experience with NVIDIA/Mellanox Infini Band or high-speed Ethernet networking equipment is strongly preferred.• Understanding of Infini Band architecture, including Subnet Management, partitions, routing, link states and performance counters.• Strong understanding of TCP/IP and data center networking.• Working knowledge of BGP, OSPF, VLAN and MLAG.• Ability to troubleshoot high-speed network connectivity, link degradation, congestion, packet loss and performance issues.• Hands-on ability to troubleshoot switches, NICs/HCAs, optical transceivers, cables and switch ports.• Familiarity with Infini Band diagnostic tools such as ibstat and ibdiagnet.• Familiarity with Linux and command-line troubleshooting.• Experience with monitoring platforms and protocols such as Prometheus, SNMP and Grafana.• Ability to write basic Python and/or Shell scripts for troubleshooting and automation.• Experience with Ansible and/or Terraform is preferred.• Good understanding of GPU cluster architecture and the relationship between GPU servers, NICs/HCAs and high-performance network fabrics.• Strong troubleshooting, documentation and Root Cause Analysis (RCA) capabilities.• Strong sense of ownership and ability to work effectively during production incidents. Preferred Qualifications• Experience supporting large-scale 1,000+ GPU clusters; experience with 10,000+ GPU environments is a strong plus.• Experience supporting NVIDIA H100/H200/B200 or AMD MI300X/MI355X GPU infrastructure.• Experience operating 100G/200G/400G or higher-speed Infini Band or Ethernet networks.• Experience troubleshooting NCCL and GPU communication issues from the network/fabric perspective.• Experience with NVIDIA/Mellanox network management and diagnostic tools.• Experience with network automation using Python, Shell, Ansible or Terraform.• CCNP, CCIE or other relevant networking certifications are a plus.• Experience working within ITIL-based incident, problem and change-management processes.
Matching similar jobs
JOB OVERVIEW
Experience level
Lead
Location
Buffalo, NY
Occupation
Network and Computer Systems Administrators
Industry
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Posted
5 days ago
Tired of running searches?
Rank the roles you'd take once, and matches like these arrive on their own.
CREATE PROFILE