Senior Network Architect – AI Infrastructure & Data Center Networks
Santa Clara County, CA - USA
Job Summary
Senior Network Architect AI Infrastructure & Data Center Networks
Location: Santa Clara CA (Hybrid 4 Days Onsite)
Job Type: Contract
STAFFXPERT LLC is seeking a Senior Network Architect AI Infrastructure & Data Center Networks on behalf of our client in Santa Clara CA.
We are seeking a highly experienced and technically accomplished Senior Network Architect to lead the design architecture and evolution of large-scale AI/ML data center and backbone network infrastructure. This role is ideal for an expert in high-performance networking who thrives in hyperscale environments and has a strong background in cloud-scale infrastructure supporting AI workloads.
The ideal candidate will bring deep expertise in high-capacity WAN architectures EVPN/VXLAN fabrics network automation and low-latency networking technologies designed for AI training and GPU-intensive environments.
Key Responsibilities-
Design and architect large-scale AI/ML data center networks and high-capacity WAN infrastructure.
-
Lead deployment and optimization of EVPN/VXLAN fabrics supporting GPU clusters and AI training environments.
-
Drive initiatives focused on network scalability reliability performance and automation across enterprise-scale infrastructure.
-
Design and optimize low-latency high-throughput networks supporting RDMA/RoCE workloads.
-
Develop and implement network automation solutions using Python Ansible Terraform/OpenTofu and CI/CD methodologies.
-
Define network standards operational frameworks observability practices and reliability engineering processes.
-
Collaborate closely with infrastructure cloud systems and AI engineering teams on strategic architecture initiatives.
-
Lead troubleshooting performance tuning and optimization efforts for large-scale production networks.
-
Provide technical mentorship contribute to architecture reviews and support engineering best practices.
-
15 years of experience in Network Architecture Network Engineering or Network Reliability Engineering.
-
Deep expertise in:
-
BGP OSPF IS-IS MPLS
-
EVPN/VXLAN
-
Data Center Networking
-
WAN and Backbone Architecture
-
AI/ML Infrastructure Networking
-
Network Performance and Capacity Planning
-
-
Strong hands-on experience with multi-vendor networking environments including Juniper Arista and Cisco technologies.
-
Experience with Linux systems administration and infrastructure automation.
-
Strong scripting/programming skills in Python Go Bash or similar languages.
-
Hands-on experience with Infrastructure-as-Code tools such as Ansible Terraform/OpenTofu or Pulumi.
-
Proven ability to design and support highly available scalable cloud and data center networks.
-
Experience supporting AI training clusters GPU fabrics or HPC environments.
-
Knowledge of PTP RDMA RoCEv2 and low-latency networking technologies.
-
Experience with network observability platforms such as Kentik ThousandEyes Zabbix Nagios or equivalent monitoring solutions.
-
Exposure to AWS GCP and hybrid cloud networking architectures.
-
Experience leading architecture reviews and cross-functional infrastructure programs.
-
Experience operating within large-scale hyperscaler environments.
-
Participation in industry organizations or technical communities focused on networking and infrastructure.
-
Background supporting multi-terabit AI research or high-performance infrastructure environments.