Back to jobs

SENIOR HPC NETWORKING ENGINEER

Mirantis
Full-timesenior

Job description

<p>Location: US<br> Employment Type: Full-time</p><p><strong>Role Overview:</strong><br> We are seeking a highly skilled <strong>Senior HPC Networking Engineer</strong> to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure.&#xa0;</p><p><strong>Key Responsibilities:</strong></p><ul><li><p>Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.</p></li><li><p>Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.</p></li><li><p>Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.</p></li><li><p>Perform performance tuning, monitoring, and capacity planning for HPC networking systems.</p></li><li><p>Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).</p></li><li><p>Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.</p></li><li><p>Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.</p></li><li><p>Develop and maintain documentation for network architecture, configurations, and operational procedures.</p></li><li><p>Participate in on-call rotations and provide escalation support for critical incidents.</p></li><li><p>Lead or contribute to network upgrades, migrations, and new deployments.</p></li><li><p>You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving.</p></li></ul><p><strong>Build, operate, and scale next-generation GPU infrastructure:</strong></p><ul><li>You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world’s most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads.</li></ul><p><strong>Build the playbook, then grow into the platform:</strong></p><ul><li>Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.</li></ul> <p><strong>Required:</strong></p><ul><li><p>5+ years of experience in network engineering, with a focus on HPC or data center environments.</p></li><li><p>Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).</p></li><li><p>Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.</p></li><li><p>Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).</p></li><li><p>Experience with network performance analysis and troubleshooting tools.</p></li><li><p>Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).</p></li><li><p>Strong analytical and problem-solving skills.</p></li></ul><p><strong>Preferred:</strong></p><ul><li><p>Experience with large-scale HPC clusters or AI/ML infrastructure.</p></li><li><p>Knowledge of RDMA, MPI, and low-latency networking concepts.</p></li><li><p>Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.</p></li><li><p>Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).</p></li></ul><p><strong>Soft Skills:</strong></p><ul><li><p>Strong communication and collaboration skills.</p></li><li><p>Ability to work independently and handle complex technical challenges.</p></li><li><p>Detail-oriented with a proactive approach to problem-solving.</p></li></ul><p><strong>What We Offer:</strong></p><ul><li><p>Opportunity to work on cutting-edge HPC infrastructure.</p></li><li><p>Collaborative and innovative work environment.</p></li><li><p>Competitive salary and benefits package.</p></li></ul><p>&#xa0;</p> <p><strong>What does Mirantis offer you?</strong></p><ul><li>Work with an established Silicon Valley leader in the cloud infrastructure industry;</li><li>Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;</li><li>Be a part of cutting-edge, open-source innovation;</li><li>Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;</li><li>Professional development and training;</li><li>Attend conferences and working groups;</li><li>Company outings, happy hours, hackathons, and tech talks;</li><li>Receive a competitive compensation package with a strong benefits plan.</li></ul><div sr-tagline=""></div><p>We are a&#xa0;<a target="_blank" href="https://www.g2.com/reports/grid-report-for-container-management-spring-2022.embed?featured=mirantis-kubernetes-engine-formerly-docker-enterprise&amp;secure%5Bgated_consumer%5D=7ed17484-74e3-4ce8-8ad6-b48b395fbf56&amp;secure%5Btoken%5D=f4b909a5c1a2d1aa71dee93761486db5732a5b82abd47aa75f3353da41e3b92c&amp;utm_campaign=gate-817340" rel="noopener noreferrer">Leader for Container Management</a>&#xa0;in G2 (#2 after AWS)!</p>

Skills

InfiniBandMellanoxNVIDIAFortiGateFortiManagerFortiAnalyzerBGPOSPFVLANsQoSTCP/IPBashPythonAnsibleTerraform