SENIOR AI INFRASTRUCTURE & PLATFORM OPERATIONS ENGINEER (REMOTE IN THE US)
Mirantis
Full-timesenior
Job description
<p>Our organization is establishing an Americas-based AI Infrastructure & Platform Operations unit dedicated to the management of expansive AI ecosystems utilizing NVIDIA GPU acceleration, high-speed interconnects, Kubernetes, and bleeding-edge platform frameworks.</p><p>This team maintains the reliability, efficiency, and architectural integrity of vital AI service platforms across a global datacenter footprint. Positioned at the nexus of core infrastructure and network engineering, you will sustain the high-performance environments essential for contemporary AI application suites.</p><p>This position offers the chance to engage with pioneering AI hardware while driving the development of automated operational capabilities via the k0rdent AI platform.</p><p><strong>Responsibilities</strong></p><p><strong>Technical Operations & Service Reliability</strong></p><ul><li><p>Lead the investigation and resolution of complex infrastructure, networking, and platform-related incidents.</p></li><li><p>Act as a senior escalation point for operational teams during critical service-impacting events.</p></li><li><p>Support large-scale NVIDIA GPU infrastructure and high-performance networking environments.</p></li><li><p>Troubleshoot complex Linux, Kubernetes, networking, storage, and hardware-related issues.</p></li><li><p>Analyze platform performance, capacity, stability, and reliability trends to proactively identify risks.</p></li><li><p>Lead root cause analysis activities and drive long-term corrective actions.</p></li><li><p>Collaborate with engineering teams, hardware vendors, and datacenter personnel to resolve complex technical challenges.</p></li><li><p>Participate in major incident management and service restoration activities.</p></li></ul><p><strong>Platform Operations & Engineering</strong></p><ul><li><p>Provide technical leadership for Kubernetes platform operations and supporting infrastructure services.</p></li><li><p>Drive improvements in platform reliability, observability, monitoring, and operational processes.</p></li><li><p>Identify opportunities to automate repetitive operational activities and improve operational efficiency.</p></li><li><p>Contribute to operational readiness reviews, infrastructure changes, upgrades, and service introductions.</p></li><li><p>Support the adoption and operation of AI-powered infrastructure services and operational capabilities through k0rdent AI.</p></li><li><p>Evaluate emerging technologies and operational practices to improve service delivery and platform resilience.</p></li></ul><p><strong>Technical Leadership</strong></p><ul><li><p>Mentor and support AI Infrastructure & Platform Operations Engineers.</p></li><li><p>Share technical knowledge through documentation, training sessions, and operational reviews.</p></li><li><p>Develop and maintain operational standards, runbooks, troubleshooting guides, and best practices.</p></li><li><p>Help define operational processes, escalation paths, and service reliability standards.</p></li><li><p>Act as a trusted technical advisor during operational planning and service improvement initiatives.</p></li></ul>
<ul><li><p>7+ years of experience in infrastructure operations, platform operations, site reliability engineering, network operations, cloud operations, datacenter operations, or related technical roles.</p></li><li><p>Expert-level Linux administration and troubleshooting skills.</p></li><li><p>Strong networking expertise, including experience diagnosing complex performance, connectivity, and reliability issues.</p></li><li><p>Strong experience operating Kubernetes in production environments.</p></li><li><p>Experience supporting large-scale production infrastructure and distributed systems.</p></li><li><p>Proven experience leading technical investigations and managing complex incidents.</p></li><li><p>Experience performing root cause analysis and driving long-term operational improvements.</p></li><li><p>Strong understanding of observability, monitoring, and service reliability practices.</p></li><li><p>Excellent troubleshooting and analytical skills across multiple infrastructure domains.</p></li><li><p>Strong communication, collaboration, and stakeholder management skills.</p></li></ul><p><strong>Preferred Experience</strong></p><p><strong>Experience in one or more of the following areas is highly desirable:</strong></p><ul><li><p>NVIDIA GPU infrastructure and accelerated computing platforms.</p></li><li><p>InfiniBand networking and NVIDIA UFM.</p></li><li><p>AI infrastructure environments.</p></li><li><p>HPC environments.</p></li><li><p>Platform Engineering or Site Reliability Engineering (SRE).</p></li><li><p>Large-scale Kubernetes operations.</p></li><li><p>Infrastructure automation technologies and Infrastructure-as-Code practices.</p></li><li><p>Observability platforms such as Grafana, Prometheus, ELK, or OpenTelemetry.</p></li><li><p>Performance analysis and optimisation of distributed infrastructure platforms.</p></li><li><p>Technical leadership, mentoring, or team lead responsibilities.</p></li></ul><p><strong>Why Join Us?</strong></p><ul><li><p>Operate some of the most advanced AI infrastructure environments in production today.</p></li><li><p>Work with the latest NVIDIA GPU technologies, Kubernetes platforms, and high-performance networking environments.</p></li><li><p>Help define operational standards and reliability practices for next-generation AI infrastructure services.</p></li><li><p>Influence the adoption of AI-powered operational capabilities through k0rdent AI.</p></li><li><p>Work alongside highly skilled engineers solving complex infrastructure and platform challenges at scale.</p></li><li><p>Join a growing organisation investing heavily in AI infrastructure, platform services, and operational innovation.</p></li></ul>
<p><strong>What does Mirantis offer you?</strong></p><ul><li>Work with an established Silicon Valley leader in the cloud infrastructure industry;</li><li>Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;</li><li>Be a part of cutting-edge, open-source innovation;</li><li>Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;</li><li>Professional development and training;</li><li>Attend conferences and working groups;</li><li>Company outings, happy hours, hackathons, and tech talks;</li><li>Receive a competitive compensation package with a strong benefits plan.</li></ul><div sr-tagline=""></div><p>We are a <a target="_blank" href="https://www.g2.com/reports/grid-report-for-container-management-spring-2022.embed?featured=mirantis-kubernetes-engine-formerly-docker-enterprise&secure%5Bgated_consumer%5D=7ed17484-74e3-4ce8-8ad6-b48b395fbf56&secure%5Btoken%5D=f4b909a5c1a2d1aa71dee93761486db5732a5b82abd47aa75f3353da41e3b92c&utm_campaign=gate-817340" rel="noopener noreferrer">Leader for Container Management</a> in G2 (#2 after AWS)!</p>