[8SN] SENIOR SITE RELIABILITY ENGINEER (SRE) – KUBERNETES
Softwaremind
Full-timesenior
Job description
<p><strong>About the Role</strong></p><p>This is a Senior SRE role supporting production reliability for a Kubernetes-based UI service / AI experience framework stack.</p><p>This is not general infrastructure, and it is not a front-end developer role. The strongest candidates will have production SRE experience across Kubernetes operations, observability, Node.js runtime troubleshooting, JVM / Java service troubleshooting, Splunk, and incident ownership.</p><p><strong>What You’ll Do</strong></p><ul><li>Support the deployment, operation, and reliability of production services running on Kubernetes.</li><li>Monitor service health and investigate production incidents across distributed applications.</li><li>Participate in on-call support, incident response, root cause analysis, postmortems, and reliability improvements.</li><li>Troubleshoot application runtime, networking, and service-to-service issues in collaboration with engineering teams.</li><li>Support CI/CD, GitOps-based deployments, observability, and production monitoring.</li><li>Work within a client-directed backlog and established priorities.</li></ul>
<p><strong>Required Qualifications</strong></p><ul><li>5+ years of experience in<strong> Site Reliability Engineering, DevOps, Platform Engineering, Production Engineering</strong>, or a closely related role, including strong recent hands-on experience supporting Kubernetes-based production services.</li><li>3+ years of hands-on production <strong>Kubernetes </strong>experience strongly preferred. Kubernetes production operations, including deployment, scaling, rollout / rollback, resource tuning, and service-to-service troubleshooting</li></ul><ul><li>Strong production incident response experience, including on-call, runbooks, postmortems, and paging hygiene</li><li><strong>Splunk </strong>experience for log aggregation, search, and production troubleshooting</li><li><strong>Prometheus and Grafana </strong>experience, specifically building alert rules and dashboards, not only using existing dashboards</li><li><strong>CI/CD</strong> and infrastructure-as-code for containerized deployments, including Helm and GitOps tools such as ArgoCD or Flux</li><li>Strong Linux and networking fundamentals, including DNS, load balancing, TCP / HTTP, HTTP/2, and Kubernetes networking</li><li><strong>Production troubleshooting experience across Node.js and JVM/Java services, with strong depth in at least one runtime environment.</strong> Experience may include Node.js heap snapshots, CPU profiling, event-loop and memory analysis, as well as JVM GC log analysis, thread dumps, JVM tuning, and Java service latency investigation.</li><li>Service-to-service authentication experience, including mTLS, certificate rotation, certificate format conversion, and JWT-based service authentication</li></ul>
<p><strong>Nice to Have</strong></p><ul><li>Web Components / Lit experience, to perform first-level debugging of UI-related issues</li><li>Server-side rendering or isomorphic runtime experience</li><li>Canary rollout / multi-version production operations</li><li>Distributed tracing and request-context correlation</li><li>KEDA or event-driven autoscaling</li><li>Experience with enterprise platform integration layers</li></ul><p><strong>What We Offer</strong></p><ul><li>Competitive salary and laptop</li><li>Professional development and training opportunities</li><li>Work with cutting-edge cloud and container technologies</li><li>Flexible work arrangements and collaborative team environment</li><li>Impact on organization-wide digital transformation initiatives</li></ul>
Skills
KubernetesSplunkPrometheusGrafanaCI/CDHelmArgoCDFluxNode.jsJVMJavamTLSJWTLinuxNetworkingDNSLoad BalancingTCPHTTPHTTP/2