Back to jobs

SENIOR SITE RELIABILITY ENGINEER

Ttms
senior

Job description

We are looking for a highly experienced Senior Site Reliability Engineer (SRE) to design and operate a highly automated, multi-cloud provisioning platform. You will drive reliability, scalability, and automation while making independent architectural decisions and shaping SRE best practices across the team. Your responsibilities: • Architect and improve the reliability, scalability, security, and performance of a multi-cloud provisioning platform. • Design and maintain production-grade cloud infrastructure, with a strong focus on AWS . • Manage, scale, and provision Kubernetes environments. • Design end-to-end automation and CI/CD pipelines to minimize manual operations and technical toil. • Define, monitor, and continuously improve SLIs and SLOs , including latency, throughput, error rates, and capacity. • Identify platform bottlenecks, single points of failure, and opportunities to simplify complex systems. • Lead incident response, troubleshooting, Root Cause Analysis (RCA) , and implementation of long-term preventive solutions. • Collaborate with Product, Technical Leads, and engineering teams to embed reliability and security throughout the SDLC. • Make independent architectural and technical decisions for production environments. • Promote engineering best practices and mentor other engineers. We are looking for you, if you have: • 10+ years of hands-on experience in Software Engineering, DevOps, SRE, Platform Engineering, or Systems Infrastructure. • 3+ years of advanced, hands-on AWS experience, including designing and troubleshooting production cloud environments. • 3+ years of production-level Kubernetes experience, including managing, scaling, and provisioning clusters. • Strong experience with Infrastructure as Code and infrastructure automation. • Experience designing and maintaining automated CI/CD pipelines. • Strong understanding of reliability engineering, including SLIs, SLOs, monitoring, capacity, and performance management. • Proven experience handling critical production incidents and conducting RCA. • Strong knowledge of cloud architecture, networking, security, and highly available systems. • Proven ability to work with a high degree of autonomy and make sound architectural decisions. • Strong problem-solving and troubleshooting skills. • Excellent communication skills and experience collaborating with technical leads and cross-functional engineering teams. • Ability to mentor engineers and promote engineering best practices. Nice to have: • Experience with Terraform. • Experience with Azure and/or Google Cloud in addition to AWS. • Strong scripting or programming skills in Python, Go, Bash, or similar languages. • Experience building self-healing or highly automated infrastructure platforms. • Experience with modern monitoring and observability solutions. • DevSecOps and cloud security experience. • Experience working with product-driven platform engineering teams. We offer: • Participation in interesting and demanding projects. • Flexible working hours. • A great, non-corporate atmosphere. • Possibility to work remote or hybrid (2 days per week from the office). • Opportunities for development and promotion. • Attractive package of benefits. We reserve the right to contact the selected candidates.

Skills

AWSKubernetesInfrastructure as CodeCI/CD pipelinesTerraformAzureGoogle CloudPythonGoBashMonitoringObservabilityDevSecOpsCloud SecuritySLIsSLOs