SENIOR SITE RELIABILITY ENGINEER
Ttms
senior
Job description
We are looking for a highly experienced Senior Site Reliability Engineer (SRE) to design and operate a highly automated, multi-cloud provisioning platform. You will drive reliability, scalability, and automation while making independent architectural decisions and shaping SRE best practices across the team.
Your responsibilities:
• Architect and improve the reliability, scalability, security, and performance of a multi-cloud provisioning platform.
• Design and maintain production-grade cloud infrastructure, with a strong focus on AWS .
• Manage, scale, and provision Kubernetes environments.
• Design end-to-end automation and CI/CD pipelines to minimize manual operations and technical toil.
• Define, monitor, and continuously improve SLIs and SLOs , including latency, throughput, error rates, and capacity.
• Identify platform bottlenecks, single points of failure, and opportunities to simplify complex systems.
• Lead incident response, troubleshooting, Root Cause Analysis (RCA) , and implementation of long-term preventive solutions.
• Collaborate with Product, Technical Leads, and engineering teams to embed reliability and security throughout the SDLC.
• Make independent architectural and technical decisions for production environments.
• Promote engineering best practices and mentor other engineers.
We are looking for you, if you have:
• 10+ years of hands-on experience in Software Engineering, DevOps, SRE, Platform Engineering, or Systems Infrastructure.
• 3+ years of advanced, hands-on AWS experience, including designing and troubleshooting production cloud environments.
• 3+ years of production-level Kubernetes experience, including managing, scaling, and provisioning clusters.
• Strong experience with Infrastructure as Code and infrastructure automation.
• Experience designing and maintaining automated CI/CD pipelines.
• Strong understanding of reliability engineering, including SLIs, SLOs, monitoring, capacity, and performance management.
• Proven experience handling critical production incidents and conducting RCA.
• Strong knowledge of cloud architecture, networking, security, and highly available systems.
• Proven ability to work with a high degree of autonomy and make sound architectural decisions.
• Strong problem-solving and troubleshooting skills.
• Excellent communication skills and experience collaborating with technical leads and cross-functional engineering teams.
• Ability to mentor engineers and promote engineering best practices.
Nice to have:
• Experience with Terraform.
• Experience with Azure and/or Google Cloud in addition to AWS.
• Strong scripting or programming skills in Python, Go, Bash, or similar languages.
• Experience building self-healing or highly automated infrastructure platforms.
• Experience with modern monitoring and observability solutions.
• DevSecOps and cloud security experience.
• Experience working with product-driven platform engineering teams.
We offer:
• Participation in interesting and demanding projects.
• Flexible working hours.
• A great, non-corporate atmosphere.
• Possibility to work remote or hybrid (2 days per week from the office).
• Opportunities for development and promotion.
• Attractive package of benefits.
We reserve the right to contact the selected candidates.
Skills
AWSKubernetesInfrastructure as CodeCI/CD pipelinesTerraformAzureGoogle CloudPythonGoBashMonitoringObservabilityDevSecOpsCloud SecuritySLIsSLOs