SENIOR SITE RELIABILITY ENGINEER
Jobgether
Full-timesenior
Job description
Accountabilities:
• Design, build, and scale Kubernetes-based infrastructure supporting secure, multi-tenant, and highly available applications.
• Develop and operate AI tooling infrastructure, including secure AI access patterns, MCP servers, and governance frameworks for production environments.
• Optimize CI/CD pipelines to improve deployment speed, reliability, automation, and rollback safety.
• Implement progressive delivery practices such as blue/green deployments and canary releases.
• Advance Infrastructure as Code practices using tools such as Terraform, Helm, and GitOps workflows to create reusable infrastructure patterns.
• Operate and improve streaming and analytics infrastructure, including Kafka, Flink, and ClickHouse environments.
• Establish and enhance observability practices through monitoring, SLOs, alerting systems, and operational dashboards.
• Lead incident response activities, perform root cause analysis, and drive long-term reliability improvements.
• Build automated testing practices into the software delivery lifecycle.
• Mentor engineers and promote best practices across Kubernetes, cloud infrastructure, automation, and reliability engineering.
Requirements:
• 6+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related roles with significant production Kubernetes experience.
• Hands-on experience integrating AI/LLM tools into engineering or operational workflows, including understanding security, governance, and access control considerations.
• Proven experience designing and maintaining CI/CD pipelines using tools such as GitHub Actions, Jenkins, GitLab CI, or similar technologies.
• Strong knowledge of Kubernetes internals and managed cloud Kubernetes services such as EKS, GKE, or AKS.
• Experience with Infrastructure as Code tools including Terraform, Helm, Pulumi, or equivalent solutions.
• Proficiency in scripting or programming languages such as Python, Bash, or Go.
• Experience with observability platforms such as Prometheus, Grafana, Datadog, or OpenTelemetry.
• Production experience working with distributed systems, streaming technologies, and analytics platforms such as Kafka, Flink, and ClickHouse.
• Strong understanding of cloud infrastructure, automation, system reliability, and operational excellence.
• Excellent communication and collaboration skills with the ability to work effectively across engineering teams.
• Experience with multi-region Kubernetes environments, chaos engineering, security automation, policy-as-code, or MLOps workflows is a plus.
Benefits:
• Competitive compensation package ranging from BRL 422,500 – BRL 485,000 total compensation (base salary plus bonus).
• Stock options and equity opportunities.
• Health benefits and country-specific employee support programs.
• Unlimited paid time off and flexible leave policies.
• Paid parental leave.
• Tuition reimbursement and learning and development opportunities.
• Flexible remote working environment.
• Additional employee benefits designed to support professional growth and well-being.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Why Apply Through Jobgether?
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1