SENIOR SITE RELIABILITY ENGINEER (SRE/DEVOPS)
Qode
Full-timesenior
Job description
As a Senior DevOps Engineer , you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems.
Your Responsibilities
• Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads
• Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort
• Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery
• Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning
• Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence
• Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving
• Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization
Requirements
• Experience & Discipline : 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering.
• Cloud Expertise : Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization.
• Core Platform Stack : Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation).
• AI/ML Infrastructure : Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems.
• System Reliability : Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments).
• Modern Delivery & Tooling : Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling).
• Networking & Edge Routing : Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions.
• Security & Hardening : Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening
• Observability : Solid understanding and practical experience with modern observability stacks.
• Mentorship & Communication : Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices.
• Development Expertise : Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.