Back to jobs

SENIOR SITE RELIABILITY ENGINEER (SRE/DEVOPS)

Qode
Full-timesenior

Job description

As a Senior DevOps Engineer , you will work closely with Product, Engineering, and AI teams to shape our infrastructure strategy, design resilient cloud architectures, and ensure our platforms are secure, scalable, and high-performing. You will play a key role in bringing AI systems into production, enabling reliable delivery, strong observability, and operational excellence across our products and internal systems. Your Responsibilities • Design and operate secure, scalable, and high-quality infrastructure that supports modern applications and advanced AI workloads • Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort • Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery • Apply strong systems engineering practices, including monitoring, incident management, performance optimization, and capacity planning • Establish and uphold DevOps best practices, ensuring reproducibility, testing, documentation, and operational excellence • Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery and effective problem-solving • Provide mentorship and technical leadership, raising the level of platform engineering, DevOps maturity, and overall engineering quality across the organization Requirements • Experience & Discipline : 6+ years of progressive experience in DevOps, Site Reliability Engineering (SRE), Platform Engineering, or Infrastructure Engineering. • Cloud Expertise : Strong, hands-on experience across multi-cloud environments (AWS, GCP, Azure), including expertise in networking, compute, storage, security, and cost optimization. • Core Platform Stack : Deep expertise in containerization and orchestration and extensive experience with Infrastructure as Code (IaC) (e.g., Terraform, Pulumi, CloudFormation). • AI/ML Infrastructure : Experience supporting or deploying AI/ML workloads (e.g., model inference, vector databases, GPU workloads), or strong familiarity with the infrastructure requirements for these systems. • System Reliability : Proven ability to design, build, and operate highly reliable, scalable production systems utilizing advanced Zero-Downtime Deployment Patterns (e.g., Blue/Green, Canary, progressive delivery, Preview Environments). • Modern Delivery & Tooling : Expertise in modernizing deployments via GitOps practices (e.g., ArgoCD, Flux) and building Self-Service Developer Platforms that enable engineering efficiency (e.g., environment automation, internal tooling). • Networking & Edge Routing : Experience implementing and managing Multi-Cloud API Gateways and Edge Routing solutions. • Security & Hardening : Strong background in platform security, including secrets management, Identity and Access Control (IAM), and Runtime/Security Hardening • Observability : Solid understanding and practical experience with modern observability stacks. • Mentorship & Communication : Excellent communication and collaboration skills with a proven ability to describe complex infrastructure decisions clearly and a background in mentoring engineers and driving improvements in engineering practices. • Development Expertise : Familiarity with modern programming languages like Node.js, NestJS, and Python is highly desirable for extending DevOps capabilities or integrating tooling.