Back to jobs

SENIOR SITE RELIABILITY ENGINEER (SRE/DEVOPS)

Qode
Full-timesenior

Job description

Senior Site Reliability Engineer (SRE/DevOps) Location: Vietnam Workplace Type: Remote About the Role We are hiring on behalf of our client, a growing technology company, for a Senior Site Reliability Engineer to help shape their infrastructure strategy, design resilient cloud architectures, and ensure their platforms are secure, scalable, and high-performing. This role is critical to bringing AI systems into production with reliable delivery, strong observability, and operational excellence. Our client is specifically looking for someone with a genuine SRE mindset - not a traditional DevOps background. Candidates should have hands-on experience optimizing systems for reliability and building zero-downtime production systems , not just automating deployments. Responsibilities • Design and operate secure, scalable, high-quality infrastructure supporting modern applications and AI workloads. • Build and maintain robust automation across CI/CD pipelines, infrastructure provisioning, and operational processes to improve reliability and minimize manual effort. • Integrate AI-driven solutions into operational workflows to enhance efficiency, detect anomalies, and accelerate delivery. • Apply strong systems engineering practices: monitoring, incident management, performance optimization, and capacity planning. • Establish and uphold SRE/DevOps best practices, including reproducibility, testing, documentation, and operational excellence. • Communicate technical decisions clearly and collaborate cross-functionally to support predictable delivery. • Provide mentorship and technical leadership, raising the bar on platform engineering and DevOps maturity across the organization. Requirements Must-Have • 6+ years of progressive experience in Site Reliability Engineering, Platform Engineering, or Infrastructure Engineering (SRE background strongly preferred over traditional DevOps). • Proven track record designing, building, and operating highly reliable, zero-downtime production systems, using patterns such as Blue/Green, Canary, progressive delivery, or Preview Environments. • Deep, hands-on expertise in Kubernetes (or equivalent container orchestration) running in production. • Strong experience with Infrastructure as Code (Terraform, Pulumi, or CloudFormation). • Solid experience with at least one major cloud provider (AWS, GCP, or Azure), including networking, compute, storage, and security. • Experience building CI/CD pipelines from the ground up (not just using pre-built templates). • Practical experience with modern observability stacks (e.g., Prometheus, Grafana, Datadog). • Some hands-on experience supporting or deploying AI/ML workloads (model inference, vector databases, or GPU workloads). • Strong background in platform security: secrets management, IAM, and runtime/security hardening. • Excellent communication skills, with a proven ability to explain complex infrastructure decisions and mentor other engineers. Nice-to-Have • Experience with GitOps practices (ArgoCD, Flux). • True multi-cloud experience across AWS/GCP/Azure. • Experience with Multi-Cloud API Gateways and Edge Routing. • Experience building Self-Service Developer Platforms. • Familiarity with Node.js, NestJS, or Python for extending DevOps tooling. • Experience collaborating with QA/IT/ISRM teams on vulnerability remediation and incident investigation. Benefits • Attractive salary range and open to negotiate for strong fits. • Hybrid/Remote-friendly culture. Work where you grow best! • Flexible hours, async teamwork. Focus time is respected. • Work equipment support. • Allowance for certification & skill development. • Year-end bonus & performance-based rewards. • 22 paid leaves from your 5th year. Take a full month off. • Career growth with personal coaching sessions. • Open, collaborative team culture. No micromanagement, only trust. • Tools & AI-powered workflows that make remote work easier.