Back to jobs

SITE RELIABILITY ENGINEER

ontrac-solutions-llc
Contractmid

Job description

Our client's Cloud Operations team is expanding its SRE function. Site Reliability Engineers keep all user-facing services and production systems running smoothly. SREs here are a blend of pragmatic operators and software craftspeople who apply sound engineering principles, operational discipline and mature automation to the environment and the codebase. The team specialises in systems — networking, the Linux kernel, and scaling, algorithms and distributed systems. As an SRE you will • Be on an on-call rotation responding to production availability incidents, and support service engineers with customer incidents • Use your on-call shift to prevent incidents from ever happening • Run infrastructure with Ansible, Puppet, Terraform and Kubernetes • Make monitoring and alerting alert on symptoms, not outages • Document every action, so findings turn into repeatable actions — and then into automation • Improve the deployment process to make it as boring as possible • Design, build and maintain core infrastructure that scales to hundreds of thousands of concurrent users • Debug production issues across services and levels of the stack • Plan the growth of the infrastructure You may be a fit if you • Think cloud-first, regardless of the flavour of public cloud • Think security-first • Think about systems — edge cases, failure modes, behaviours, specific implementations • Know your way around Linux and Windows • Know the use of config-management systems like Ansible or Puppet • Have strong programming skills — Python, Java, Golang, Node.js • Collaborate and communicate asynchronously, and document so nothing is learned twice • Have a go-for-it attitude: when you see something broken, you fix it • Have experience with Nginx, HAProxy, Docker, Kubernetes, Terraform or similar technologies Projects you could work on • Coding infrastructure automation with Ansible and Terraform • Improving Prometheus monitoring or building new metrics • Helping release managers deploy and fix new versions of application software • Planning and executing the migration from AWS virtual machines to cloud-native, container-based deployments on Kubernetes (EKS) • Developing a relationship with a product group and defining their SRE KPIs — the SRE practice here is early in its journey

Skills

AnsiblePuppetTerraformKubernetesPythonJavaGolangNode.jsNginxHAProxyDockerPrometheus