SENIOR SITE RELIABILITY ENGINEER (PERFORMANCE AND SCALABILITY)
Digital Zone
Full-timeseniorTop-of-the-market compensation packages
Job description
Your mission is to make DigitalZone able to scale. You will build the platform's capacity to absorb campaign-level traffic spikes, and you will give every engineering team the tools, standards, and practices to load- and failure test their own systems. This is an enablement role at its core: you raise the reliability bar across the org by building capability, not by owning every service yourself.
What you'll do
• Build the platform's scalability foundation: capacity planning, autoscaling, caching, queueing, and graceful degradation designed for large campaign spikes rather than steady-state load.
• Establish load and failure testing as a standard engineering practice, giving teams the frameworks, tooling, and runbooks to test their own services and act on the results.
• Own SLOs, error budgets, and the observability stack (metrics, logs, traces, alerting) across TypeScript, Go, and PHP/Laravel services, and standardize how teams instrument for scale.
• Harden Postgres and AWS infrastructure for performance and availability, and reduce toil through automation and IaC.
• Lead incident response and blameless postmortems, and drive the systemic fixes upstream into design and campaign planning so reliability is built in, not bolted on.
• Partner with engineering teams early on capacity and resilience, acting as the multiplier that makes them self-sufficient at scaling their own systems.
What you'll bring
• 5+ years in SRE, platform, or backend engineering, with strong production ownership of large-scale systems operating at 10s of thousands of requests per minute.
• A track record of scaling systems through real traffic spikes, and of designing and running load and failure testing programs that other teams adopted.
• Deep AWS experience and a solid grasp of Postgres performance and scaling.
• Fluency with observability tooling and infrastructure-as-code, plus scripting in Go, TypeScript, or similar.
• A calm, systematic approach to incidents, and the communication skills to influence and enable other teams rather than gatekeep.
• Immediate, large-scale impact on a high-growth business
• Top-of-the-market compensation packages
• Work alongside top regional talent, with team members from Talabat, Careem, Etisalat, and more
Skills
AWSPostgresGoTypeScriptPHP/LaravelObservabilityInfrastructure-as-codeLoad testingFailure testingSLOsError budgetsMetricsLogsTracesAlertingAutoscalingCachingQueueingGraceful degradation