Back to jobs

SENIOR SITE RELIABILITY ENGINEER

Lodgify
Full-time Contrato Indefinidosenior
Sign in to applyFree account, takes a minute.

Job description

⭐ Who we are   Lodgify is a fast-growing scale-up company leading the vacation rental industry. Backed by $30M in funding, our platform empowers property owners and managers worldwide to efficiently manage and grow their business through technology.   Headquartered in sunny Barcelona, we're now a team of 380+ people representing over 60 nationalities, united by a passion for transforming the future of short-term rentals. ⭐ How will you make an impact? • Define meaningful SLIs, SLOs, and reliability targets for the platform. • Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability. • Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness. • Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure, including how systems scale during growth, traffic spikes, and dependency failures. • Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana. • Implement operational and security best practices through guidelines, policies and automation. • Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly. • Automate repetitive operational work using Python or other languages, turning recurring manual work into safer automation and clearer runbooks. • Implement self-service Internal Developer Platform features via APIs and Kubernetes operators. • Improve deployment safety, rollbackability, and release observability. • Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms. • Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements. • Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability. ⭐ What makes you a great fit? • You have 7+ years of production experience operating Kubernetes-based platforms and cloud infrastructure. • You understand and apply SRE practices: SLIs, SLOs, error budgets, production readiness, incident response, post-incident learning, toil reduction, scalability, capacity planning, high availability, backups, and disaster recovery. • You can design and improve observability and alerting for critical systems using metrics, logs, traces, and golden signals, and are comfortable troubleshooting complex distributed systems to identify systemic reliability improvements. • You can write maintainable software to automate operational tasks and reduce manual intervention. • You have experience with stateful production systems such as relational databases, caches, queues, or streaming platforms. • You know how to balance reliability, performance, cost, and delivery speed pragmatically. • You are comfortable working in a transitional environment where SRE practices are being introduced while critical infrastructure and delivery systems still need hands-on reliability support. • You collaborate effectively with Engineering, Platform, Security, and Product stakeholders. • You communicate clearly, document well, and enjoy coaching teams toward stronger production ownership. • You model initiative and accountability, raising risks early and driving improvements through to completion. ⭐ What does success look like? • Critical services have clear owners, meaningful SLIs/SLOs, actionable alerts, dashboards, runbooks, and production readiness coverage. • Reliability targets are consistently met across critical infrastructure and services. • Operational toil and manual intervention are measurably reduced through automation and safer workflows. • MTTR improves through reduced alert noise, better signal quality, stronger observability, and clear incident response playbooks and escalation paths. • Post-incident actions are tracked, completed, and used to reduce repeat incidents. • Disaster recovery exercises validate that critical services and infrastructure can recover within agreed expectations. • Cloud and infrastructure resources are optimised without sacrificing performance, elasticity, or resilience. Why you’ll love us: You’ll be part of a growing, dynamic company with a truly international team. At Lodgify, we are full of contagious energy, hard work, and passion for what we do. We celebrate diversity and are proud to acknowledge a variety of backgrounds, perspectives and skills in our team; committed to creating a workplace where everyone is heard and feels a sense of belonging.   What's in it for you?* 🏠 Remote Flexibility: The freedom to work from home any day that works for you. 🌴 Time to Recharge: 25 working days of paid vacation and Jornada Intensiva in August.. 💊 Alan Health Insurance: Premium health, dental, and mental health support via Alan. Pre-existing conditions are covered. 😋 Meal Perk:€150/month allowance on your Alan card + 50% off Ametller Origen prepared dishes at the office. 💸 Tax-Free Savings: Increase your take-home pay by using Flexible Remuneration for extra meal costs (up to €70/mo) and public transport (up to €136/mo). 🖥️ Home Office Gear: We provide a table, ergonomic chair, and monitor for your home setup. 🇪🇸 Language Learning: Free Spanish classes. 🤑 Referrals: Cash rewards for bringing in new talent.  🌟 Social Life: Daily office breakfast and monthly team events 🎯 Dynamic Hub: A high-energy, inclusive environment designed for collaboration and connection with a team that represents over 60 countries. *Benefits offered may differ based on the type of contract that is issued   So, what are you waiting for? Apply now! All applications and CVs must be submitted in English 😉

Skills

KubernetesCloud InfrastructureDatadogPrometheusGrafanaPythonRelational DatabasesCachesQueuesStreaming PlatformsAPIsKubernetes Operators