Τι θα κάνεις
- Ηγείσαι και αναπτύσσεις την ομάδα SRE, παρέχοντας καθημερινή διαχείριση ανθρώπων, καθοδήγηση και ανάπτυξη καριέρας
- Διαχειρίζεσαι τη στρατηγική παρατηρησιμότητας στην πλατφόρμα, περιλαμβάνοντας μετρικές, καταγραφή, εντοπισμό ιχνών και ειδοποιήσεις
- Διαχειρίζεσαι την αξιολόγηση, τον προϋπολογισμό και τις σχέσεις με προμηθευτές για εργαλεία παρατηρησιμότητας
- Εξερευνάς και υιοθετείς τεχνολογίες AI για παρατηρησιμότητα, όπως ανίχνευση ανωμαλιών και AI-driven ανάλυση αιτίας
- Διαχειρίζεσαι την πρακτική επιδόσεων, συμπεριλαμβανομένων ελέγχων φορτίου και σχεδιασμού δυναμικότητας
- Ορίζεις και προωθείς SLIs, SLOs και error budgets για κρίσιμες υπηρεσίες
- Λειτουργείς ως σημείο κλιμάκωσης για σημαντικά περιστατικά, προσφέροντας ανάλυση αιτίας και βελτιώσεις
- Συνεργάζεσαι με ομάδες DevOps, Infrastructure και Database για να γεφυρώσεις κενά παρατηρησιμότητας σε υβριδικά περιβάλλοντα
- Διασφαλίζεις την ύπαρξη σαφών οδηγιών για κοινά σενάρια αποτυχίας
- Προωθείς μια πολιτιστική αλλαγή προς την πρόληψη μέσω καλύτερης αυτοματοποίησης, εργαλείων και προτύπων
Βασικές προϋποθέσεις
- Ηγεσία και διαχείριση ομάδων Site Reliability Engineering και Performance Engineering
- Εμπειρία σε παρατηρησιμότητα: μετρικές, logging, tracing και alerting
- Ικανότητες στον performance engineering: load testing, capacity planning, benchmarking
- Καθορισμός και λειτουργικοποίηση SLIs, SLOs και error budgets
- Αντιμετώπιση περιστατικών και ανάλυση αιτίων
- Εμπειρία με cloud περιβάλλοντα, AWS, Kubernetes και Rancher
- Γνώση εργαλείων AI για παρατηρησιμότητα και χρήση AI για ανίχνευση ανωμαλιών και root-cause analysis
- Ισχυρές επικοινωνιακές δεξιότητες και ικανότητα μετάφρασης τεχνικών δεδομένων σε μη τεχνικούς
- Διαχείριση ανθρώπων και καλλιέργεια ομαδικού πνεύματος
Παροχές
- Ελκυστικά οφέλη, ανοιχτό και υποστηρικτικό περιβάλλον, σύγχρονος και ελκυστικός χώρος εργασίας.
- Ευκαιρίες αλληλεπίδρασης με παγκόσμιες ομάδες και ουσιαστική ανάπτυξη καριέρας.
- Ευέλικτη εργασία με καθοδήγηση και ανάπτυξη δεξιοτήτων.
- Συνεργατικό περιβάλλον γραφείου με 2 ημέρες ανά εβδομάδα παρουσίας στο χώρο.
Περιγραφή Θέσης
We're looking for an SRE Manager to lead our Site Reliability Engineering team, with a particular focus on observability and performance engineering. You'll build and lead the team responsible for how we see, measure, and understand the health of our platform — setting the standards for monitoring, alerting, and performance testing that keep our systems running reliably. You'll balance hands-on technical leadership with people management, working closely with DevOps, Infrastructure, and Database Engineering to raise the bar on reliability across the estate.
What you’ll be doing
Lead and grow the SRE team, providing day-to-day people management, coaching, and career development for a group of reliability and performance engineers
Own the observability strategy across the platform — metrics, logging, tracing, and alerting — equipping DevOps, Infrastructure, and Engineering teams with self-service dashboards and insight rather than gatekeeping the data
Own the evaluation, budget, and vendor relationship for observability tooling, balancing capability, cost, and operational fit
Explore and adopt AI-assisted observability capabilities — anomaly detection, predictive alerting, and AI-driven root-cause analysis — to help the team spot issues earlier and resolve them faster
Own the performance engineering practice, including load testing, capacity planning, and performance benchmarking
Define and drive SLIs, SLOs, and error budgets for critical platform services, working with Engineering to embed reliability targets into delivery
Act as an escalation point for major incidents, contributing root cause analysis and ensuring learnings feed back into monitoring, alerting, and performance improvements
Partner with DevOps, Infrastructure, and Database Engineering to close observability gaps across the hybrid on-premises and cloud estate
Ensure the team maintains clear runbooks and playbooks for common failure scenarios, making reliability knowledge repeatable rather than dependent on individual expertise
Champion a proactive reliability culture, shifting the team from reactive firefighting toward prevention through better tooling, automation, and standards
The Player
Proven experience leading a Site Reliability Engineering or Performance Engineering team, including direct people management responsibility
Deep hands-on background in observability — metrics, logging, tracing, and alerting
Strong experience with performance engineering practices, including load testing, capacity planning, and performance benchmarking
A track record of defining and operationalising SLIs, SLOs, and error budgets in a production environment
Experience acting as an escalation point for major incident response, contributing to root cause analysis through to concrete reliability improvements
Comfort working across hybrid on-premises and cloud environments, ideally within a high-availability, transaction-heavy domain
Hands-on AWS experience, with a focus on cost optimization and leveraging AI-powered tools to drive efficiency across cloud infrastructure
Working knowledge of Kubernetes and Rancher, enough to operate confidently across our hybrid on-premises and cloud platforms
An AI-native approach to observability, with exposure to AI or LLM-assisted tools for anomaly detection and root-cause analysis a plus
Strong communication skills, with the ability to translate technical reliability data into insight for non-technical stakeholders
A coaching mindset, with genuine enthusiasm for developing engineers and building a strong team culture
What’s the Score?
Why OpenBet?
The Playground: Join a team of innovators, disruptors, and game-changers who are reshaping the future of betting and gaming.
The Mission: Be part of a mission-driven organization that's committed to revolutionizing the way the world plays.
The Impact: Make a real impact on the world stage, leaving a lasting legacy that transcends boundaries and inspires generations to come.
The Culture: Immerse yourself in a culture of creativity, collaboration, and curiosity, where every idea is welcomed, every voice is heard, and every dream is encouraged.
The Future: Join us on the journey to build the future of betting and gaming, one game-changing innovation at a time.
What we can offer YOU:
Attractive benefits, an open and supportive environment as well as a modern and exciting workplace
The opportunity to interact with global teams on a regular basis as you and our business continues to develop & grow
Tangible and genuine development - at OpenBet, you can take your career where you want it to go!
And if that’s not enough; enjoy flexible working whilst we provide you with the guidance and development skills you need to progress and enhance your career
We have a collaborative office environment with our team members in office 2 days per week.



