What you'll do
- Lead and grow the SRE team with people management, coaching, and career development
- Own the observability strategy including metrics, logging, tracing, and alerting
- Provide self-service dashboards and insights for Engineering teams
- Manage evaluation, budget, and vendor relationships for observability tools
- Adopt AI-assisted observability tools for anomaly detection and root-cause analysis
- Lead performance engineering including load testing, capacity planning, and benchmarking
- Define and drive SLIs, SLOs, and error budgets for platform services
- Act as an escalation point for major incidents with root cause analysis
- Partner with DevOps, Infrastructure, and Database Engineering to improve observability
- Maintain clear runbooks and playbooks for failure scenarios
- Champion proactive reliability culture focused on prevention and automation
Key requirements
- Experience in People Management and leading SRE or Performance Engineering teams
- Expertise in Observability with metrics, logging, tracing, alerting
- Strong background in Performance Engineering practices
- Knowledge of SLI/SLO/Error Budgets operationalization
- Incident Response experience and root cause analysis
- Work with hybrid Cloud Environments including AWS
- Hands-on experience with Kubernetes and Rancher
- Familiarity with AI-assisted Observability Tools
- Strong Communication Skills
Benefits
- Attractive benefits, an open and supportive environment, modern and exciting workplace.
- Opportunities to interact with global teams and genuine career development.
- Flexible working with guidance and skills development.
- Collaborative office environment with 2 days per week on-site.
About the job
We're looking for an SRE Manager to lead our Site Reliability Engineering team, with a particular focus on observability and performance engineering. You'll build and lead the team responsible for how we see, measure, and understand the health of our platform — setting the standards for monitoring, alerting, and performance testing that keep our systems running reliably. You'll balance hands-on technical leadership with people management, working closely with DevOps, Infrastructure, and Database Engineering to raise the bar on reliability across the estate.
What you’ll be doing
Lead and grow the SRE team, providing day-to-day people management, coaching, and career development for a group of reliability and performance engineers
Own the observability strategy across the platform — metrics, logging, tracing, and alerting — equipping DevOps, Infrastructure, and Engineering teams with self-service dashboards and insight rather than gatekeeping the data
Own the evaluation, budget, and vendor relationship for observability tooling, balancing capability, cost, and operational fit
Explore and adopt AI-assisted observability capabilities — anomaly detection, predictive alerting, and AI-driven root-cause analysis — to help the team spot issues earlier and resolve them faster
Own the performance engineering practice, including load testing, capacity planning, and performance benchmarking
Define and drive SLIs, SLOs, and error budgets for critical platform services, working with Engineering to embed reliability targets into delivery
Act as an escalation point for major incidents, contributing root cause analysis and ensuring learnings feed back into monitoring, alerting, and performance improvements
Partner with DevOps, Infrastructure, and Database Engineering to close observability gaps across the hybrid on-premises and cloud estate
Ensure the team maintains clear runbooks and playbooks for common failure scenarios, making reliability knowledge repeatable rather than dependent on individual expertise
Champion a proactive reliability culture, shifting the team from reactive firefighting toward prevention through better tooling, automation, and standards
The Player
Proven experience leading a Site Reliability Engineering or Performance Engineering team, including direct people management responsibility
Deep hands-on background in observability — metrics, logging, tracing, and alerting
Strong experience with performance engineering practices, including load testing, capacity planning, and performance benchmarking
A track record of defining and operationalising SLIs, SLOs, and error budgets in a production environment
Experience acting as an escalation point for major incident response, contributing to root cause analysis through to concrete reliability improvements
Comfort working across hybrid on-premises and cloud environments, ideally within a high-availability, transaction-heavy domain
Hands-on AWS experience, with a focus on cost optimization and leveraging AI-powered tools to drive efficiency across cloud infrastructure
Working knowledge of Kubernetes and Rancher, enough to operate confidently across our hybrid on-premises and cloud platforms
An AI-native approach to observability, with exposure to AI or LLM-assisted tools for anomaly detection and root-cause analysis a plus
Strong communication skills, with the ability to translate technical reliability data into insight for non-technical stakeholders
A coaching mindset, with genuine enthusiasm for developing engineers and building a strong team culture
What’s the Score?
Why OpenBet?
The Playground: Join a team of innovators, disruptors, and game-changers who are reshaping the future of betting and gaming.
The Mission: Be part of a mission-driven organization that's committed to revolutionizing the way the world plays.
The Impact: Make a real impact on the world stage, leaving a lasting legacy that transcends boundaries and inspires generations to come.
The Culture: Immerse yourself in a culture of creativity, collaboration, and curiosity, where every idea is welcomed, every voice is heard, and every dream is encouraged.
The Future: Join us on the journey to build the future of betting and gaming, one game-changing innovation at a time.
What we can offer YOU:
Attractive benefits, an open and supportive environment as well as a modern and exciting workplace
The opportunity to interact with global teams on a regular basis as you and our business continues to develop & grow
Tangible and genuine development - at OpenBet, you can take your career where you want it to go!
And if that’s not enough; enjoy flexible working whilst we provide you with the guidance and development skills you need to progress and enhance your career
We have a collaborative office environment with our team members in office 2 days per week.



