Senior Reliability Engineer - FanDuel, Hybrid & Remote
About Betfair Romania Development:
Betfair Romania Development is the largest technology hub of Flutter Entertainment, with over 2,000 people powering the world’s leading sports betting and iGaming brands. Exciting, immersive and safe experiences are delivered to over 18 million customers worldwide, from our office in Cluj-Napoca. Driven by relentless innovation and commitment to excellence, we operate our own unbeatable portfolio of diverse proprietary brands such as FanDuel, PokerStars, SportsBet, Betfair, Paddy Power, or Sky Betting & Gaming.
Our Values:
The values we share at Betfair Romania Development define what makes us unique as a team. They empower us by giving meaning to our contributions, and they ensure that we consistently strive for excellence in everything we do. We are looking for passionate individuals who align with our values and are committed to making a difference.
Win together | Raise the bar | Got your back | Own it | Positive impact
About FanDuel:
FanDuel is a leading force in the sports-tech entertainment industry, redefining how fans engage with their favorite sports, teams, and leagues. As the premier gaming destination in North America, FanDuel operates across multiple verticals, including sports betting, daily fantasy sports, online gaming, advance-deposit wagering, and media.
Role Overview:
As a Senior Reliability Engineer / SRE, you’ll help improve how FanDuel’s services are built, operated, and continuously improved from a reliability perspective.
You’ll work closely with application and platform teams to understand critical customer journeys, identify reliability risks, establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs), improve operational readiness, and turn production learnings into measurable engineering improvements.
This is a hands-on engineering role. You’ll use production telemetry, incident data, service architecture, and reliability practices to understand how systems behave and where improvements are needed. You’ll help teams adopt SRE practices such as SLOs, error budgets, production readiness, incident learning, toil reduction, and operational automation, while ensuring application teams continue to own the reliability of their services.
You’ll also work closely with Observability, Performance and Resilience Engineering to connect telemetry, performance testing, Gamedays, chaos testing, and incident findings into a broader view of service reliability. Where recurring operational problems exist, you’ll help identify opportunities for automation and self-service rather than relying on repeated manual intervention.
Key Accountabilities & Responsibilities:
-
Partner with application and platform teams to understand service architecture, dependencies, critical customer journeys, and reliability risks.
-
Help teams define meaningful SLIs and SLOs that connect technical service behaviour with customer and business outcomes.
-
Support the adoption of error budgets and help teams understand how reliability performance should influence engineering priorities.
-
Assess services against reliability and production-readiness expectations, identifying gaps across observability, alerting, SLOs, runbooks, dependencies, capacity, and recovery.
-
Use reliability maturity assessments and scorecards to help teams understand their current position and prioritise improvements with the greatest impact.
-
Support incident response and investigation for complex production issues, using logs, metrics, traces, service dependencies, and deployment information to understand system behaviour.
-
Review incidents and recurring operational issues to identify patterns, systemic risks, and opportunities for longer-term engineering improvements.
-
Improve runbooks, alerting, escalation paths, and operational practices based on production learnings.
-
Identify repetitive operational work and help design automation or self-service capabilities to reduce engineering toil.
-
Build tooling and automation that improves reliability workflows, operational readiness, investigation, and remediation.
-
Partner with Observability Engineering to ensure services have the telemetry required to measure reliability and troubleshoot production issues effectively.
-
Partner with Performance and Resilience Engineering to understand findings from performance tests, capacity assessments, Gamedays, chaos experiments, and failover testing, and translate them into actionable reliability improvements.
-
Support peak-event readiness by reviewing service health, SLOs, dependencies, operational readiness, capacity risks, and outstanding reliability findings.
-
Help application teams improve the reliability of critical customer journeys through better monitoring, dependency understanding, failure handling, and recovery practices.
-
Contribute to Reliability Engineering standards, patterns, documentation, and golden paths that can be reused across engineering teams.
-
Support the integration of reliability capabilities into CI/CD and developer workflows, including SLO-as-Code, telemetry validation, production-readiness checks, and automated operational controls.
-
Use AI-assisted investigation and automation capabilities to accelerate troubleshooting, identify recurring patterns, and reduce manual operational effort.
-
Share knowledge and support other engineers in developing stronger SRE and production-engineering practices.
Skills, Capabilities & Experience Required:
-
Strong hands-on experience in Site Reliability Engineering, Reliability Engineering, Platform Engineering, DevOps, or Production Engineering.
-
Good understanding of core SRE principles, including SLIs, SLOs, error budgets, incident management, operational readiness, toil reduction, and automation.
-
Experience defining or working with SLIs and SLOs for production services.
-
Strong production troubleshooting skills, with experience investigating issues across applications, infrastructure, networks, databases, and service dependencies.
-
Good understanding of distributed systems concepts and common reliability patterns such as retries, timeouts, circuit breakers, graceful degradation, redundancy, backpressure, and failure isolation.
-
Experience working with observability platforms such as Datadog or equivalent, with a good understanding of logs, metrics, traces, APM, dashboards, monitors, and synthetic monitoring.
-
Experience participating in incident response, post-incident reviews, and the implementation of meaningful follow-up actions.
-
Hands-on experience with Kubernetes and cloud infrastructure, preferably AWS.
-
Experience with infrastructure-as-code tools such as Terraform and modern CI/CD environments.
-
Strong automation and software engineering skills, with proficiency in at least one modern programming language such as Go, Java, Python, or JavaScript.
-
Experience identifying repetitive operational activities and replacing them with automation or self-service where appropriate.
-
Working knowledge of performance engineering, capacity management, resilience testing, failover, or chaos engineering.
-
Ability to understand application architecture and identify reliability risks across service and infrastructure dependencies.
-
Strong analytical and problem-solving skills, with the ability to turn technical signals and production behaviour into actionable engineering improvements.
-
Good communication skills and the ability to explain reliability concepts and recommendations to engineers and engineering leadership.
-
Ability to work across multiple teams, balance competing priorities, and drive work through to measurable outcomes.
-
A mindset focused on automation, continuous improvement, knowledge sharing, and solving systemic problems rather than repeatedly addressing the same symptoms.
A Sneak Peek Into Our Tech Stack:
AWS, Kubernetes, Terraform, Helm, Ansible, Vault
Datadog, OpenTelemetry, PagerDuty
Buildkite, GitHub and infrastructure-as-code workflows
Bits AI SRE and AI-assisted investigation capabilities
Locust, AWS Resilience Hub, AWS Fault Injection Service
Benefits:
-
Hybrid & remote working options
-
€1,000 per year for self-development
-
Company share scheme
-
25 days of annual leave per year
-
20 days per year to work abroad
-
5 personal days/year
-
Flexible benefits: travel, sports, hobbies
-
Extended health, dental and travel insurances
-
Customized well-being programmes
-
Career growth sessions
-
Thousands of online courses through Udemy
-
A variety of engaging office events
Disclaimer:
We are an inclusive employer. By embracing diverse experiences and perspectives, we create a lasting, positive impact for our employees, customers, and the communities we’re part of. You don't have to meet all the requirements listed to apply for this role. If you need any adjustments to make this role work for you, let us know, and we’ll see how we can accommodate them.
We thank all applicants for their interest; however, only the candidates who best meet the job requirements will be contacted for an interview.
By submitting your application online, you agree that your details will be used to progress your application for employment. If your application is successful, your details will be used to administer your personnel record. If your application is unsuccessful, we will retain your details for a period no longer than three years, to consider you for prospective roles within the company.
Required Skills
Required Languages
🇬🇧 English