Position:
Lead Site Reliability Engineer
Company:
Luxoft
Location:
USA, Wilmington, Delaware
Short Summary:
Responsible at the expert level for ensuring the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. Serves as a senior individual contributor responsible for designing, implementing, and improving Site Reliability Engineering (SRE) practices across the software development lifecycle.
Responsibilities:
- Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure following enterprise technology standards and SRE best practices.
- Lead initiatives to improve system reliability, availability, performance, and operational maturity through automation and engineering excellence.
- Define, implement, and monitor Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets for critical business services.
- Develop comprehensive observability strategies leveraging Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics, logging, dashboards, and alerting solutions.
- Analyze production telemetry to proactively identify performance bottlenecks, reliability risks, and capacity constraints.
- Lead incident response activities for high-severity production events, coordinating cross-functional teams to restore services and minimize customer impact.
- Drive operational excellence through automation of repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls.
- Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational best practices.
Requirement:
- Strong experience in observability and monitoring, including hands-on expertise with Dynatrace, OpenTelemetry (OTel), distributed tracing, metrics collection, and analysis.
- Proven experience designing and executing automated regression testing frameworks and test suites to ensure application and platform stability following deployments.
- Strong proficiency in Infrastructure as Code (IaC) using Terraform.
- Expert knowledge of production systems monitoring, incident management, and operational troubleshooting.
- Strong understanding of application performance management, distributed systems, and modern cloud-native architectures.
- Demonstrated experience implementing and operating SRE practices, including SLOs, SLIs, error budgets, incident management, and RCA.
Benefits:
- Not specified.