Site Reliability Engineer
Site Reliability Engineer - SAP Cloud Ops
Software-Development Operations Apply now »We help the world run better
At SAP, we keep it simple: you bring your best to us, and we'll bring out the best in you. We're builders touching over 20 industries and 80% of global commerce, and we need your unique talents to help shape what's next. The work is challenging – but it matters. You'll find a place where you can be yourself, prioritize your wellbeing, and truly belong. What's in it for you? Constant learning, skill growth, great benefits, and a team that wants you to grow and succeed.
As an SRE (Site Reliability Engineer), you will be to part of a high-performance team which continuously improves the reliability of critical systems, working closely with development and operations teams. The SRE is responsible for monitoring, troubleshooting, and developing tolling and automation to optimize system performance and efficiency.
The ideal candidate will challenge the status quo. To be successful in this role, you will need to thrive in an agile environment where teams work together toward a common goal:
· Identify engineering defects in the existing code base and continuously improve the code quality.
· Collaborate with development teams to implement and deploy new features and enhancements, ensuring they meet reliability and performance standards.
· Performs code reviews and pair program with other engineers on the team.
· Define and implement efficient end-to-end provisioning of automation solutions.
· Build CI/CD pipeline configurations to orchestrate provisioning and deployment.
· Automate monitoring tools to monitor system health and reliability to support high uptime requirements.
· Automate repetitive tasks and processes to improve efficiency and reduce manual intervention.
· Collaborate with cross-functional teams to define and establish service level indicators (SLIs), service level objectives (SLOs) and key software engineering metrics.
· Automate infrastructure in AWS and in private data centers with CloudFormation, Terraform, Ansible, and AWS DevOps tools.
· Conduct post-incident analyses to identify root causes and implement preventive measures to avoid future incidents.
· Perform capacity planning and resource allocation to ensure optimal system performance and scalability.
· Stay up to date with industry best practices, new technologies, and emerging trends in site reliability engineering.
· Create and maintain documentation for system architecture, configuration, and troubleshooting procedures.
Requirements
· Full understanding of DevOps, SRE and agile software development roles and concepts.
· Ability to use one or more of these languages: Python, Typescript/Javascript, Golang, Java, or C#.
· Full understanding of Git (code version control) and software development best practices (GitOps).
· Strong knowledge of Microsoft and Linux/Unix systems and command line tools.
· Knowledge of IaC and Configuration Management, using technologies such as Cloud Formation, Terraform, Puppet and Ansible.
· AWS main resources (VPC, EC2, IAM, API Gateway, autoscaling, availability zones, Lambda...) and deploying and running systems at scale.
· Understanding of microservices architecture (concepts).
· Full understanding of observability best practices, and monitoring and logging tools such as Dynatrace, New Relic, Prometheus, Grafana, ELK stack, Splunk…
· Jenkins, AWS Code Deploy or similar CI/CD tools and pipelines.
· Prior experience with containerized deployment
AI Usage in the Recruitment Process
For information on the responsible use of AI in our recruitment process, please refer to our Guidelines for Ethical Usage of AI in the Recruiting Process.
Please note that any violation of these guidelines may result in disqualification from the hiring process.
Requisition ID: 455626 | Work Area: Software-Development Operations | Expected Travel: 0 - 10% | Career Status: Professional | Employment Type: Regular Full Time | Additional Locations: #LI-Hybrid
Required Skills
Required Languages
🇬🇧 English