Site Reliability Engineer
The Site Reliability Engineer maintains and improves the availability, performance, resilience, and operational recoverability of enterprise identity, credential, and access-management services. This role supports cloud identity and directory platforms, including Microsoft Entra ID, Active Directory, and integrated authentication services, by applying site reliability engineering practices to measure service health, identify degradation, reduce operational toil, and strengthen reliability for mission-critical authentication and access services.
The engineer implements monitoring, alerting, dashboards, log-analysis capabilities, automation, and infrastructure-as-code solutions that support repeatable operations and recovery. Working with senior engineers, identity engineers, and operations teams, the Site Reliability Engineer supports incident response, root-cause analysis, preventive remediation, observability standards, and continuous service improvements across the enterprise identity environment.
Key Responsibilities
- Monitor service-level indicators, service-level objectives, service-level agreements, availability, latency, authentication success rates, provisioning outcomes, directory synchronization status, endpoint availability, and service dependency health for enterprise identity services.
- Implement and maintain observability solutions for Microsoft Entra ID, Active Directory, and integrated identity services using Azure Monitor, Log Analytics, Splunk, or comparable monitoring and security-event platforms.
- Collect, query, correlate, and analyze identity audit, sign-in, provisioning, directory, and infrastructure logs to investigate service failures, identify abnormal trends, validate remediation, and support compliance-oriented operational reporting.
- Support incident response for identity-service disruptions by assessing scope and business impact, executing documented runbooks, communicating status, engaging appropriate escalation teams, preserving diagnostic evidence, and validating service recovery.
- Perform root-cause analysis for recurring authentication, authorization, provisioning, directory synchronization, monitoring, capacity, and configuration failures; document corrective and preventive actions that improve long-term service reliability.
- Develop and support automated remediation workflows that detect known service conditions, validate safeguards, execute approved recovery actions, record results, and escalate when automation does not restore service health.
- Use infrastructure-as-code tools, including Terraform, Ansible, ARM templates, or comparable technologies, to define, version, review, deploy, and maintain repeatable cloud and identity-supporting infrastructure configurations.
- Apply Git-based version-control practices, including branching, pull requests, peer review, release tagging, change history, configuration rollback, and controlled promotion of infrastructure code across environments.
- Support Microsoft Entra ID, Active Directory, or comparable enterprise identity platforms by troubleshooting authentication flows, conditional access effects, application integrations, directory objects, service principals, synchronization dependencies, and access-policy impacts.
- Create and maintain operational runbooks, alert-response procedures, dashboards, recovery guides, post-incident records, and reliability-improvement backlogs that enable consistent support and continuous improvement.
Required Qualifications
- High school diploma or equivalent and at least 3 years of experience in identity, credential, and access-management engineering, infrastructure operations, cloud operations, DevOps, site reliability engineering, or related enterprise IT roles.
- At least 1 year of hands-on experience supporting Microsoft Entra ID, Active Directory, or another enterprise cloud identity or directory platform.
- Working knowledge of site reliability engineering concepts, including service-level indicators, service-level objectives, service-level agreements, alert thresholds, availability monitoring, latency measurement, on-call response, reliability reporting, and operational toil reduction.
- Experience implementing or supporting monitoring, alerting, log analysis, dashboards, or observability capabilities for cloud, directory, identity, or integrated authentication services.
- Experience using infrastructure-as-code tools, such as Terraform, Ansible, ARM templates, or comparable technologies, to define, deploy, and maintain repeatable infrastructure configurations.
- Experience using Git-based version control, including branching, pull requests, peer review, change history, tagging, rollback, and controlled promotion of code or configuration across environments.
- Demonstrated ability to support incident response, technical triage, root-cause analysis, corrective actions, preventive measures, and recovery validation for service disruptions.
- U.S. citizenship and ability to obtain and maintain a Public Trust background investigation.
Preferred Qualifications
- Microsoft certification in identity and access administration, cloud administration, security operations, or Azure infrastructure.
- HashiCorp Terraform Associate certification or demonstrated experience developing and operating production infrastructure-as-code pipelines.
- Experience using Splunk, Azure Monitor, Log Analytics, Microsoft Sentinel, or comparable observability and security-event platforms.
- Experience with automated orchestration or scripting using PowerShell, Python, Bash, or comparable technologies.
Compensation Ranges
Compensation ranges for ASM Research positions vary depending on multiple factors; including but not limited to, location, skill set, level of education, certifications, client requirements, contract-specific affordability, government clearance and investigation level, and years of experience. The compensation displayed for this role is a general guideline based on these factors and is unique to each role. Monetary compensation is one component of ASM's overall compensation and benefits package for employees.
EEO Requirements
It is the policy of ASM that an individual's race, color, religion, sex, disability, age, sexual orientation or national origin are not and will not be considered in any personnel or management decisions. We affirm our commitment to these fundamental policies.
All recruiting, hiring, training, and promoting for all job classifications is done without regard to race, color, religion, sex, disability, or age. All decisions on employment are made to abide by the principle of equal employment.
Physical Requirements
The physical requirements described in "Knowledge, Skills and Abilities" above are representative of those which must be met by an employee to successfully perform the primary functions of this job. (For example, "light office duties' or "lifting up to 50 pounds" or "some travel" required.) Reasonable accommodations may be made to enable individuals with qualifying disabilities, who are otherwise qualified, to perform the primary functions.
Disclaimer
The preceding job description has been designed to indicate the general nature and level of work performed by employees within this classification. It is not designed to contain or be interpreted as a comprehensive inventory of all duties, responsibilities and qualifications required of employees assigned to this job.
Required Skills
Required Languages
🇬🇧 English