Senior Site Reliability Engineer
Our mission at Duolingo is to develop the best education in the world and make it universally available. It’s a big mission, and that’s where you come in!
At Duolingo, you’ll join a team that cares about finding innovative solutions to complex technical problems, running countless experiments (300+ at a time!) with our massive user base to make data-driven decisions, and educating our users and employees alike. You’ll have limitless learning opportunities, mentorship and collaboration with world-class minds, and a variety of projects with large scopes — while doing work that’s both fun and meaningful.
Join our life-changing mission to develop education for our half a billion (and growing!) learners around the world.
About the role...
As a Senior Site Reliability Engineer, you will work closely with both product and platform engineering teams to ensure Duolingo’s sophisticated distributed systems and products are built and maintained with extraordinary quality, and operated in measurable and scalable ways.
🧠 You will...
- Collaborate with internal teams to identify sources of instability in distributed systems and drive operational excellence
- Support core infrastructure (i.e understand, diagnose, and debug these systems in production)
- Provide system design consulting, develop software platforms/frameworks, and conduct launch reviews and root cause analysis
- Maintain and document sustainable postmortem/incident response practices
- Advocate for and implement changes that improve reliability, scalability, and velocity
- Reduce the burden of toil with iterative development of tooling and automation
- Collaborate with engineering teams to release new features and become an authority on our services
✅ You have...
- 5+ years of experience within site reliability engineering/DevOps of a product with millions of users
- Experience identifying and solving issues in large-scale distributed systems
- Experience with Java, Kotlin, Python or Go
- An understanding of containerization toolsets and container orchestration technologies (Docker, Mesos, Kubernetes, Nomad, etc)
⭐ Exceptional candidates will have...
- Experience in improving automation and tooling to reduce service maintenance toil
- Proven experience driving improvements to incident response processes
- Experience assessing reliability and troubleshooting issues in Dynamo, MySQL, and/or PostgreSQL databases
The offered salary is dependent upon several factors, including work experience, skills, and internal peer comparisons. The posted range is subject to change in the future. For this role, base salary is supplemented by equity compensation. We encourage you to talk with your recruiter for more information related to compensation for this role!
Salary Range: $182,800 — $247,300 USDAI in Hiring: Duolingo uses AI technology to assist in our hiring process. We want to be transparent about how we use this technology and your rights. We use Greenhouse Software's Talent Matching feature to assist with reviewing applications for the skills and experience we have decided are important for this role.
How it works:
- Talent Matching reviews your resume and responses to application questions to determine the extent to which your qualifications match the job requirements selected for this role. These requirements and their desired weighting are established by our recruiters in collaboration with hiring managers.
- All Talent Matching output is reviewed by our recruiting and hiring teams, who make the final decision on each application. Talent Matching does not automatically disposition any candidate or make hiring decisions.
Required Skills
Required Languages
🇬🇧 English