Operations Engineer
Zoom · Technology
- Location: San Jose (United States of America)
- Work arrangement: hybrid
- Employment type: Full Time
- Salary: $98.9k - $228.7k USD/yr
- Seniority: middle
- Posted:
Job description
Position:
ZfG Operations Engineer
Company:
Zoom
Compensation:
Minimum: $98,900.00
Maximum: $228,700.00
Location:
United States, San Jose, California
Employment type:
Full time
Work Arrangement:
Hybrid
Short Summary:
You will develop and maintain reliable infrastructure solutions that power distributed systems serving millions of users globally. You will implement automation frameworks, monitoring tools, and performance optimization techniques while collaborating with cross-functional teams to resolve complex technical challenges.
Responsibilities:
- Designing and implementing automation frameworks and monitoring solutions that enhance system reliability, reduce manual interventions, and optimize performance across distributed production environments
- Analyzing system performance metrics to identify bottlenecks, recommend scalability improvements, and implement proactive solutions that prevent service degradation
- Leading incident response efforts by coordinating with cross-functional teams, conducting root cause analysis, and implementing preventive measures that reduce future outage risk
- Developing and maintaining operational documentation, runbooks, and service level objectives that standardize procedures and improve team efficiency
- Mentoring team members on troubleshooting methodologies, system optimization techniques, and infrastructure best practices while facilitating knowledge sharing across teams
Requirement:
- Demonstrate proficiency in at least one programming or scripting language (Python, Go, Bash, or similar) for building automation tools and infrastructure solutions
- Apply knowledge of distributed systems architecture, including scalability patterns, fault tolerance mechanisms, and performance optimization principles
- Utilize monitoring tools, observability platforms, and metrics collection systems to maintain visibility into production environments
- Execute incident management protocols, including response coordination, root cause analysis, and remediation planning
- Collaborate effectively with cross-functional teams to solve complex technical problems and drive infrastructure improvements
- Operate independently on large-scale projects with minimal supervision while maintaining accountability for deliverables and timelines
- Possess equivalent practical experience in site reliability engineering, DevOps, or infrastructure operations roles
- Contribute to on-call rotations and demonstrate experience supporting production systems in high-availability environments
Benefits:
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways.
ZfG Operations Engineer
Company:
Zoom
Compensation:
Minimum: $98,900.00
Maximum: $228,700.00
Location:
United States, San Jose, California
Employment type:
Full time
Work Arrangement:
Hybrid
Short Summary:
You will develop and maintain reliable infrastructure solutions that power distributed systems serving millions of users globally. You will implement automation frameworks, monitoring tools, and performance optimization techniques while collaborating with cross-functional teams to resolve complex technical challenges.
Responsibilities:
- Designing and implementing automation frameworks and monitoring solutions that enhance system reliability, reduce manual interventions, and optimize performance across distributed production environments
- Analyzing system performance metrics to identify bottlenecks, recommend scalability improvements, and implement proactive solutions that prevent service degradation
- Leading incident response efforts by coordinating with cross-functional teams, conducting root cause analysis, and implementing preventive measures that reduce future outage risk
- Developing and maintaining operational documentation, runbooks, and service level objectives that standardize procedures and improve team efficiency
- Mentoring team members on troubleshooting methodologies, system optimization techniques, and infrastructure best practices while facilitating knowledge sharing across teams
Requirement:
- Demonstrate proficiency in at least one programming or scripting language (Python, Go, Bash, or similar) for building automation tools and infrastructure solutions
- Apply knowledge of distributed systems architecture, including scalability patterns, fault tolerance mechanisms, and performance optimization principles
- Utilize monitoring tools, observability platforms, and metrics collection systems to maintain visibility into production environments
- Execute incident management protocols, including response coordination, root cause analysis, and remediation planning
- Collaborate effectively with cross-functional teams to solve complex technical problems and drive infrastructure improvements
- Operate independently on large-scale projects with minimal supervision while maintaining accountability for deliverables and timelines
- Possess equivalent practical experience in site reliability engineering, DevOps, or infrastructure operations roles
- Contribute to on-call rotations and demonstrate experience supporting production systems in high-availability environments
Benefits:
As part of our award-winning workplace culture and commitment to delivering happiness, our benefits program offers a variety of perks, benefits, and options to help employees maintain their physical, mental, emotional, and financial health; support work-life balance; and contribute to their community in meaningful ways.
Skills
- bash
- distributed_systems
- golang
- performance_optimization
- python
Languages
EN
Apply
Open this job in our interactive board to apply, save it, or sign up for matched alerts on similar roles.
View & apply