Position:
Observability Engineer
Company:
Weekday's Client
Compensation:
Rs 500000 - Rs 1700000 (INR 5 - 17 LPA)
Location:
India (Bengaluru, Chennai, Hyderabad, Mumbai, Pune)
Employment type:
Full-time
Work Arrangement:
On-site
Short Summary:
We are looking for a highly skilled Senior Observability Engineer to design, implement, and scale enterprise-grade observability platforms that provide deep visibility into distributed systems. This role is ideal for professionals passionate about improving system reliability, performance, and operational excellence through modern observability practices.
Responsibilities:
- Design, develop, and manage scalable end-to-end observability solutions covering metrics, logs, traces, and alerting across enterprise environments.
- Lead the migration from legacy monitoring platforms to modern observability frameworks and cloud-native monitoring solutions.
- Deploy, administer, and optimize observability platforms running on Kubernetes or OpenShift environments.
- Build reusable dashboards, alerts, and monitoring standards to improve operational visibility and incident response.
- Develop and maintain Helm charts for deployment and lifecycle management of observability components.
- Implement automation for deployment, configuration management, and operational workflows using Python or Bash scripting.
- Collaborate with engineering and application teams to define observability standards and integrate monitoring into development workflows.
- Analyze system performance, identify bottlenecks, and recommend improvements that enhance platform reliability and scalability.
- Provide technical leadership, architectural guidance, and strategic recommendations for observability initiatives.
- Support production operations by troubleshooting complex monitoring and infrastructure issues.
- Contribute to continuous improvement initiatives and drive adoption of observability best practices across engineering teams.
Requirement:
- Strong hands-on experience with OpenTelemetry for instrumentation and telemetry collection.
- Expertise in the Grafana Enterprise Stack, including Mimir, Loki, and Tempo.
- Experience administering and scaling ITRS Geneos in enterprise environments.
- Strong knowledge of Prometheus and PromQL.
- Hands-on experience with Grafana, including dashboard creation, alerting, and data source management.
- Experience administering OpenShift or Kubernetes clusters.
- Expertise in developing and managing Helm Charts for Kubernetes deployments.
- Experience designing, deploying, and scaling enterprise observability platforms.
Benefits:
- Not specified.