Position:
Site Reliability Engineer
Company:
Kong Inc.
Location:
Italy, Milan
Employment type:
Full time
Work Arrangement:
Hybrid
Short Summary:
The Site Reliability Engineering team at Kong is responsible for architecting and operating large-scale infrastructure, ensuring world-class reliability and performance for critical applications.
Responsibilities:
- Build and maintain core infrastructure as code using tools like Terraform and Ansible.
- Implement robust monitoring, logging, and alerting systems to ensure services meet and exceed 99.99% uptime.
- Resolve production incidents through systematic debugging and drive the blameless post-mortem process.
- Write automation to reduce operational toil and improve system efficiency.
- Collaborate with developers to embed reliability and scalability best practices.
- Contribute to capacity planning, disaster recovery drills, and security hardening processes.
- Participate in a fair and sustainable on-call rotation.
Requirement:
- Experience operating production workloads on a major cloud provider (AWS, GCP, Azure).
- Proficiency in at least one programming or scripting language (Golang, Python, or Bash).
- Hands-on experience with containerization and orchestration technologies (Docker, Kubernetes).
- Knowledge of Infrastructure as Code principles and tools (Terraform is a plus).
- Familiarity with CI/CD concepts and pipeline tools (e.g., GitLab CI, Jenkins).
- Understanding of modern observability stacks (e.g., Prometheus, Grafana, ELK).
Benefits:
- Join a leading developer of API and AI connectivity technologies.
- Work with trusted organizations, including Fortune 500 companies.
- Contribute to building infrastructure that powers the agentic era.