Position:
SRE Engineering Manager - GPU Cloud
Company:
Scaleway
Location:
France, Paris
Employment type:
Full-time (long term)
Work Arrangement:
Hybrid
Short Summary:
Join Scaleway and shape the sovereign cloud of tomorrow! Lead the Site Reliability Engineering (SRE) team to build, automate, and maintain a highly reliable GPU cluster infrastructure.
Responsibilities:
- Lead and manage a team of 6 Site Reliability Engineers, supporting their career growth and technical execution
- Design and implement automated solutions for server lifecycle management across GPU clusters
- Design and implement observability, logging, and monitoring solutions for large-scale GPU clusters
- Plan, prioritize, and manage the technical development roadmap for the SRE team
- Collaborate and coordinate closely with software engineering, product, and cross-functional teams
- Handle recruitment and career management for team members
- Maintain, scale, and optimize high-availability production systems under heavy load
- Participate in on-call rotations to ensure production reliability and fast incident resolution
Requirement:
- Strong experience managing engineering teams in high-constraint production environments
- Proven expertise with Kubernetes container orchestration
- Direct experience with cluster management and virtualization tools (Proxmox, Warewulf)
- Experience with monitoring, metrics, and observability stacks (Prometheus, Grafana)
- Exposure to modern GPU hardware ecosystems (Nvidia, AMD) and high-speed networking fabric (InfiniBand, Spectrum-X, Tomahawk)
- Knowledge of distributed and high-performance storage solutions (Lustre DDN, VAST)
- Strong engineering leadership and team management capabilities
- Technical rigor and high attention to detail in production-critical environments
- Ability to handle high-pressure operational situations and manage incident stress pragmatically
- Excellent communication skills with the ability to convey challenging messages effectively
- Collaborative mindset with a focus on empowering engineers rather than micromanaging
Benefits:
- Hybrid work: Up to 3 days of remote work per week
- Spacious, dynamic workspaces with outdoor spaces and bike parking
- Healthy meal service at headquarters and lunch support for regional sites
- Well-being commitments including gym access and daycare services
- International environment with a diverse workforce
- Opportunities for internal mobility within the Iliad Group