Senior Storage Engineering Manager
At Verda, we're building a full-stack AI cloud, covering everything from data centers and hardware to our own cloud platform that the world's leading AI teams use to do serious AI work.
We strive to make a positive mark on the world through the infrastructure we build and give leading teams a service they can truly depend on. Headquartered in Helsinki, we operate globally with offices in London and San Francisco.
Join Verda while it’s still being built - not once it’s finished.
Why Verda
Cash and equity compensation along with local benefits.
40+ nationalities, with 6 different ones on the management team.
A real chance to make an impact and work alongside world class engineers, researchers, and partners across the global AI ecosystem.
Practicalities
Storage that holds real AI workloads at scale, in datacenters we own, with a new site opening roughly every month.
A senior team that is growing, and a leadership role in defining its structure and practices.
Hybrid work from our Helsinki/ London office or fully remote work within EU.
About the role
The team
The Storage team builds and operates Verda's storage platform. It runs Ceph across every Verda datacenter: block storage (RBD) behind every VM volume, CephFS and NFS behind our Shared File System, and RGW behind S3-compatible Object Storage, which is launching to customers with multi-site replication. The team operates shared cloud clusters and dedicated storage clusters for customers running large-scale ML training, brings up a new storage cluster with every datacenter we open, and provides storage to internal engineering teams, including the Kubernetes platform. It also owns storage performance and capacity, the shared-filesystem strategy for AI workloads (including evaluation of high-performance parallel filesystems), security response for the storage stack, and 24/7 storage on-call.
Core stack: Ceph (RBD, CephFS, RGW multisite), NVMe, HAProxy, Ubuntu, containers, high speed fabric storage networks, Gitlab, VictoriaMetrics, Grafana, PagerDuty.
The role
You will lead the Storage team as its Engineering Manager, reporting to the VP of Infrastructure. The team is small, senior, and growing as Verda opens new datacenters at a rate of greater than one a month. Your job is to turn deep individual expertise into a mature team that can scale: you own its people, its delivery, and the reliability and durability of the data it holds.
You own the team and its growth. Hiring, onboarding, career development, and building depth so the team's knowledge does not live in one or two heads. You decide what the team should look like in a year, and you build it.
You own priorities and projects. Storage sits underneath almost everything Verda sells: VM volumes, shared filesystems, object storage, dedicated clusters, new sites. Every team needs something from Storage. You set the priorities, give requests a clear front door, say "not yet" when it's the right answer, and drive multi-quarter projects to completion. When someone needs to know what Storage will deliver and when, you are the answer.
You own maturity and reliability as an outcome. A storage incident can mean lost data and lost customer trust, and you are accountable for it not happening. That means defining durability and availability targets for each service the team runs, running the sheriff and on-call rotation and escalation path, owning incident response and postmortems through to closed actions, managing capacity, upgrade and hardware risk, and deciding when reliability work takes precedence over feature work. You replace heroics with process: runbooks, change management, upgrade playbooks, and a repeatable storage recipe for every new datacenter.
The tech lead and Staff Engineer own technical architecture and direction. How the storage platform is built, which backends it uses, and how services are designed to meet the targets you set. You decide the technical roadmap together; on questions of how, they have the final say, and on who, what, when and whether, you do.
This role requires real storage engineering depth. You should be able to read a Ceph design doc, a rebalancing plan or an incident timeline and argue with it: you know why etcd does not belong on distributed storage, and what a read hotspot on a handful of OSDs looks like. But your primary output is a team that ships reliably, keeps growing, and stays healthy.
Your responsibilities
People and growth
Manage the Storage team directly: 1:1s, coaching, career development, performance, and compensation input.
Own the hiring plan. Define profiles and levels, balancing strong generalist SRE and infrastructure engineers who have storage at heart against deep specialists, and knowing when expertise is better bought from vendors and consultants than hired. Run a fast and fair process and onboard new engineers to productivity quickly.
Spread storage knowledge through pairing, internal training and documentation, so every subsystem has more than one person who can own it at 3am.
Maturity and process
Establish clear ownership areas within the team across block storage and VM volumes, shared filesystems, object storage, dedicated customer clusters, new-site buildout, and performance and capacity.v
Run a sustainable sheriff and on-call rotation with escalation policies, runbooks, blameless postmortems and written handovers.
Introduce change management for high-risk operations (major Ceph upgrades, rolling node maintenance, rebalancing, multi-site changes) so routine work does not cause customer-visible impact.
Own the storage security and compliance posture with the Security team: vulnerability response, encryption, data lifecycle and sanitisation, and customer-data handling rules.
Track service health and delivery metrics (durability, availability, latency, capacity headroom, incident trends) and act on them.
Priorities
Introduce lightweight triage and prioritisation so requests from Cloud, Provisioning, Platform, AI/ML and customer-facing teams have a clear front door and unplanned work is visible and managed.
Represent the team to Infrastructure peers (SRE, Provisioning, Network, Platform, Observability), to product teams, and to customers. Negotiate scope and priorities on the team's behalf.
Work with DC Ops and procurement on capacity planning, hardware and vendor evaluations, turning a technical recommendation into a decision with cost, risk and timeline attached.
Projects
Turn the technical roadmap into quarterly plans with owners, milestones and honest status, and drive multi-quarter efforts to completion: the Object Storage customer launch, fleet-wide Ceph upgrades, a single shared-filesystem strategy for AI workloads, VM volume features (snapshots, cloning, multi-attach, resize, encryption), and repeatable storage turn-up for every new datacenter.
Your key competencies
3+ years managing infrastructure or software engineers, including hiring, performance management, and leading a team through significant growth or change.
5+ years hands-on with distributed storage in production, ideally Ceph, including cluster lifecycle: upgrades, recovery and rebalancing, capacity expansion, hardware failures, and the incidents that come with them.
A solid grasp of block, file and object storage and their trade-offs: RBD vs CephFS vs NFS vs S3, replication vs erasure coding, what drives latency and throughput, and how storage interacts with the network.
Linux systems and performance depth: NVMe, filesystems, storage networking (MTU, bonding), and performance analysis.
Experience designing or fixing an on-call rotation and incident process, and clear views on what makes one humane.
A track record of taking a team from heroics to process: runbooks, change management, planning and metrics.
Enough technical credibility that engineers trust your judgement because you understand the work.
Strong written communication; much of this role is turning ambiguity into plans others can act on.
Nice to have
Ceph RGW multisite, S3 compatibility, and running object storage as a customer-facing product.
High-performance shared or parallel filesystems for AI: Lustre, VAST, DDN, WEKA, GPUDirect Storage, RDMA or NVMe-oF.
Storage for GPU clusters and ML training or inference (checkpointing, dataset throughput), including NVIDIA reference architectures.
Storage for virtualisation (QEMU/KVM) and Kubernetes (CSI, Rook).
Running storage across many datacenters, or bringing up new sites repeatably.
Vendor management and structured hardware or software evaluations.
Storage security and compliance: encryption at rest, data sanitisation, ISO 27001 or SOC 2.
What's next
We're building fast and this role needs the right person behind it. There's no artificial deadline, but when we find who we're looking for, we move. If this sounds like your next move, apply now.
Please submit your application through our Careers page. We don't accept applications sent by email.
Required Skills
Required Languages
🇬🇧 English