Senior DevOps Engineer
Highlights:
- Location: Spain Remote
- Language: Russian-speaking team
About Fundraise Up
We're Fundraise Up — a global fundraising platform built to make donating to nonprofits fast, seamless, and accessible to all. Every month, our technology powers tens of millions of dollars in donations across the globe. Our platform is trusted by many of the world's leading nonprofits, including UNICEF, the Alzheimer's Association, and a wide range of global NGOs, with a 4.9/5 rating across top software review platforms. Our mission is to unlock the world's generosity potential by optimizing how people give.
We're a distributed team of 160+ product professionals, including 80+ engineers, running a large ecosystem: checkout and embeddable widgets, donor, organization and partner portals, admin tools and internal apps. The stack is Node.js (Koa, NestJS), MongoDB, TypeScript, Vue.js and React, with Kafka and Bull (Redis) for messaging and background jobs, ClickHouse for analytics and Elasticsearch for search.
About the Role
On a busy giving day, ten minutes of downtime can cost around $500k — we need the engineer who makes sure it doesn't happen.
You'll join the DevOps team that owns the platforms every engineering team relies on daily: CI/CD, observability, logging and developer tooling. This is a senior-only role. We'll hand you an entire area — observability or CI/CD — and trust you to run it, from requirements and design through rollout, operations and mentoring others. You'll be the go-to technical reference for that area and the senior escalation point when things get hard.
What You’ll Do
- Own one of our core platform areas end-to-end: observability (VictoriaMetrics, Grafana, Graylog / VictoriaLogs, fluent bit, exporters, alerting) or CI/CD (Jenkins scripted pipelines, Harbor, Nexus, build agents) — you drive its architecture, reliability, and roadmap.
- Drive technical initiatives end-to-end: gather requirements, write the design doc, decompose into tasks, implement, deliver to production, and own the operational health afterwards.
- Drive clarity in ambiguous situations by defining requirements, assumptions, and next steps.
- Design for reliability and scale: evolve the architecture of our platforms — topology, integration points, scaling approach, and reliability model.
- Support developers: deploy and monitor applications on both on-premise servers and Kubernetes (Helm), troubleshoot builds and deploys, help teams with metrics, alerts, and logs; participate in chat duty in developer support channels.
- Automate away toil: repetitive operations, provisioning, and maintenance should be codified, not performed by hand.
- Investigate production incidents as the senior escalation point for your area: drive resolution, lead post-mortems, implement systemic fixes. Participate in on-call rotations and raise the bar for how on-call works.
- Mentor less experienced engineers through design discussions, reviews, and pairing; catch debt-inducing shortcuts at the review stage.
- Use AI in all aspects of day-to-day work: researching, troubleshooting, developing.
Requirements
- 6+ years as a DevOps Engineer / SRE (or very close responsibilities).
- Track record of owning technical initiatives end-to-end — from requirements and technical design through production delivery. You can showcase initiatives that were yours, not just tasks you completed.
- Confident Linux skills (we use Ubuntu).
- Working knowledge of the Prometheus stack: metric types, exporters, and how alerting works — enough to navigate and extend an existing setup.
- Hands-on experience with CI/CD: pipeline design, build orchestration, artifact delivery.
- Containers: Docker, image building, registries.
- Ansible.
- Git.
- Experience with Bash or Python scripting for automation and observability (writing exporters, eliminating routine work).
- Production/on-call experience: diagnosing incidents, restoring service, leading post-mortems.
- Experience mentoring less experienced engineers.
- Ownership and attention to detail. Downtime is expensive: during busy events 10 minutes of downtime can cost us around $500k.
Must Have
We understand it’s impossible to be an expert in everything, but it’s important to have solid hands-on experience in two or more of the areas below:
- VictoriaMetrics / Prometheus stack at scale: architecture, cardinality control, exporters, alerting infrastructure.
- Log pipelines at scale: Graylog / VictoriaLogs / ELK — collection (fluent bit or similar), retention, sharding, performance.
- Jenkins scripted pipelines: shared libraries, pipeline infrastructure, build agent fleets.
- Container registries and artifact management: Harbor, Nexus, base images, image policies.
- Operating applications on Kubernetes: Helm, workload monitoring and log delivery, deploy troubleshooting.
- Grafana: dashboards as code, alerting, performance at scale.
Bonus points
Great if you’ve worked with any of the following:
- Analytics & DS platforms: JupyterHub, Airflow, Tableau, MLflow, Airbyte — deployment, maintenance, resource limits. Building platform around these tools to improve Quality of Life for Analytics.
- Remote development environments and AI agent execution environments. E.g. Coder/Telepresence.
- Bare-metal Kubernetes: provisioning, networking, scaling.
- Flux and GitOps.
- Terraform.
- Sentry on-premise: operating self-hosted error tracking.
- ClickHouse, MongoDB.
You'll Thrive Here If You…
- Want an area that's truly yours, and the accountability that comes with it.
- Commit to your own estimates, and flag early when something is slipping instead of hoping it won't.
- Ask for help at 2 a.m. rather than quietly going down the rabbit hole alone.
- Explain a complex incident in two minutes, not twenty.
- Disagree openly, make your case with data, then commit and move.
- Treat every manual step as a bug waiting to be automated.
- Get more satisfaction from making other engineers self-sufficient than from being the only one who knows how it works.
- Enjoy a fast-moving team that's rebuilding how it works around AI, and stay curious as the tools change.
Compensation
- The compensation range for this role is: €5,800 - €6,600 gross per month.
- Final compensation will be determined based on relevant experience, skills, qualifications, and alignment with the role's requirements.
Perks & Benefits
- Private medical insurance for the employee and their family
- 23 paid vacation days per year
- 11 paid public holidays per year
- 5 company-paid sick leave days
- English learning courses
- Relevant professional education
- Gym or swimming pool
- Home Office Setup Assistance: the company offers assistance with purchasing furniture (office chair, office desk, monitor) and other items to create a comfortable workspace
- Co-working
- Remote working
We are an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, color, religion, gender, sexual orientation, gender identity or expression, age, national origin, disability, or any other characteristic protected by applicable law in the countries where we operate.
Required Skills
Required Languages
🇬🇧 English