Data Engineer

PulseRise Technologies

Job description

We're looking for a Middle+ Data/Databricks Engineer to join our data team and take ownership of pipelines built on Databricks. You'll design and run production-grade ETL/ELT workloads, work with the lakehouse (Delta Lake) architecture, and partner closely with analysts, data scientists, and business stakeholders who rely on the data you deliver. "Middle+" here means roughly 3–5 years of hands-on data engineering experience: you can own a pipeline end to end with minimal supervision, make sound design decisions, and are starting to mentor more junior engineers — without yet carrying full architectural ownership of the platform. Contract & Logistics Contract type: B2B (valid Czech IČO / trade license required). Location: Remote, candidate must be based in the Czech Republic; occasional visit to Prague office. Start date: ASAP English: Fluent Czech: Native/Fluent What You'll Do Design, build, and maintain scalable ETL/ELT pipelines on Databricks using PySpark and SQL. Implement and evolve Delta Lake tables following a medallion (bronze/silver/gold) architecture. Develop batch and, where needed, streaming data processing jobs (Structured Streaming). Build and maintain CI/CD for Databricks notebooks and jobs (Databricks Asset Bundles, Repos, Git-based workflows). Orchestrate pipelines with Databricks Workflows, Azure Data Factory, Airflow, or similar tools. Monitor job performance and cluster costs; troubleshoot failures and optimize for reliability and spend. Implement data quality checks, testing, and monitoring for critical datasets. Collaborate with analysts, data scientists, and business stakeholders to translate requirements into data models. Participate in code reviews, document technical decisions, and help shape engineering best practices. Provide guidance and informal mentorship to junior team members. What We're Looking For 3+ years of hands-on experience with Apache Spark, ideally on Databricks specifically. Strong PySpark and SQL skills, including performance tuning of Spark jobs. Practical experience with Delta Lake and lakehouse architecture concepts. Solid grasp of data modeling (dimensional modeling, star schema) and ETL/ELT design patterns. Experience with at least one major cloud platform — Azure, AWS, or GCP (Azure Databricks is a plus). Familiarity with pipeline orchestration tools (Databricks Workflows, ADF, Airflow, or equivalent). Working knowledge of Git and CI/CD practices for data engineering. Understanding of data governance, access control, and security basics (Unity Catalog is a plus). Good written and spoken English; Czech or Slovak is a plus but not required. Comfortable working independently and communicating clearly in a remote, distributed team. Nice to Have Databricks certification (Data Engineer Associate or Professional). Experience with streaming technologies such as Kafka or Azure Event Hubs. Experience with dbt for transformation and testing. Solid Python software-engineering practices — testing, packaging, typing. Infrastructure-as-code experience (Terraform). Exposure to MLOps or ML pipeline support. Experience in a regulated or large enterprise data environment. What We Offer Fully remote work with flexible hours, based anywhere in the Czech Republic. Long-term B2B cooperation with a stable pipeline of work. A modern, actively evolving data stack rather than legacy maintenance work. Direct collaboration with senior engineers and architects — real input on design decisions. Budget for training, certifications, and conference attendance. A flat structure with fast decision-making and minimal bureaucracy.

Skills

  • airflow
  • aws
  • azure
  • ci/cd
  • databricks
  • etl
  • gcp
  • git
  • pyspark
  • sql

Languages

EN, CS

Apply

Open this job in our interactive board to apply, save it, or sign up for matched alerts on similar roles.

View & apply
Data Engineer — PulseRise Technologies | RVC