RVC
JobsFor Employers
  1. Jobs
  2. /
  3. Kafka Expert

Kafka Expert

MeteorOps | IT consulting
12h 55m ago
Worldwide Remote
Freelance

MeteorOps is looking for a freelance Kafka troubleshooting & modernization specialist to step into an older on-prem Kafka environment that currently has no dedicated Kafka owner and limited observability. The cluster supports real-time market quote / HFT tick data at very high throughput (potentially millions of messages/sec) and feeds downstream systems including downsampling services and a SQL Server writer, eventually supporting trading execution workflows.

The Kafka setup is 6–7 years old, deployed on VMware on-prem VMs with 10 Kafka brokers and 5 ZooKeepers, running Kafka 2.13-3.0.0. Each broker has multiple data disks (currently stated as 7 disks ~1TB each; prior notes mention higher disk counts—part of the engagement will be to verify actual layout). Historically disk usage sits around ~10%, but recently one or more brokers spiked toward 100%, coinciding with application Kafka errors and broker/topic instability (e.g., missing leader, invalid partition, impaired topic failover).

You’ll diagnose the incident and underlying risks, produce a clear findings + recommendations report, and help the engineering team implement pragmatic improvements: monitoring, tooling, operational runbooks, resilience/failover improvements, and an assessment of upgrade options (including a path away from ZooKeeper).

What you'll do

  • Rapid triage & incident diagnosis
    • Confirm scope of the recent issue (disk saturation, broker health, controller/ZK health, partition leadership, ISR, replication, rebalances).
    • Determine why disk utilization jumped from ~10% to near 100% (retention changes, log segment growth, stuck cleanup, partition skew, under-replication, data directory imbalance, etc.).
    • Identify root causes of missing leader, topic access failures, and invalid partition behavior.
  • Cluster assessment & hardening plan
    • Review broker configuration, topic settings (replication factor, min.insync.replicas, retention), partition distribution, rack awareness (if any), and failover behavior.
    • Evaluate ZooKeeper reliability and operational risks; document current failure domains and bottlenecks.
  • Observability & operations uplift
    • Propose and/or implement proper Kafka monitoring (broker + ZK + OS/disk), dashboards, and alerting (lag, under-replication, disk, controller events, request latency, GC, network, etc.).
    • Recommend stronger GUI/management tooling beyond read-only usage (currently Kafdrop and Zabbix are used but limited).
  • Engineering enablement
    • Produce a Findings Report and Recommendations / Roadmap (quick wins + medium/long-term).
    • Create runbooks for safe operations: broker restart procedure, partition reassignments, capacity checks, config backups, upgrades with minimal disruption, and recovery steps.
    • Coach the engineering team and IT on day-to-day Kafka ops and troubleshooting patterns.
  • Optional improvement work
    • Execute selected remediations (e.g., storage rebalancing, retention tuning, partition reassignment, leader imbalance fixes).
    • Assess and plan Kafka upgrade strategy, including ZooKeeper removal (KRaft migration path) if appropriate for their risk tolerance and timelines.
    • Improve resilience posture toward minimized RTO/RPO (goal: “as low as practical,” possibly ~1 minute max data loss tolerance).

What we're looking for

Must-have

  • Proven hands-on experience operating Kafka in production, including high-throughput clusters.
  • Strong troubleshooting of:
    • Partition leadership issues, missing leaders, ISR shrinkage, under-replicated partitions
    • Broker restarts and safe recovery without “sledgehammer” approaches
    • Storage/disk issues on multi-disk broker layouts (JBOD patterns, partition skew, log retention/cleanup behavior)
  • Linux systems competence: disk/IO analysis, filesystem saturation, process/resource analysis, networking basics.
  • Experience with ZooKeeper-based Kafka clusters and operational best practices.
  • Ability to deliver clear, actionable documentation: findings, recommendations, and runbooks.
  • Strong communication skills for working with a mixed team (engineering + IT unfamiliar with Kafka).

Nice-to-have

  • Experience with Kafka monitoring stacks (e.g., JMX metrics pipelines, Prometheus/Grafana, lag monitoring, alerting design).
  • Experience with GUI/admin tooling and governance practices (RBAC, auditing approach, safer topic/config workflows).
  • Experience planning Kafka upgrades and migrations, including evaluation of KRaft readiness and risk.
  • Familiarity with workloads involving market data / trading systems and latency-sensitive pipelines.
  • Experience with VMware-based on-prem operations and capacity planning.

Apply for Kafka Expert

Upload your CV and links. A person reviews every application and we'll get back to you either way.

Frequently Asked Questions

Required Skills

kafkavmwarelinuxprometheusgrafanazabbix

Required Languages

🇬🇧 English

Related searches

  • Fully Remote Jobs
1 jobs
Sort by
12h 55m ago

Kafka Expert

MeteorOps·IT consulting
kafkavmwarelinuxprometheusgrafana+1
📡Worldwide Remote
General Infra/Network
No similar jobs match these filters. Try changing or clearing a filter.
Remote roles
PythonJavaReactGoDevOpsNode.jsC# / .NET
Countries
United StatesUnited KingdomCanadaUS & EMEAGermanyPolandSpainNetherlandsPortugal
Experience
SeniorMid-levelJunior
R© 2026 RVC Globalbuild 4417945a
AboutMCPPricingContactPrivacyCookiesRefundsTerms & ConditionsFor Employers