Senior Batch Data Engineer (Alibaba Cloud Stack)
OneBullEx is looking for a highly accountable, self-driven Senior Batch Data Engineer to take end-to-end ownership of our enterprise batch data warehouse on Alibaba Cloud. In this role, you will design, maintain, and optimise batch pipelines (T-1) supporting executive dashboards, business analytics, and AI/ML services — leveraging modern AI-assisted engineering techniques to rapidly build data layers (ODS, DWD, DWS, ADS) and enforcing strict data quality, automated validation, and proactive monitoring across the platform.
Key Responsibilities
Pipeline Architecture & Data Warehouse Modelling
- Design, build, and maintain scalable ETL/ELT batch pipelines using Alibaba Cloud DataWorks and MaxCompute.
- Rapidly design and build enterprise data warehouse layers (ODS, DWD, DWS, ADS) using modular SQL/Python scripts and AI generation tools.
- Develop and optimise data ingestion, transformation, and batch processing workflows across Lakehouse and Data Warehouse architectures (MaxCompute, Hologres, DataWorks, DTS).
- Ensure data reliability, high query performance, schema stability, and cost-effective cloud resource usage.
AI-Accelerated Engineering & Testing
- Integrate modern AI tools and techniques (e.g. Claude, GitHub Copilot, prompt engineering) into daily workflow to accelerate code generation, SQL refactoring, and data validation.
- Apply new AI testing methodologies to rapidly write unit tests, simulate data scenarios, and audit data accuracy prior to production release.
Data Quality, Monitoring & Alerting
- Implement strict data quality rules, automated reconciliation scripts, and validation checks directly inside DataWorks jobs.
- Build proactive alerting and monitoring workflows (integrated with Lark) to detect data drift, schema breaks, or pipeline failures before business impact occurs.
- Support downstream analytics, BI reporting, and AI/LLM services with clean, trusted datasets.
End-to-End Ownership & Collaboration
- Take full ownership of daily T-1 batch pipeline stability, job scheduling, historical backfills, and legacy issue cleanup.
- Act as a bilingual technical bridge, collaborating fluently in both English and Chinese with cross-functional technical leads, product managers, and developers.
Required Qualifications
- 5+ years of hands-on expertise with Alibaba Cloud data services: DataWorks, MaxCompute, DTS, Hologres, FC Functions, Data Agents, and MaxCompute Lakehouse architectures.
- Advanced proficiency in SQL and Python for data engineering, performance tuning, and database optimisation.
- Proven experience modelling multi-layer enterprise data structures (ODS, DWD, DWS, ADS) and managing batch Lakehouse architectures.
- Fluent in both English and Chinese, spoken and written, for seamless technical collaboration with regional teams.
- Strong sense of ownership, high velocity, proactive communication, and a "deliver fast, iterate continuously" attitude.
Preferred
- Active experience leveraging AI coding assistants (Copilot, LLMs) to speed up ETL development, code reviews, and query optimisation.
- Experience building alerting and monitoring integrations with Lark or a similar collaboration platform.
- Track record applying AI-driven testing methodologies ahead of production release.
Required Skills
Required Languages
🇬🇧 English, 🇨🇳 Chinese