A collection of real-world Data Engineering experience - from building pipelines and migrating data warehouses, to leading teams and building industry-scale data platforms.
Overview of cross-industry data engineering patterns covered in these use cases
In agritech HORECA, fast and massive data collection is required with distributed sources and no official APIs.
Optimized data pipelines in an agritech environment to reduce BigQuery costs and improve performance through data usage audits, real-time vs batch segmentation, and query and architecture efficiency improvements.
Led migration from Oracle Data Warehouse to Hadoop Ecosystem at one of Indonesia's largest laboratory companies, without disrupting existing analytical workflows.
Transformed 5 System Analysts into productive Data Engineers while keeping operations running in a fast-paced agritech environment.
Built an end-to-end data platform on the Hadoop ecosystem to integrate SCADA, Kafka, and external data sources for mining operational analytics.
The Story
The HORECA sector — hotels, restaurants, and cafés — holds enormous commercial data, but none of it is available through official APIs. The business team needed comprehensive market coverage, fast. I designed a distributed web crawling system using Python and Selenium, built to run across multiple devices simultaneously using multi-threading. Data flowed directly into PostgreSQL with consistent schema and deduplication. What used to take days completed in hours. The output — clean, structured CSV exports — went straight into the canvassing team's workflow.
In agritech HORECA, fast and massive data collection is required with distributed sources and no official APIs.
📊 Impact
🧩 Tech Stack
Python, Selenium, Threading/Concurrent Execution, PostgreSQL, CSV Export Pipeline
⚡ Problem Statement
🧠 Solution Overview
🏗️ Architecture
🔥 Challenges & Solutions
Optimize data pipeline in Agritech to lower cost and improve performance.
📊 Impact
🧩 Tech Stack
Python, BigQuery, Fivetran, Apache Airflow, Advanced SQL
⚡ Problem Statement
🧠 Solution Strategy
🏗️ Architecture
🔥 Challenges & Solutions
I led as Lead Data Engineer on a data warehouse migration from Oracle to the Hadoop Ecosystem at one of Indonesia's largest laboratory companies, without disrupting existing analytical workflows.
📊 Impact
🧩 Tech Stack
Oracle Database, Hadoop (HDFS), Hive, Parquet, Oozie, Sqoop/NiFi
⚡ Problem Statement
🧠 Solution Overview
🏗️ Architecture
🔥 Challenges & Solutions
During a scaling phase in agritech, I led the build-up of an internal Data Engineering capability - transforming 5 System Analysts into productive Data Engineers, while ensuring operations continued without disruption.
📊 Impact
🧩 Tech & Concepts
Data Engineering Fundamentals, Medallion Architecture (Bronze & Silver), Data Pipeline Design, Data Quality Validation, Schema Migration Handling, Mentoring & Team Building
⚡ Problem Statement
🧠 Approach
🏗️ Architecture & Responsibilities
🔥 Key Challenges & Solutions
In the mining industry, operations rely heavily on real-time data from SCADA systems. I built an end-to-end data platform on the Hadoop ecosystem to integrate multi-source data and deliver analytical cubes that support operational decision making.
📊 Impact
🧩 Tech Stack
Apache NiFi, Apache Kafka, Hadoop (HDFS), Hive, SCADA Integration, BMKG API, Analytical Cube Design, Virtual Machine Infrastructure
⚡ Problem Statement
🧠 Solution Overview
🏗️ Architecture
🔥 Key Challenges & Solutions
Building retrieval pipelines and AI-ready data infrastructure — a natural extension of data engineering into the LLM ecosystem.
RAG (Retrieval-Augmented Generation)
Context retrieval architecture for accurate LLM output
Vector Database
Embedding storage & similarity search (Qdrant, ChromaDB)
Text Chunking Strategy
Document splitting strategies for optimal retrieval quality
Embedding Pipeline
Text-to-vector transformation for semantic search
LLM Integration
Connecting data pipelines to language models (Ollama, OpenAI)
Explore other domains