Tugas Besar Analisis Big Data β Institut Teknologi Sumatera 2026
Implementasi Medallion Architecture berbasis Apache Spark & Docker untuk Analisis Ketimpangan Pendapatan (SDG 10)
- Deskripsi Proyek
- Anggota Tim
- Arsitektur Sistem
- Teknologi
- Dataset
- Cara Menjalankan
- Struktur Folder
- Pipeline Medallion
- Dashboard Interaktif
- Benchmark & Evaluasi
- Lisensi Data
- Status Proyek
- Referensi
Proyek ini merancang dan mengimplementasikan sistem pemrosesan data skala besar untuk menganalisis ketimpangan pendapatan dalam konteks Sustainable Development Goal 10 (SDG 10 Reduced Inequalities).
Menggunakan Medallion Architecture (Bronze β Silver β Gold) yang dijalankan pada Apache Spark cluster terdistribusi dan dikontainerisasi dengan Docker, sistem ini mampu memproses ratusan ribu hingga jutaan baris data mikro sensus individu secara efisien, menghasilkan metrik ketimpangan (Gini coefficient, Palma ratio, Theil index, dan shared prosperity premium), serta menyajikannya melalui dashboard interaktif Streamlit.
Apakah arsitektur Medallion berbasis Apache Spark yang terkontainerisasi dengan Docker mampu meningkatkan throughput pemrosesan dan mengurangi latensi end-to-end pipeline untuk data mikro sensus ketimpangan pendapatan dibandingkan dengan pipeline pemrosesan sekuensial berbasis Pandas?
| No | Nama | Peran | Tanggung Jawab Utama |
|---|---|---|---|
| 1 | Ginda Fajar Riadi Marpaung | π― Ketua / Project Integrator | Koordinasi harian, merge kode, finalisasi proposal & presentasi, integrasi antar-modul |
| 2 | Vany Salsabilla Putri | ποΈ Data Engineer | Handle data IPUMS, bangun Bronze β Silver layer, data quality & profiling |
| 3 | Fathya Intami Gusd | β‘ Spark Analytics Developer | Bangun Gold layer (Gini UDF, kuintil, Theil index), optimasi query Spark |
| 4 | Malika Azzahra Salsabila | π Baseline & Benchmark Specialist | Pipeline Pandas sekuensial, ukur throughput & latensi, bandingkan Spark vs Pandas |
| 5 | Luthfia Laila Ramadhani | π₯οΈ Dashboard & DevOps Engineer | Setup Docker Compose (Spark + Streamlit), bangun dashboard interaktif, diagram arsitektur |
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β DOCKER CONTAINER ORCHESTRATION β
β ββββββββββββ ββββββββββββ ββββββββββββ ββββββββββββ ββββββββββ β
β β Jupyter β β Spark β β Spark β β Spark β βStreamlitβ β
β β (Driver) β β Master β β Worker 1 β β Worker 2 β βDashboardβ β
β ββββββ¬ββββββ ββββββ¬ββββββ ββββββ¬ββββββ ββββββ¬ββββββ βββββ¬βββββ β
β β β β β β β
β βββββββββββββββ΄βββββββ¬βββββ΄ββββββββββββββ β β
β βΌ β β
β βββββββββββββββββββ β β
β β SHARED VOLUME β β β
β β /data/ β β β
β β βββ raw/ β β IPUMS CSV (ignore) β β
β β βββ bronze/ β β Parquet hasil ingest β β
β β βββ silver/ β β Parquet hasil clean β β
β β βββ gold/ β β Parquet hasil agregasiβ β
β βββββββββββββββββββ β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MEDALLION ARCHITECTURE PIPELINE β
β β
β [IPUMS CSV] βββΊ [BRONZE] βββΊ [SILVER] βββΊ [GOLD] βββΊ [UI] β
β (Raw) (Ingest) (Transform) (Analytics) β
β β
β β’ Validasi skema β’ Filter null β’ Agregasi kuintil β
β β’ Parquet β’ Deduplikasi β’ Gini coefficient (UDF) β
β β’ Partisi β’ Normalisasi β’ Palma ratio β
β PPP β’ Theil index β
β β’ Shared prosperity β
β β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
| Kategori | Teknologi | Versi | Fungsi |
|---|---|---|---|
| Big Data Engine | Apache Spark | 3.5 | Pemrosesan terdistribusi, PySpark API, Spark SQL |
| Containerization | Docker | Latest | Kontainerisasi Spark cluster + Jupyter + Streamlit |
| Orchestration | Docker Compose | Latest | Manajemen multi-container 1 perintah |
| Bahasa | Python | 3.11 | PySpark, Pandas, Streamlit |
| Dashboard | Streamlit | Latest | UI interaktif, filter, visualisasi real-time |
| Storage Format | Apache Parquet | β | Columnar storage, kompresi Snappy, efisiensi query |
| Lakehouse (opsional) | Delta Lake | Latest | ACID transactions, time travel, schema enforcement |
| Version Control | Git + GitHub | β | Kolaborasi tim, tracking perubahan |
| Atribut | Spesifikasi |
|---|---|
| Nama | IPUMS International (Integrated Public Use Microdata Series) |
| Pengelola | Minnesota Population Center, University of Minnesota |
| Negara | Brazil 2010, Mexico 2010 |
| Unit Observasi | Individu (person records) |
| Estimasi Baris | 32,257,874 baris |
| Ukuran File Mentah | ~1.2 β 2.5 GB (CSV hasil ekstraksi dari .csv.gz) |
| Variabel Inti | INCTOT, INCEARN, PERWT, AGE, SEX, EDATTAIN, EMPSTAT, OCCISCO, INDGEN |
| Lisensi | Academic/Research Use Only dan redistribution dilarang |
β οΈ Peringatan: Data mentah IPUMS tidak boleh di-push ke GitHub publik sesuai ketentuan lisensi. Hanya kode pipeline, hasil agregat (Gold layer), dan dokumentasi yang dipublikasikan.
- Docker Desktop terinstall
- Git terinstall
- Minimal RAM 8 GB (16 GB direkomendasikan)
git clone https://github.com/JARS-17/sdg10-bigdata-itera.git
cd sdg10-bigdata-iteracd docker
docker-compose up -dVerifikasi semua service berjalan:
docker-compose ps| Service | URL | Keterangan |
|---|---|---|
| Jupyter Notebook | http://localhost:8888 | Development & submit Spark jobs |
| Spark UI (Master) | http://localhost:8080 | Monitor cluster, job, stage, task |
| Streamlit Dashboard | http://localhost:8501 | Dashboard interaktif hasil analisis |
# Di terminal Jupyter container atau terminal lokal dengan Spark
python scripts/bronze_layer.py # Ingesti CSV β Parquet
python scripts/silver_layer.py # Cleaning β Transformasi
python scripts/gold_layer.py # Agregasi β Metrik SDG# Dashboard otomatis berjalan jika docker-compose sudah up
# Atau jalankan manual:
streamlit run dashboard/dashboard.py --server.address=0.0.0.0docker-compose down
# atau hapus semua data volume:
docker-compose down -vsdg10-bigdata-itera/
βββ π data/
β βββ π raw/ β DATA MENTAH IPUMS (excluded dari Git)
β βββ π bronze/ β Hasil ingest: Parquet terpartisi
β βββ π silver/ β Hasil transformasi
β βββ π gold/ β Hasil agregasi
βββ π notebooks/
β βββ 01_baseline_pandas.ipynb β Pipeline baseline Pandas
β βββ 02_eda_ipums.ipynb β Eksplorasi data
β βββ 03_benchmark_analysis.ipynb β Analisis perbandingan Spark vs Pandas
βββ π scripts/
β βββ bronze_layer.py β Ingesti & validasi skema
β βββ silver_layer.py β Pembersihan & transformasi
β βββ gold_layer.py β Agregasi & perhitungan metrik
β βββ utils.py β Fungsi Gini, Theil, Palma ratio
β βββ config.py β Konstanta: path, variabel, threshold
βββ π dashboard/
β βββ dashboard.py β Aplikasi Streamlit utama
β βββ π components/
β β βββ gini_chart.py β Visualisasi Gini coefficient
β β βββ quintile_table.py β Tabel kuintil interaktif
β β βββ income_dist.py β Histogram distribusi income
β βββ π assets/
β βββ logo.png β Logo/logo tim
βββ π docker/
β βββ docker-compose.yml β Definisi 5 service container
β βββ Dockerfile.spark β Custom image Spark + Delta Lake
βββ π docs/
β βββ proposal.docx β Proposal tugas besar
β βββ arsitektur-diagram.png β Diagram arsitektur sistem
βββ π tests/
β βββ test_gini.py β Unit test perhitungan Gini
βββ .gitignore β Exclude data mentah & file besar
βββ README.md β Dokumentasi ini
βββ Makefile (opsional) β Perintah otomatisasi
# Contoh: bronze_layer.py
from pyspark.sql import SparkSession
spark = SparkSession.builder \
.appName("SDG10-Bronze") \
.getOrCreate()
df = spark.read.csv("/data/raw/ipums_brazil_mexico_2010.csv",
header=True, inferSchema=True)
df.write.parquet("/data/bronze/ipums_bronze.parquet",
partitionBy=["COUNTRY", "YEAR"])# Contoh: silver_layer.py
from pyspark.sql.functions import col, when
df_silver = df_bronze \
.filter(col("INCTOT").isNotNull()) \
.dropDuplicates(["SAMPLE", "SERIAL", "PERNUM"]) \
.withColumn("income_ppp", col("INCTOT") * 0.85) # normalisasi PPP# Contoh: gold_layer.py
from scripts.utils import calculate_gini
df_gold = df_silver.groupBy("COUNTRY", "YEAR") \
.agg(calculate_gini("INCTOT", "PERWT").alias("gini_coeff"))| Fitur | Deskripsi | Interaktivitas |
|---|---|---|
| Filter Negara | Pilih Brazil atau Mexico | Dropdown sidebar |
| Bar Chart Gini | Perbandingan Gini coefficient | Hover tooltip |
| Histogram Income | Distribusi pendapatan per kuintil | Slider rentang |
| Tabel Kuintil | Q1βQ5 dengan jumlah populasi | Sort & search |
| Box Plot Sektir | Income per industri (INDGEN) | Drill-down |
| Anomaly Flag | Highlight jika bottom 40% < 50% median | Auto-detect |
Dashboard akan di-deploy di
localhost:8501setelah pipeline Gold layer berjalan.
| Metrik | Baseline (Pandas) | Target (Spark) | Speedup |
|---|---|---|---|
| Throughput Ingesti | 2.500 baris/dtk | β₯ 25.000 baris/dtk | 10Γ+ |
| Latensi End-to-End | ~18 menit | β€ 4 menit | 4.5Γ+ |
| Rasio Kompresi | 1.0Γ (CSV) | β₯ 3.0Γ (Parquet) | 3Γ |
| Akurasi Gini | Β±0.015 (exact) | Β±0.020 (approximate) | Acceptable |
Evaluasi diukur pada hardware identik: RAM 16 GB, 4-core CPU, SSD.
Data IPUMS International digunakan berdasarkan Academic Use License dari Minnesota Population Center dan kantor statistik mitra nasional. Ketentuan utama:
- β Penggunaan untuk penelitian dan pendidikan
- β Redistribution data mentah dilarang
- β Commercial use dilarang
- β Re-identification individu dilarang
- β Publikasi hasil agregat diperbolehkan dengan sitasi
Setiap anggota tim harus memiliki akun IPUMS International yang teregistrasi secara individual.
| Milestone | Status | Hari Target |
|---|---|---|
| Setup Infrastruktur | π’ Done | Hari 1 |
| Bronze Layer | π’ Done | Hari 2 |
| Silver Layer | π’ Done | Hari 3 |
| Gold Layer | π’ Done | Hari 4 |
| Dashboard Streamlit | π’ Done | Hari 5 |
| Benchmark & Polish | π’ Done | Hari 6 |
| Final Testing & Submit | π‘ Progress | Hari 7 |
Timeline: 7 Hari (Senin - Minggu)
Metodologi: Agile Daily Sync (19:00 WIB)
[1] Y. Liu et al., "A big data approach to assess progress towards Sustainable Development Goals for cities of varying sizes," Communications Earth & Environment, vol. 4, no. 1, p. 82, 2023.
[2] Steven Ruggles, Lara Cleveland, Rodrigo Lovaton, Sula Sarkar, Matthew Sobek, Derek Burk, Dan Ehrlich, Jane Lee, and Nate Merrill. Integrated Public Use Microdata Series, International: Version 7.6 [dataset]. Minneapolis, MN: IPUMS, 2025. https://doi.org/10.18128/D020.V7.7
[3] World Bank, Atlas of Sustainable Development Goals 2020: From World Development Indicators, Washington, DC: World Bank, 2020.
[4] M. Armbrust et al., "Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores," Proc. VLDB Endowment, vol. 13, no. 12, pp. 3411β3424, 2020.
[5] LIS Cross-National Data Center in Luxembourg, Luxembourg Income Study Database: Inequality and Poverty Key Figures, 1967-2020, Colchester, Essex: UK Data Service, 2022.
[6] M. Zaharia et al., "Apache Spark: A unified engine for big data processing," Commun. ACM, vol. 59, no. 11, pp. 56β65, 2016.
Built with β€οΈ by Team SDG10-ITERA | Institut Teknologi Sumatera 2026