An end-to-end Azure data engineering project that ingests AdventureWorks sales data, transforms it through a medallion-style data lake, and makes the resulting datasets available for Power BI reporting.
The pipeline combines Azure Data Factory, Azure Data Lake Storage, Azure Databricks, Azure Synapse Analytics, and Power BI.
| Area | Implementation |
|---|---|
| Source | AdventureWorks CSV datasets hosted in GitHub and included in Datasets/ |
| Orchestration | Azure Data Factory (ADF) using metadata-driven file mappings |
| Storage | Azure Data Lake Storage Gen2 with Bronze and Silver zones |
| Processing | Azure Databricks with PySpark |
| Serving | Azure Synapse Analytics serverless SQL and external views/tables |
| Reporting | Power BI |
| Data format | CSV at ingestion; Parquet after transformation |
| Security approach | Managed identity/database-scoped credential pattern in Synapse scripts |
The high-level and low-level designs below summarize the flow implemented in the project.
flowchart LR
A[GitHub / AdventureWorks CSV files] -->|HTTP copy| B[Azure Data Factory]
B -->|Raw files| C[(ADLS Gen2<br/>Bronze zone)]
C -->|Read CSV| D[Azure Databricks<br/>PySpark transformations]
D -->|Cleaned Parquet| E[(ADLS Gen2<br/>Silver zone)]
E --> F[Azure Synapse Analytics<br/>Serverless SQL]
F --> G[Curated SQL views<br/>and external tables]
G --> H[Power BI reports]
I[ADF-Script/git.json] -. metadata .-> B
flowchart TB
subgraph Ingestion[ADF ingestion]
M[git.json metadata]
L[ForEach dataset]
H[HTTP source using p_rel_url]
S[Sink using p_sink_folder and p_sink_file]
M --> L --> H --> S
end
subgraph Lake[ADLS Gen2]
B[(bronze/<dataset>/*.csv)]
SV[(silver/<dataset>/*.parquet)]
S --> B
end
subgraph Transform[Databricks / PySpark]
R[Read CSV]
T[Standardize types and dates<br/>Create derived columns<br/>Clean selected fields]
W[Write Parquet]
B --> R --> T --> W --> SV
end
subgraph SQL[Synapse Serverless SQL]
O[OPENROWSET over Silver Parquet]
V[gold schema views]
X[External Parquet table]
SV --> O --> V
V --> X
end
P[Power BI]
V --> P
X --> P
Create a reusable analytics pipeline for AdventureWorks that supports questions such as:
- How are sales and order volumes changing over time?
- Which products, categories, territories, and customers contribute most to performance?
- What products are being returned, and how does that relate to sales activity?
The result is a structured data foundation for BI dashboards and ad hoc analysis, while keeping ingestion, transformation, and serving responsibilities separate.
To start, the following Azure resources were provisioned:
- Azure Data Factory (ADF): Used for data orchestration and automation.
- Azure Storage Account: Acts as the data lake, storing raw (bronze), transformed (silver), and curated (gold) data.
- Azure Databricks: Performs data transformations and computations.
- Azure Synapse Analytics: Handles data warehousing for BI use.
All resources were configured with proper Identity and Access Management (IAM) roles to ensure seamless integration and security.

Azure Data Factory (ADF) serves as the backbone for orchestrating the data pipeline.
- Dynamic Copy Activity:
The raw data is now securely stored and ready for transformation.
Using Azure Databricks, the raw data from the bronze container was transformed into a structured format.
-
Cluster Setup: A Databricks cluster was created to process the data efficiently.
-
Data Lake Integration: Databricks connected to Azure Storage to access the raw data.
-
Normalized date formats for consistency.
-
Cleaned and filtered invalid or incomplete records.
-
Grouped and concatenated data to make it more usable for analysis.
-
Saved the transformed data in the silver container in Parquet format for optimal storage and query performance.
Azure Synapse Analytics structured the processed data for analysis and BI reporting.
- Connection to Silver Container: Configured Synapse to query data directly from Azure Storage.
- Serverless SQL Pools: Enabled querying without provisioning upfront resources.
- Database and Schema Creation:
The cleaned, structured data was then moved to the gold container for reporting purposes.
The final step involved integrating the data with a BI tool to visualize and generate insights.
- Power BI Integration:
This project demonstrates the power of Azure’s ecosystem in creating a robust data engineering pipeline. By combining tools like ADF, Databricks, Synapse Analytics, and Power BI, the solution achieves:
- Automation: Seamlessly moves data through different stages.
- Scalability: Handles large datasets with ease.
- Efficiency: Optimizes storage and querying with Parquet format and serverless SQL pools.
- Actionable Insights: Delivers data to stakeholders through interactive BI dashboards.
This end-to-end solution exemplifies how modern data-driven businesses can leverage Azure to transform raw data into meaningful insights, driving informed decision-making. ✅
ADF-Script/git.json— source-to-sink metadata used by the ingestion pattern.NoteBook/Data Transformations(Bronze_to_Silver).ipynb— PySpark exploration and Bronze-to-Silver transformations.Synapse SQL Scripts/Create Views Gold.sql— Synapse views over Silver Parquet datasets.Synapse SQL Scripts/Create External Table.sql— external data source, file format, credential, and table definitions.PowerBI/AdventureWorks.pbix— Power BI report artifact.Datasets/— local AdventureWorks CSV sample files.









