Welcome to my SQL data warehouse project! Feel free to explore.
This project is sourced from Data with Baraa, a YT channel that give comprehensive tutorial for projects that specializes in data. For this project, I followed the tutorial from scratch and but then created my own files and databases, thus having my own version. You may notice differences in structure and naming conventions compared to the tutorial.
This project demonstrates a comprehensive data warehousing and analytics solution, from building a data warehouse to generating actionable insights. Designed as a portfolio project, it highlights industry best practices in data engineering and analytics.
The data architecture for this project follows Medallion Architecture Bronze, Silver, and Gold layers: 
- Bronze Layer: Stores raw data as-is from the source systems. Data is ingested from CSV Files into SQL Server Database.
- Silver Layer: This layer includes data cleansing, standardization, and normalization processes to prepare data for analysis.
- Gold Layer: Houses business-ready data modeled into a star schema required for reporting and analytics.
This project involves:
- Data Architecture: Designing a Modern Data Warehouse Using Medallion Architecture Bronze, Silver, and Gold layers.
- ETL Pipelines: Extracting, transforming, and loading data from source systems into the warehouse.
- Data Modeling: Developing fact and dimension tables optimized for analytical queries.
- Analytics & Reporting: Creating SQL-based reports and dashboards for actionable insights.
🎯 This repository showcases expertise in:
- SQL Development
- Data Architect
- Data Engineering
- ETL Pipeline Development
- Data Modeling
- Data Analytics
Develop a modern data warehouse using SQL Server to consolidate sales data, enabling analytical reporting and informed decision-making.
- Data Sources: Import data from two source systems (ERP and CRM) provided as CSV files.
- Data Quality: Cleanse and resolve data quality issues prior to analysis.
- Integration: Combine both sources into a single, user-friendly data model designed for analytical queries.
- Scope: Focus on the latest dataset only; historization of data is not required.
- Documentation: Provide clear documentation of the data model to support both business stakeholders and analytics teams.
Develop SQL-based analytics to deliver detailed insights into:
- Customer Behavior
- Product Performance
- Sales Trends These insights empower stakeholders with key business metrics, enabling strategic decision-making.
data-warehouse-project/
│
├── datasets/ # Raw datasets used for the project (ERP and CRM data)
│
├── documents/ # Project documentation and architecture details
│ ├── sql project diagram.drawio # Main Draw.io file shows the project's architecture
│ ├── HIGH LEVEL ARCHITECHTURE.png # Image shows the project's architecture
│ ├── SALES DATA MART.png # Data Mart Architechture
│ ├── Data Flow Diagram.png # Image for the data flow diagram
│ ├── INTEGRATION MODEL.png # Image file for data models (star schema)
│
│
├── scripts/ # SQL scripts for ETL and transformations
│ ├── bronze/ # Scripts for extracting and loading raw data
│ ├── silver/ # Scripts for cleaning and transforming data
│ ├── gold/ # Scripts for creating analytical models
│
├── tests/ # Test scripts and quality files
│
├── README.md # Project overview and instructions
├── LICENSE # License information for the repository
└── requirements.txt # Dependencies and requirements for the project
This project is licensed under the MIT License. You are free to use, modify, and share this project with proper attribution.