I turn messy, multi-source data into pipelines and dashboards that people actually trust and use. I design for the 3am pager alert, not just the demo — with tested SCD2 dimensional models, automated data quality gates, and disaster-recovery runbooks — while keeping the end report or dashboard the whole point of the pipeline, not an afterthought.
Currently seeking Data Engineer / Analytics Engineer roles. 3 production-grade pipelines shipped with CI, 26+ automated tests, and dbt/Airflow/Kafka experience across both cloud and local-hybrid stacks.
Data Engineering & Analytics Intern National Telecommunication Institute (NTI) · April 2026 – July 2026
- Designed and shipped 3 interactive Power BI dashboards tracking customer-segment KPIs and telecom performance, adopted for weekly use by operations and management teams
- Analyzed 100K+ records across multi-source retail and telecom data (SQL, Pandas, NumPy) to surface behavioral patterns that shaped the current quarter's campaign targeting strategy
- Ran A/B testing frameworks to measure segment lift — results were adopted as the marketing team's campaign baseline
- Automated recurring KPI reporting (Python + Power Query/PivotTables), cutting manual reporting time by 4+ hours per cycle
| 🔧 Pipeline Engineering | Airflow, Kafka, and PySpark across batch and real-time workloads |
| 🏛️ Dimensional Modeling | Star schema, SCD Type 1/2, in SQL Server, Snowflake, and DuckDB |
| ✅ Data Quality & Testing | dbt tests, automated DQ frameworks that fail builds on critical issues |
| 📊 The Last Mile | Power BI dashboards and DAX measures business users actually rely on |
| 🛡️ Production Reliability | Documented SLOs, tested DR runbooks, access-control design |
| Category | Tools |
|---|---|
| Orchestration & Streaming | Apache Airflow · Apache Kafka · Debezium (CDC) |
| Processing & Transformation | PySpark · dbt · Python (Pandas, NumPy) |
| Warehousing & Modeling | SQL Server · Snowflake · DuckDB · Star Schema · SCD Type 1 & 2 · Medallion Architecture |
| BI & Reporting | Power BI · DAX · SSAS (OLAP Cubes) · Excel |
| Infrastructure & Monitoring | Docker · Grafana |
| Languages | Python · T-SQL · DAX |
Airflow · Kafka · dbt · PySpark · Databricks · Snowflake
- Hybrid platform: identical code runs locally (DuckDB/PySpark) or in the cloud (Databricks/ADLS/Snowflake) — one environment variable switches modes
- Real-time Kafka streaming + CDC replication feeding a dbt-built Medallion (Bronze/Silver/Gold) dimensional model
- Shipped with SLOs, a tested disaster-recovery runbook, and a PAN-tokenization security design — the ops layer most portfolio projects skip
Airflow · PySpark · PostgreSQL · Snowflake
- Bronze/Silver/Gold pipeline with watermark-based incremental extraction and SCD Type 2 history
- Automated data quality framework that fails the build on critical checks, not just logs them
- 26 unit + integration tests running in GitHub Actions CI, with bronze load and DQ checks on a live Airflow schedule
Snowflake · dbt · Airflow · Kafka
- Real-time CDC ingestion and Kafka streaming pipelines feeding dbt-modeled, BI-ready marts
Power BI · DAX · Power Query
- Full BI workflow — data cleaning, star-schema modeling, 11+ custom DAX measures — into an interactive 3-page dashboard covering attrition, attendance, and performance
More projects — Sales Analytics DW · Hotel Booking DW (SSIS/SSAS, SCD2, OLAP cubes)
SSIS · SSAS · SQL Server · Power BI — Star-schema warehouse with SCD Type 2 tracking and a multidimensional OLAP cube powering enterprise-wide sales reporting.
SSIS · SSAS · Power BI — End-to-end BI pipeline with SCD Type 2 historical tracking, OLAP cube modeling, and automated ETL for booking and occupancy analytics.
- Data Engineer in Python — DataCamp
- Google Data Analytics Professional Certificate
- IBM Data Science Professional Certificate
- Data Analyst in Power BI — DataCamp
- Data Analyst in Python — DataCamp