End-to-end Azure Data Engineering project for Healthcare Revenue Cycle Management
-
Updated
Apr 10, 2026 - Python
End-to-end Azure Data Engineering project for Healthcare Revenue Cycle Management
A Python package for Azure Datalake Storage adlsgen2 [abfss://] and Microsoft Fabric Lakehouse [abfss://], Google Cloud Storage [gs://bucket], AWS S3 bucket [s3://bucket] enables format detection and schema retrieval for Iceberg, Delta, and Parquet formats. It helps identify partitioned columns for parquet datasets. It also supports querying Delta,
ETL project with Spark and Airflow
End-to-end Azure Data Engineering Pipeline using Azure Data Factory, Azure Databricks, ADLS Gen2, Medallion Architecture, Metadata-Driven ETL, PySpark, SQL and Power BI.
Implemented Azure Databricks for real-time data processing and governance using Unity Catalog, Spark Structured Streaming, Delta Lake features, Medallion Architecture, and end-to-end CI/CD pipelines. Focused on incremental loading, compute cluster management, maintaining data quality, and creating workflows.
🚀 Production ETL pipeline with Apache Airflow, Spark & Azure Data Lake
High-Throughput PyTorch Sequential Data Loaders for GPU Starvation Reduction
Databricks medallion architecture pipeline for NFL Big Data Bowl 2026 prediction using PySpark, Delta Lake, SparkML, and Azure ADLS Gen2
"Explore Formula 1 data analytics with this project. Leveraging the Ergast API, it utilizes Databricks Spark for ingestion, transformation, and analysis. ADLS acts as the storage layer, while Power BI visualizes the ADLS presentation layer. Uncover insights in the world of Formula 1 through powerful data analytics."
Contract-first Azure batch data product using Synapse Spark with deterministic recompute guarantees and audit-style run evidence.
This project demonstrates a comprehensive data engineering pipeline capable of transforming vast amounts of historical F1 data into valuable insights. By leveraging modern cloud-based technologies and automating key processes, the project is designed to be both scalable and adaptable to future requirements.
This is an End to End Azure Data Engineering project copying data from Rest API to Azure cloud.
End-to-end Azure Data Engineering pipeline using Databricks, PySpark, and Azure Data Factory with Medallion Architecture, incremental processing, and Delta Lake MERGE.
AirBnB CDC Ingestion Pipeline: Near Real-Time Change Data Capture (CDC) Pipeline on Azure for Seamless Integration of Continuous Data Streams
Add a description, image, and links to the adlsgen2 topic page so that developers can more easily learn about it.
To associate your repository with the adlsgen2 topic, visit your repo's landing page and select "manage topics."