Welcome to my digital workspace! I specialize in engineering scalable, fault tolerant modern data stacks. I don't just move data; I build resilient pipelines that survive real world API chaos, schema drift, and enterprise scale.
- Lakehouse Engineering: Building end to end Medallion architectures (Bronze, Silver, Gold) using Databricks and Delta Lake.
- Enterprise Governance: Implementing strict Data Quality (DQ) constraints and managing data assets natively within Unity Catalog.
- Orchestration: Crafting optimized, parallel executing DAGs via Databricks Workflows and external orchestrators.
- Enterprise 1TB Medallion Lakehouse: A petabyte-ready Databricks architecture processing 1TB of TPC-DS retail data. Engineered with automated CI/CD via Databricks Asset Bundles (DABs) and GitHub Actions. Features advanced Spark memory optimization (zero disk spill), SCD Type 2 tracking, and Incremental MERGE strategies for sub-second executive BI rendering.
- Enterprise Lakehouse: An orchestrated batch processing pipeline that ingests raw JSON, applies a "Universal Bronze Shield" for schema protection, performs complex cross currency joins, and serves aggregated financial data via Databricks AI/BI Genie spaces.
- Core: Python, PySpark, SQL
- Cloud & Compute: Azure, Databricks, Apache Spark
- Storage & Formats: Azure Data Lake Storage (ADLS Gen2), Delta Lake, Parquet
- Ecosystem: Unity Catalog, Databricks Workflows, Azure Data Factory (ADF), GitHub Actions (CI/CD)
I am always open to discussing complex data engineering challenges, optimization strategies, or how to handle dirty API data without breaking downstream facts.
- Email: sasidharpurum@gmail.com
- Location: Tirupati, India (Open to global challenges)

