Skip to content
View sasidhar-de's full-sized avatar
πŸ’­
Data Engineer
πŸ’­
Data Engineer

Block or report sasidhar-de

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sasidhar-de/README.md

⚑ Sasidhar | Data Engineer & Lakehouse Specialist

Welcome to my digital workspace! I specialize in engineering scalable, fault tolerant modern data stacks. I don't just move data; I build resilient pipelines that survive real world API chaos, schema drift, and enterprise scale.

πŸš€ What I Do

  • Lakehouse Engineering: Building end to end Medallion architectures (Bronze, Silver, Gold) using Databricks and Delta Lake.
  • Enterprise Governance: Implementing strict Data Quality (DQ) constraints and managing data assets natively within Unity Catalog.
  • Orchestration: Crafting optimized, parallel executing DAGs via Databricks Workflows and external orchestrators.

πŸ† Featured Projects

  • Enterprise 1TB Medallion Lakehouse: A petabyte-ready Databricks architecture processing 1TB of TPC-DS retail data. Engineered with automated CI/CD via Databricks Asset Bundles (DABs) and GitHub Actions. Features advanced Spark memory optimization (zero disk spill), SCD Type 2 tracking, and Incremental MERGE strategies for sub-second executive BI rendering.
  • Enterprise Lakehouse: An orchestrated batch processing pipeline that ingests raw JSON, applies a "Universal Bronze Shield" for schema protection, performs complex cross currency joins, and serves aggregated financial data via Databricks AI/BI Genie spaces.

πŸ’» Tech Stack

  • Core: Python, PySpark, SQL
  • Cloud & Compute: Azure, Databricks, Apache Spark
  • Storage & Formats: Azure Data Lake Storage (ADLS Gen2), Delta Lake, Parquet
  • Ecosystem: Unity Catalog, Databricks Workflows, Azure Data Factory (ADF), GitHub Actions (CI/CD)

πŸ“« Let's Connect

I am always open to discussing complex data engineering challenges, optimization strategies, or how to handle dirty API data without breaking downstream facts.

Pinned Loading

  1. Enterprise-Medallion-Lakehouse Enterprise-Medallion-Lakehouse Public

    Python 3

  2. E2E-Medillian-Architecture-LakeHouse E2E-Medillian-Architecture-LakeHouse Public

    Jupyter Notebook

  3. DataEngineeringProjects DataEngineeringProjects Public

    repo consists for data engineering projects

    Python 1

  4. Tensorflow_Advance_Techniques Tensorflow_Advance_Techniques Public

    deeplearning.ai Tensorflow advance techniques specialization

    Jupyter Notebook 73 69