Data Engineer | Data Analyst | Databricks Certified Data Engineer Associate | SQL • Python • PySpark • Databricks • Tableau
Recent MS Computer Science graduate from the University of Southern California, focused on building reliable data pipelines, analytics-ready datasets, and decision-support solutions.
I enjoy turning messy raw data into clean, validated, structured datasets that support business decisions, reporting, BI, machine learning, and operational workflows.
| Area | What I’m Building / Learning | Tools |
|---|---|---|
| Data Engineering | End-to-end lakehouse pipelines, ETL/ELT workflows, data validation, incremental ingestion, and medallion architecture | Databricks, PySpark, Spark SQL, Delta Lake |
| Analytics & BI | Dashboard-ready datasets, KPI reporting, business analysis, customer segmentation, and data storytelling | SQL, Tableau, Power BI, Excel |
| Machine Learning | Classification, feature engineering, model evaluation, risk scoring, and decision-support systems | Python, scikit-learn, pandas |
| Cloud & AI Certifications | Preparing for AWS Certified Cloud Practitioner and AWS Certified AI Practitioner certifications | AWS Skill Builder, Cloud Fundamentals, Generative AI, Machine Learning |
| Continuous Learning | Expanding my Databricks, Lakeflow, Unity Catalog, and cloud data engineering skills | Databricks Academy, Lakeflow, Unity Catalog |
| Project | Description | Tech Stack |
|---|---|---|
| HealthPulse - Smart Patient Monitoring and Readmission Analytics | Built an end-to-end healthcare analytics platform in Databricks using Bronze, Silver, and Gold architecture. Processed patient, admission, condition, medication, lab, and vital-sign data to create Patient 360 profiles, hospital operations metrics, monitoring alerts, and readmission-risk datasets. Developed classification models, explainable risk factors, audit logging, validation checks, watermark tracking, and incremental ingestion workflows. | Python, PySpark, Spark SQL, Databricks, Delta Lake, Machine Learning |
| RetentionIQ - Telecom Churn and Retention Analytics | Built an end-to-end retention decision-support solution that scores customers by churn risk, customer value, and revenue exposure. Developed a recommendation engine, campaign prioritization workflow, SQLite analytics layer, budget simulation, and interactive Tableau dashboard. | Python, SQL, scikit-learn, SQLite, Tableau, pandas |
| U.S. EPA Emissions Analytics Dashboard | Built Databricks SQL workflows to clean, transform, and analyze EPA emissions data, including state and county-level metrics and emissions-per-person calculations. | Databricks SQL, SQL, Data Cleaning, Data Visualization |
| DuckDB Zonemap Optimization | Extended DuckDB query execution by integrating column imprint bitmaps with zonemap pruning to improve predicate filtering efficiency. | C++, DuckDB, SQL, Query Optimization |
| Cropable - AI Crop Disease Detection | Published research project on an AI-powered crop disease detection system using CNN and VGG16-based image classification to identify tomato leaf diseases and recommend treatment measures. | Python, CNN, VGG16, Image Processing, Machine Learning |
| Category | Tools / Skills |
|---|---|
| Programming & Querying | Python, SQL, PySpark, Spark SQL |
| Data Engineering | ETL/ELT, batch processing, incremental ingestion, data transformation, data validation, pipeline orchestration, watermarking, schema enforcement |
| Big Data & Lakehouse | Apache Spark, Databricks, Delta Lake, Delta tables, Unity Catalog, Auto Loader, medallion architecture |
| Analytics & Business Intelligence | Exploratory data analysis, KPI development, dashboard design, business reporting, customer segmentation, trend analysis, data storytelling |
| Machine Learning | scikit-learn, classification, feature engineering, model evaluation, risk scoring, predictive analytics |
| Databases & Data Modeling | MySQL, MariaDB, SQLite, Databricks SQL, relational modeling, dimensional modeling, joins, window functions |
| BI & Visualization | Tableau, Power BI, Excel, interactive dashboards, calculated fields, filters, executive reporting |
| Data Quality & Governance | Data profiling, missing-value handling, duplicate detection, validation rules, audit logging, data lineage, access governance |
| Developer Tools | Git, GitHub, Jupyter Notebook, VS Code, virtual environments, CI/CD fundamentals |
| Cloud & Platforms | AWS fundamentals, S3 concepts, cloud data services, Databricks platform |
🎓 University of Southern California
M.S. in Computer Science
🎓 Vellore Institute of Technology
B.Tech in Computer Science and Engineering
- Databricks Certified Data Engineer Associate — Issued July 18, 2026
- Preparing for AWS Certified Cloud Practitioner
- Preparing for AWS Certified AI Practitioner
- Knowledge Badge - Data Ingestion with Lakeflow Connect
- Knowledge Badge - Deploy Workloads with Lakeflow Jobs
- Knowledge Badge - Build Data Pipelines with Lakeflow Spark Declarative Pipelines
- Knowledge Badge - DevOps Essentials for Data Engineering
- Knowledge Badge - SQL Analytics on Databricks
- Knowledge Badge - AI/BI for Data Analysts
⭐ Open to Data Engineer, Data Analyst, Analytics Engineer, Business Intelligence Analyst, and related data roles.