Aspiring Machine Learning Engineer, currently building expertise through data analytics and engineering projects.
Recent CS grad (Data Science concentration) seeking Data Analyst / Junior Data Engineer roles, building toward a long-term goal in machine learning.
End-to-end data pipeline that generates synthetic healthcare claims data, validates records, loads them into PostgreSQL, and produces analytics-ready reports.
Concepts
- ETL pipeline design
- SQL analytics views
- data validation & reject handling
- indexing and query optimization
- automated pipeline execution
https://github.com/cybereptilia/healthcare-claims-pipeline
Time-series database experiment (~600k rows) demonstrating PostgreSQL query planner behavior and index selectivity.
Concepts
- composite index design
- cost-based query optimization
- sequential vs index scan behavior
- EXPLAIN ANALYZE benchmarking
https://github.com/cybereptilia/predictive-maintenance-sql
- MedTracker Frontend — React + Tailwind integration for Front-end
- ML Study Series — Implementing algorithms from Hands-On Machine Learning with Scikit-Learn, Keras & TensorFlow
- Data Visualization Projects — Interactive dashboards using Matplotlib, Seaborn, and Plotly
Manual Softmax classifier using gradient descent and cross-entropy optimization achieving 96.67% accuracy.
Complete regression pipeline including EDA, feature engineering, model training, and evaluation.
Binary classification evaluation including precision-recall tradeoffs and ROC analysis (AUC ≈ 0.99).
Languages
Python • SQL • JavaScript • C++ • Java
Libraries & ML Frameworks
NumPy • Pandas • Scikit-Learn • TensorFlow • Keras • Matplotlib • Seaborn
Data Tools
PostgreSQL • JupyterLab • Google Colab • Excel
Workflow
Git • GitHub • Virtual Environments
Operating Systems
Rocky Linux • Ubuntu • Windows





