Skip to content

About

A Machine Learning system that detects credit card fraud in real time using Random Forest, Neural Networks and Streamlit web app with 99.99% accuracy.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

ย 

History

97 Commits

Folders and files

NameName
Last commit message
Last commit date
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 
ย 

Repository files navigation

๐Ÿ›ก๏ธ Fraud Monitor โ€” FinWise Secure

Real-time credit card fraud detection powered by machine learning.

Python FastAPI scikit-learn JavaScript Render Vercel License Status

Live Demo ยท API ยท Report Bug ยท Request Feature


๐Ÿ“‘ Table of Contents


๐Ÿ“Œ About

Fraud Monitor is a real-time credit card fraud detection system that classifies transactions as Legitimate or Fraudulent using a machine learning model trained on anonymized, PCA-transformed transaction data.

Problem it solves: Credit card fraud costs billions annually, and manual transaction review doesn't scale. This project demonstrates an end-to-end ML pipeline โ€” from a trained classification model to a live, interactive risk-scoring dashboard โ€” that can flag suspicious transactions instantly based on transaction amount, timing, and behavioral features.

Who it's for:

  • Recruiters and engineers reviewing an end-to-end applied ML project
  • Developers learning how to take a trained model from notebook โ†’ production API โ†’ deployed frontend
  • Anyone exploring fraud detection techniques on the classic anonymized credit card dataset

Project type: Full-Stack Machine Learning Application Status: ๐Ÿšง In Development


โœจ Features

  • ๐Ÿ” Real-time transaction analysis โ€” submit a transaction and get an instant classification
  • โšก Quick-load sample transactions โ€” test the model against known transaction profiles without manual entry
  • ๐Ÿ“Š Risk scoring readout โ€” Fraud Risk % and Legitimate Score % displayed as visual meters
  • ๐Ÿšฆ Risk level tagging โ€” transactions are labeled Low / Medium / High risk
  • ๐Ÿง  ML-backed classification โ€” Random Forest model trained on PCA-transformed features (V1โ€“V28), Amount, and Time
  • ๐ŸŒ REST API backend โ€” FastAPI service exposing a prediction endpoint, decoupled from the frontend
  • ๐Ÿ–ฅ๏ธ Custom dashboard UI โ€” dark-themed, purpose-built HTML/CSS/JS interface (no framework overhead)
  • ๐Ÿ›Ÿ Graceful degradation โ€” if the trained model isn't available, the API falls back to a heuristic estimate rather than crashing
  • โ˜๏ธ Split cloud deployment โ€” frontend on Vercel, backend on Render, model artifacts served via GitHub Releases

๐Ÿ› ๏ธ Tech Stack

Frontend:

  • HTML5, CSS3, Vanilla JavaScript (custom dashboard, no framework)

Backend:

  • Python 3.13
  • FastAPI
  • Uvicorn (ASGI server)

Machine Learning:

  • scikit-learn (Random Forest Classifier)
  • pandas, NumPy
  • pickle (model serialization)

Deployment:

  • Backend โ†’ Render
  • Frontend โ†’ Vercel
  • Model artifacts (model.pkl, scalers.pkl) โ†’ hosted as GitHub Release assets, downloaded at server startup

Dataset:

  • Kaggle โ€” Credit Card Fraud Detection dataset (anonymized, PCA-transformed features V1โ€“V28)

๐Ÿ“‚ Folder Structure

Machine-Learning-Project/
โ”œโ”€โ”€ backend/
โ”‚   โ”œโ”€โ”€ main.py              # FastAPI app entry point, model loading, prediction endpoint
โ”‚   โ”œโ”€โ”€ fraud_detection.py   # Model training pipeline (Random Forest + scalers)
โ”‚   โ”œโ”€โ”€ requirements.txt     # Backend Python dependencies
โ”‚   โ””โ”€โ”€ (model.pkl, scalers.pkl generated locally / downloaded at runtime โ€” not committed)
โ”œโ”€โ”€ frontend/
โ”‚   โ”œโ”€โ”€ index.html           # Dashboard UI
โ”‚   โ”œโ”€โ”€ style.css
โ”‚   โ””โ”€โ”€ script.js            # Calls backend API, renders analysis readout
โ”œโ”€โ”€ README.md
โ””โ”€โ”€ .gitattributes           # Ensures binary files (e.g. .pkl) aren't corrupted by line-ending conversion

Note: model.pkl, scalers.pkl, and the training dataset (creditcard.csv) are intentionally not committed to Git due to file size (100MB+ limits on standard GitHub repos). See Machine Learning below for how they're provisioned.


โš™๏ธ Installation

Prerequisites

  • Python 3.10+
  • pip
  • Git

1. Clone the repository

git clone https://github.com/slashthose/Machine-Learning-Project.git
cd Machine-Learning-Project

2. Set up a virtual environment

python -m venv venv
# Windows
venv\Scripts\activate
# macOS/Linux
source venv/bin/activate

3. Install backend dependencies

cd backend
pip install -r requirements.txt

4. Obtain the trained model

The model isn't stored in this repo. Either:

  • Download it automatically โ€” the backend fetches model.pkl and scalers.pkl from the project's GitHub Release the first time it starts (see load_artifacts() in main.py), or
  • Train it yourself โ€” download the Kaggle Credit Card Fraud dataset, place creditcard.csv in backend/, and run:
python fraud_detection.py

5. Run the backend

uvicorn main:app --reload --port 8000

API will be available at http://localhost:8000.

6. Run the frontend

Open frontend/index.html directly in a browser, or serve it locally:

cd frontend
python -m http.server 5500

Then visit http://localhost:5500. Update the API base URL in script.js to point to your local backend (http://localhost:8000) if testing locally.


๐Ÿ” Environment Variables

Create a .env file inside backend/ (not committed to Git):

MODEL_URL=https://github.com/slashthose/Machine-Learning-Project/releases/download/v1.0-model/model.pkl
SCALERS_URL=https://github.com/slashthose/Machine-Learning-Project/releases/download/v1.0-model/scalers.pkl
MODEL_PATH=model.pkl
SCALERS_PATH=scalers.pkl
PORT=8000

No API keys or secrets are currently required โ€” this project doesn't use external paid APIs or authentication.


๐Ÿš€ Usage

  1. Visit the live dashboard.
  2. Either:
    • Select a known sample transaction from the dropdown, or
    • Manually enter an Amount and Time (seconds since first transaction).
  3. Click Analyze Transaction.
  4. The Analysis Readout panel returns:
    • Classification: LEGITIMATE or FRAUDULENT
    • Risk level: LOW RISK / MEDIUM RISK / HIGH RISK
    • Fraud Risk % and Legitimate Score % (visual meters)
    • A recommended action (e.g., "Approve and process the transaction" or "Flag for manual review")

๐Ÿ–ผ๏ธ Screenshots

Add screenshots to a /screenshots folder and update the paths below.

Dashboard Analysis Readout
Dashboard Analysis Readout

๐Ÿ“ก API Documentation

POST /predict

Analyzes a transaction and returns a fraud classification.

Request Body:

{
  "amount": 150.00,
  "time": 50000,
  "features": [0.0, 0.0, "...", 0.0]
}

(features = V1โ€“V28 PCA-transformed values; defaults to zeroed values for manual entries not sourced from a known sample)

Response:

{
  "prediction": "legitimate",
  "fraud_risk_percent": 0,
  "legitimate_score_percent": 100,
  "risk_level": "low",
  "recommendation": "Approve and process the transaction.",
  "using_fallback_model": false
}

Status Codes:

Code Meaning
200 Prediction successful
422 Invalid request payload
500 Internal server / model loading error

Exact field names may differ slightly depending on your current main.py implementation โ€” update this section to match your actual schema before publishing.


๐Ÿค– Machine Learning

  • Model: Random Forest Classifier (scikit-learn), selected as the best-performing model after evaluation against alternative classifiers
  • Dataset: Kaggle Credit Card Fraud Detection dataset โ€” anonymized transactions with PCA-transformed features V1โ€“V28, plus Amount and Time
  • Preprocessing: Amount and Time are scaled separately using dedicated scalers (scaler_time, scaler_amount) before being passed to the model
  • Training: Handled in fraud_detection.py โ€” includes data loading, scaling, train/test split, model fitting, and evaluation
  • Inference: The FastAPI backend loads the serialized model and scalers via pickle and returns a classification + probability score per request
  • Evaluation Metric: ~99.99% accuracy on the held-out test set (severely imbalanced dataset โ€” precision/recall/F1 and confusion matrix are more meaningful than raw accuracy given the extreme class imbalance in fraud data)
  • Libraries: scikit-learn, pandas, NumPy, pickle

Model & Scaler Provisioning: Because the trained model (400โ€“600MB) and dataset exceed GitHub's standard file size limits, they are not committed to this repository. Instead:

  • model.pkl and scalers.pkl are published as assets on a GitHub Release
  • On startup, main.py's load_artifacts() downloads them automatically if not already present locally
  • If the files fail to download, the API falls back to a simple heuristic estimate rather than crashing, and flags the response accordingly

โšก Performance

  • Inference time: Sub-second per transaction (single Random Forest prediction, no batch overhead)
  • Scalability: Stateless FastAPI service โ€” horizontally scalable behind a load balancer if needed
  • Cold starts: On Render's free tier, the first request after inactivity may be slower due to model download + service spin-up
  • Optimization opportunities: Model size (400โ€“600MB) is large for a Random Forest and is a candidate for compression via reduced n_estimators/max_depth, which would also speed up cold-start downloads

๐Ÿง— Challenges Faced

  • Binary file corruption via Git line-ending conversion โ€” the model .pkl file was corrupted in transit due to Git's autocrlf normalizing line endings in a binary file, causing _pickle.UnpicklingError on deploy
  • GitHub's 100MB file size limit โ€” the trained model (400โ€“600MB) couldn't be committed directly, requiring a rethink of the deployment/model-loading architecture
  • Decoupling training from serving โ€” ensuring the API can start reliably even when the model artifact isn't bundled with the code, without silently masking failures
  • Balancing graceful degradation with correctness โ€” the app intentionally falls back to a heuristic rather than crashing when the model is unavailable, which improves uptime but requires clear signaling to the frontend/user that predictions aren't from the real model

๐Ÿ”ฎ Future Improvements

  1. Add authentication for the API (JWT-based)
  2. Add rate limiting to the /predict endpoint
  3. Log all predictions to a database for audit/history tracking
  4. Add a batch-prediction endpoint for CSV uploads
  5. Build a model retraining pipeline triggered on new labeled data
  6. Add SHAP/feature-importance explainability to the analysis readout
  7. Compress the model (fewer estimators / max depth tuning) to speed up cold starts
  8. Add automated CI/CD (GitHub Actions) for testing before deploy
  9. Add unit tests for the FastAPI endpoints
  10. Add integration tests covering the fallback heuristic path
  11. Add a Dockerfile for consistent local/prod parity
  12. Add real-time transaction streaming (WebSocket) demo mode
  13. Add user accounts to save transaction history
  14. Add multi-model comparison (e.g., XGBoost vs Random Forest vs Neural Net) in the readout
  15. Add dark/light theme toggle to the dashboard
  16. Add mobile-responsive layout improvements
  17. Add model versioning so old model artifacts can be rolled back to

๐Ÿงช Testing

  • Manual testing: Verified via the live dashboard using both known sample transactions and manual input across varying amount/time ranges
  • Edge cases to cover: zero/negative amounts, extreme time values, missing PCA features, malformed request payloads
  • Unit testing: Not yet implemented โ€” recommended: pytest + httpx.AsyncClient for FastAPI endpoint testing
  • Security testing: Not yet implemented โ€” recommended: basic input validation fuzzing and dependency vulnerability scanning (pip-audit)
  • Performance testing: Not yet implemented โ€” recommended: load testing the /predict endpoint with locust or k6

โ˜๏ธ Deployment

Component Platform Notes
Backend (FastAPI) Render Auto-deploys from main branch; downloads model artifacts from GitHub Releases at startup
Frontend (Dashboard) Vercel Static site deploy, auto-deploys from main branch
Model Artifacts GitHub Releases Hosted as binary release assets to bypass Git's 100MB file limit

To deploy your own copy:

  1. Fork this repo
  2. Create a GitHub Release and attach your own model.pkl / scalers.pkl
  3. Connect the backend/ folder as a new Web Service on Render, set the start command to uvicorn main:app --host 0.0.0.0 --port $PORT
  4. Connect the frontend/ folder as a new Vercel project
  5. Update MODEL_URL / SCALERS_URL in your backend environment variables to point to your release assets

๐Ÿค Contributing

Contributions are welcome!

  1. Fork the repository
  2. Create a feature branch: git checkout -b feature/your-feature-name
  3. Commit your changes: git commit -m "Add your feature"
  4. Push to your branch: git push origin feature/your-feature-name
  5. Open a Pull Request describing your changes

Please open an issue first for significant changes so we can discuss the approach.


๐Ÿ“„ License

Distributed under the MIT License. See LICENSE for more information.


๐Ÿ™ Acknowledgements


Built by Sakshi

โญ Star this repo if you found it useful!

About

A Machine Learning system that detects credit card fraud in real time using Random Forest, Neural Networks and Streamlit web app with 99.99% accuracy.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages