Real-time credit card fraud detection powered by machine learning.
Live Demo ยท API ยท Report Bug ยท Request Feature
- About
- Features
- Tech Stack
- Folder Structure
- Installation
- Environment Variables
- Usage
- Screenshots
- API Documentation
- Machine Learning
- Performance
- Challenges Faced
- Future Improvements
- Testing
- Deployment
- Contributing
- License
- Acknowledgements
Fraud Monitor is a real-time credit card fraud detection system that classifies transactions as Legitimate or Fraudulent using a machine learning model trained on anonymized, PCA-transformed transaction data.
Problem it solves: Credit card fraud costs billions annually, and manual transaction review doesn't scale. This project demonstrates an end-to-end ML pipeline โ from a trained classification model to a live, interactive risk-scoring dashboard โ that can flag suspicious transactions instantly based on transaction amount, timing, and behavioral features.
Who it's for:
- Recruiters and engineers reviewing an end-to-end applied ML project
- Developers learning how to take a trained model from notebook โ production API โ deployed frontend
- Anyone exploring fraud detection techniques on the classic anonymized credit card dataset
Project type: Full-Stack Machine Learning Application Status: ๐ง In Development
- ๐ Real-time transaction analysis โ submit a transaction and get an instant classification
- โก Quick-load sample transactions โ test the model against known transaction profiles without manual entry
- ๐ Risk scoring readout โ Fraud Risk % and Legitimate Score % displayed as visual meters
- ๐ฆ Risk level tagging โ transactions are labeled Low / Medium / High risk
- ๐ง ML-backed classification โ Random Forest model trained on PCA-transformed features (V1โV28), Amount, and Time
- ๐ REST API backend โ FastAPI service exposing a prediction endpoint, decoupled from the frontend
- ๐ฅ๏ธ Custom dashboard UI โ dark-themed, purpose-built HTML/CSS/JS interface (no framework overhead)
- ๐ Graceful degradation โ if the trained model isn't available, the API falls back to a heuristic estimate rather than crashing
- โ๏ธ Split cloud deployment โ frontend on Vercel, backend on Render, model artifacts served via GitHub Releases
Frontend:
- HTML5, CSS3, Vanilla JavaScript (custom dashboard, no framework)
Backend:
- Python 3.13
- FastAPI
- Uvicorn (ASGI server)
Machine Learning:
- scikit-learn (Random Forest Classifier)
- pandas, NumPy
- pickle (model serialization)
Deployment:
- Backend โ Render
- Frontend โ Vercel
- Model artifacts (
model.pkl,scalers.pkl) โ hosted as GitHub Release assets, downloaded at server startup
Dataset:
- Kaggle โ Credit Card Fraud Detection dataset (anonymized, PCA-transformed features V1โV28)
Machine-Learning-Project/
โโโ backend/
โ โโโ main.py # FastAPI app entry point, model loading, prediction endpoint
โ โโโ fraud_detection.py # Model training pipeline (Random Forest + scalers)
โ โโโ requirements.txt # Backend Python dependencies
โ โโโ (model.pkl, scalers.pkl generated locally / downloaded at runtime โ not committed)
โโโ frontend/
โ โโโ index.html # Dashboard UI
โ โโโ style.css
โ โโโ script.js # Calls backend API, renders analysis readout
โโโ README.md
โโโ .gitattributes # Ensures binary files (e.g. .pkl) aren't corrupted by line-ending conversion
Note:
model.pkl,scalers.pkl, and the training dataset (creditcard.csv) are intentionally not committed to Git due to file size (100MB+ limits on standard GitHub repos). See Machine Learning below for how they're provisioned.
- Python 3.10+
- pip
- Git
git clone https://github.com/slashthose/Machine-Learning-Project.git
cd Machine-Learning-Projectpython -m venv venv
# Windows
venv\Scripts\activate
# macOS/Linux
source venv/bin/activatecd backend
pip install -r requirements.txtThe model isn't stored in this repo. Either:
- Download it automatically โ the backend fetches
model.pklandscalers.pklfrom the project's GitHub Release the first time it starts (seeload_artifacts()inmain.py), or - Train it yourself โ download the Kaggle Credit Card Fraud dataset, place
creditcard.csvinbackend/, and run:
python fraud_detection.pyuvicorn main:app --reload --port 8000API will be available at http://localhost:8000.
Open frontend/index.html directly in a browser, or serve it locally:
cd frontend
python -m http.server 5500Then visit http://localhost:5500. Update the API base URL in script.js to point to your local backend (http://localhost:8000) if testing locally.
Create a .env file inside backend/ (not committed to Git):
MODEL_URL=https://github.com/slashthose/Machine-Learning-Project/releases/download/v1.0-model/model.pkl
SCALERS_URL=https://github.com/slashthose/Machine-Learning-Project/releases/download/v1.0-model/scalers.pkl
MODEL_PATH=model.pkl
SCALERS_PATH=scalers.pkl
PORT=8000No API keys or secrets are currently required โ this project doesn't use external paid APIs or authentication.
- Visit the live dashboard.
- Either:
- Select a known sample transaction from the dropdown, or
- Manually enter an Amount and Time (seconds since first transaction).
- Click Analyze Transaction.
- The Analysis Readout panel returns:
- Classification:
LEGITIMATEorFRAUDULENT - Risk level:
LOW RISK/MEDIUM RISK/HIGH RISK - Fraud Risk % and Legitimate Score % (visual meters)
- A recommended action (e.g., "Approve and process the transaction" or "Flag for manual review")
- Classification:
Add screenshots to a
/screenshotsfolder and update the paths below.
| Dashboard | Analysis Readout |
|---|---|
![]() |
![]() |
Analyzes a transaction and returns a fraud classification.
Request Body:
{
"amount": 150.00,
"time": 50000,
"features": [0.0, 0.0, "...", 0.0]
}(features = V1โV28 PCA-transformed values; defaults to zeroed values for manual entries not sourced from a known sample)
Response:
{
"prediction": "legitimate",
"fraud_risk_percent": 0,
"legitimate_score_percent": 100,
"risk_level": "low",
"recommendation": "Approve and process the transaction.",
"using_fallback_model": false
}Status Codes:
| Code | Meaning |
|---|---|
200 |
Prediction successful |
422 |
Invalid request payload |
500 |
Internal server / model loading error |
Exact field names may differ slightly depending on your current
main.pyimplementation โ update this section to match your actual schema before publishing.
- Model: Random Forest Classifier (scikit-learn), selected as the best-performing model after evaluation against alternative classifiers
- Dataset: Kaggle Credit Card Fraud Detection dataset โ anonymized transactions with PCA-transformed features
V1โV28, plusAmountandTime - Preprocessing:
AmountandTimeare scaled separately using dedicated scalers (scaler_time,scaler_amount) before being passed to the model - Training: Handled in
fraud_detection.pyโ includes data loading, scaling, train/test split, model fitting, and evaluation - Inference: The FastAPI backend loads the serialized model and scalers via
pickleand returns a classification + probability score per request - Evaluation Metric: ~99.99% accuracy on the held-out test set (severely imbalanced dataset โ precision/recall/F1 and confusion matrix are more meaningful than raw accuracy given the extreme class imbalance in fraud data)
- Libraries: scikit-learn, pandas, NumPy, pickle
Model & Scaler Provisioning: Because the trained model (400โ600MB) and dataset exceed GitHub's standard file size limits, they are not committed to this repository. Instead:
model.pklandscalers.pklare published as assets on a GitHub Release- On startup,
main.py'sload_artifacts()downloads them automatically if not already present locally - If the files fail to download, the API falls back to a simple heuristic estimate rather than crashing, and flags the response accordingly
- Inference time: Sub-second per transaction (single Random Forest prediction, no batch overhead)
- Scalability: Stateless FastAPI service โ horizontally scalable behind a load balancer if needed
- Cold starts: On Render's free tier, the first request after inactivity may be slower due to model download + service spin-up
- Optimization opportunities: Model size (400โ600MB) is large for a Random Forest and is a candidate for compression via reduced
n_estimators/max_depth, which would also speed up cold-start downloads
- Binary file corruption via Git line-ending conversion โ the model
.pklfile was corrupted in transit due to Git'sautocrlfnormalizing line endings in a binary file, causing_pickle.UnpicklingErroron deploy - GitHub's 100MB file size limit โ the trained model (400โ600MB) couldn't be committed directly, requiring a rethink of the deployment/model-loading architecture
- Decoupling training from serving โ ensuring the API can start reliably even when the model artifact isn't bundled with the code, without silently masking failures
- Balancing graceful degradation with correctness โ the app intentionally falls back to a heuristic rather than crashing when the model is unavailable, which improves uptime but requires clear signaling to the frontend/user that predictions aren't from the real model
- Add authentication for the API (JWT-based)
- Add rate limiting to the
/predictendpoint - Log all predictions to a database for audit/history tracking
- Add a batch-prediction endpoint for CSV uploads
- Build a model retraining pipeline triggered on new labeled data
- Add SHAP/feature-importance explainability to the analysis readout
- Compress the model (fewer estimators / max depth tuning) to speed up cold starts
- Add automated CI/CD (GitHub Actions) for testing before deploy
- Add unit tests for the FastAPI endpoints
- Add integration tests covering the fallback heuristic path
- Add a Dockerfile for consistent local/prod parity
- Add real-time transaction streaming (WebSocket) demo mode
- Add user accounts to save transaction history
- Add multi-model comparison (e.g., XGBoost vs Random Forest vs Neural Net) in the readout
- Add dark/light theme toggle to the dashboard
- Add mobile-responsive layout improvements
- Add model versioning so old model artifacts can be rolled back to
- Manual testing: Verified via the live dashboard using both known sample transactions and manual input across varying amount/time ranges
- Edge cases to cover: zero/negative amounts, extreme time values, missing PCA features, malformed request payloads
- Unit testing: Not yet implemented โ recommended:
pytest+httpx.AsyncClientfor FastAPI endpoint testing - Security testing: Not yet implemented โ recommended: basic input validation fuzzing and dependency vulnerability scanning (
pip-audit) - Performance testing: Not yet implemented โ recommended: load testing the
/predictendpoint withlocustork6
| Component | Platform | Notes |
|---|---|---|
| Backend (FastAPI) | Render | Auto-deploys from main branch; downloads model artifacts from GitHub Releases at startup |
| Frontend (Dashboard) | Vercel | Static site deploy, auto-deploys from main branch |
| Model Artifacts | GitHub Releases | Hosted as binary release assets to bypass Git's 100MB file limit |
To deploy your own copy:
- Fork this repo
- Create a GitHub Release and attach your own
model.pkl/scalers.pkl - Connect the
backend/folder as a new Web Service on Render, set the start command touvicorn main:app --host 0.0.0.0 --port $PORT - Connect the
frontend/folder as a new Vercel project - Update
MODEL_URL/SCALERS_URLin your backend environment variables to point to your release assets
Contributions are welcome!
- Fork the repository
- Create a feature branch:
git checkout -b feature/your-feature-name - Commit your changes:
git commit -m "Add your feature" - Push to your branch:
git push origin feature/your-feature-name - Open a Pull Request describing your changes
Please open an issue first for significant changes so we can discuss the approach.
Distributed under the MIT License. See LICENSE for more information.
- Kaggle Credit Card Fraud Detection Dataset
- scikit-learn
- FastAPI
- Render and Vercel for free-tier hosting
Built by Sakshi
โญ Star this repo if you found it useful!

