This project presents a comprehensive machine learning pipeline for detecting fraudulent financial transactions. The pipeline involves data preprocessing, model training, evaluation, and deployment of machine learning models to serve real-time predictions via a RESTful API. The objective is to address the challenge of identifying fraudulent transactions in a highly imbalanced dataset using advanced techniques and tools within an MLOps framework. The implementation includes the use of XGBoost and deep learning models, which are evaluated and deployed for practical use.
- Introduction
- Data Preprocessing
- Model Training
- Model Evaluation
- Deployment
- Results
- Paper
- Usage
- Contributing
- License
Fraud detection in financial transactions has emerged as a critical challenge in the modern financial landscape. This project aims to develop a robust, scalable, and interpretable fraud detection system using both XGBoost and deep learning models within an MLOps framework.
The data preprocessing steps include:
- Loading the dataset.
- Checking for missing values.
- Standardizing the features.
- Handling imbalanced data using SMOTE (Synthetic Minority Over-sampling Technique).
The preprocessing code can be found in /data_preprocessing.py.
We implemented two models:
- XGBoost Model: A powerful gradient boosting algorithm.
- Deep Learning Model: A neural network using Keras.
The training code for these models is located in /model_training.py.
The XGBoost model is trained with hyperparameter tuning using GridSearchCV.
The deep learning model is built using Keras and hyperparameters are optimized using Keras Tuner.
Model evaluation includes:
- Generating classification reports.
- Plotting confusion matrices.
- Calculating ROC AUC scores.
- Using SHAP values for model interpretability.
The evaluation code is in /model_evaluation.py.
The trained models are deployed using a Flask web application, with Docker used to containerize the application for easy deployment across different environments.
/predict_xgb: Returns the prediction made by the XGBoost model./predict_dl: Returns the prediction made by the deep learning model.
The deployment code is located in /app.py.
The results of the model evaluation, including performance metrics and SHAP analysis, are detailed in the paper.
You can read the full paper, "Fraud Detection in Financial Transactions Using Machine Learning: A Comprehensive MLOps Approach," here.
To use this project, follow these steps:
-
Clone the repository:
git clone https://github.com/your-repository-link.git cd your-repository -
Install dependencies:
pip install -r requirements.txt
-
Run the data preprocessing:
python scripts/data_preprocessing.py
-
Train the models:
python scripts/train.py
-
Run the Flask app:
python app.py
-
Make predictions: Use an API client like Postman to send POST requests to the
/predict_xgbor/predict_dlendpoints with transaction data in JSON format.
Contributions are welcome! Please open an issue or submit a pull request for any improvements or bug fixes.
This project is licensed under the Apache License - see the LICENSE file for details.