Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Cardiovascular Disease Prediction

A project that makes a data analysis and creates a model for classification of cardiovascular disease.

Dataset Information

cardiovascular-disease

This dataset was published on the Kaggle competition platform. For more information it can be seen accessing this link.

Project Structure

The project repository is structured as follows:

  • database

    • cardio.csv
    • cardio_test.csv
  • notebook

    • Heart_Disese.ipynb
  • README.md

  1. Database directory contains the datasets for training and test the model.
  2. Notebook directory contains the .ipynb file where all the data science pipelines performed form analyzing the data to final results.

Business performance

  • What is the accuracy and precision of the model?

    • Accuracy : 0.714
    • Precision : 0.733
  • How much profit will healthcare company have with the new tool?

    • The profit using the model is $11,285,500

Attributes

  1. Age | Objective Feature | age | int (days)
  2. Height | Objective Feature | height | int (cm) |
  3. Weight | Objective Feature | weight | float (kg) |
  4. Gender | Objective Feature | gender | categorical code |
  5. Systolic blood pressure | Examination Feature | ap_hi | int |
  6. Diastolic blood pressure | Examination Feature | ap_lo | int |
  7. Cholesterol | Examination Feature | cholesterol | 1: normal, 2: above normal, 3: well above normal |
  8. Glucose | Examination Feature | gluc | 1: normal, 2: above normal, 3: well above normal |
  9. Smoking | Subjective Feature | smoke | binary |
  10. Alcohol intake | Subjective Feature | alco | binary |
  11. Physical activity | Subjective Feature | active | binary |
  12. Presence or absence of cardiovascular disease | Target Variable | cardio | binary |

All of the dataset values were collected at the moment of medical examination.

Context

A healthcare company that specializes in detecting heart disease in the early stages. Its business model is service type, that is, the company offers an early diagnosis of cardiovascular disease for a certain price.

Currently, the diagnosis of cardiovascular disease is made manually by a team of specialists. The current accuracy of the diagnosis varies between 55% and 65%, due to the complexity of the diagnosis and also the fatigue of the team who take turns to minimize the risks. The cost of each diagnosis, including equipment and analysts' payroll, is around $ 1,000.00.

The price of the diagnosis, paid by the client, varies according to the precision achieved by the team of specialists, the client pays R$ 500.00 for every 5% accuracy above 50%. For example, for an accuracy of 55%, the diagnosis costs R$ 500.00 for the client, for an accuracy of 60%, the value is R$ 1000.00 and so on. If the diagnostic accuracy is 50%, the customer does not pay for it.

Objective

As a Data Scientist, objective is to create a tool that increases the accuracy of the diagnosis and that this accuracy is stable for all diagnosis.

Pipeline

  • Opening

  • Data Descriptions

  • Feature Engineering

  • Data Exploration

  • Filtering Variables

  • Exploratory Data Analysis

  • Data Preparation

  • Feature Selection

  • Machine Learning Modeling

  • Hyperparameter Fine Tuning

  • Traduction and Error's Interpretation

  • Deploy

Reporting Issues

If you encounter any issues or have any suggestions regarding the project, feel free to open an issue in the project repository.

Contribution Guidelines

Contributions to this project are welcome! If you would like to contribute, please follow these steps:

  1. Fork the repository.

  2. Create a new branch for your feature or bug fix.

  3. Commit your changes and push the branch to your fork.

  4. Submit a pull request to the main repository.

  5. Please ensure that your contributions align with the project's coding style and follow the best practices of data science.

Contact Information

For any additional information or inquiries, please go to my profile and send a mail.

Reference

Happy analyzing!

We hope you find this project insightful and informative! Happy analyzing!

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages