Overview This Jupyter notebook implements a K-Nearest Neighbors (KNN) classifier to identify iris flower species based on their sepal and petal measurements. The model achieves 95.56% accuracy on test data.
Dataset The dataset contains 150 samples of iris flowers with the following characteristics:
Features (4) Sepal Length (cm)
Sepal Width (cm)
Petal Length (cm)
Petal Width (cm)
Target Classes (3) Iris-setosa
Iris-versicolor
Iris-virginica
Dependencies The notebook requires the following Python libraries:
numpy - Numerical operations
pandas - Data manipulation
matplotlib - Data visualization
seaborn - Statistical visualizations
scikit-learn - Machine learning algorithms
Implementation Steps
- Data Loading & Exploration Loads CSV file from specified directory
Displays dataset shape, first 10 rows, and column information
Shows statistical summary and data types
Checks for missing values (none present)
- Data Preprocessing Removes unnecessary Id column
Separates features (X) and target (y)
Encodes species names to numerical values using LabelEncoder
Standardizes features using StandardScaler for optimal KNN performance
- Train-Test Split Splits data into 70% training (105 samples) and 30% testing (45 samples)
Uses stratified split to maintain class balance
- KNN Model Training Tests different k values (1, 3, 5, 7, 9, 11, 13, 15) and selects the best performing one:
k Value Train Accuracy Test Accuracy 1 100.00% 93.33% 3 97.14% 91.11% 5 98.10% 91.11% 7 97.14% 93.33% 9 97.14% 95.56% 11 97.14% 95.56% 13 97.14% 93.33% 15 98.10% 93.33% Best k = 9 with 95.56% test accuracy
- Predictions The model successfully classifies new iris flowers with confidence distances:
Flower Measurements (cm) Predicted Species Confidence 1 [5.1, 3.5, 1.4, 0.2] Iris-setosa 0.288 2 [6.5, 3.0, 5.2, 2.0] Iris-virginica 0.518 3 [6.0, 2.9, 4.5, 1.5] Iris-versicolor 0.412 4 [4.9, 3.0, 1.5, 0.1] Iris-setosa 0.245 5 [6.8, 3.2, 5.9, 2.3] Iris-virginica 0.427 6 [5.8, 2.7, 4.1, 1.0] Iris-versicolor 0.416 6. Visualizations The notebook generates three visualizations:
Accuracy vs k Value - Comparing train/test accuracy across different k values
Confusion Matrix Heatmap - Visualizing classification performance
Feature Scatter Plots - Sepal and petal characteristics by species
Results Classification Report Species Precision Recall F1-Score Support Iris-setosa 1.00 1.00 1.00 15 Iris-versicolor 0.88 1.00 0.94 15 Iris-virginica 1.00 0.87 0.93 15 Confusion Matrix text [[15 0 0] [ 0 15 0] [ 0 2 13]] Key Findings The model perfectly identifies Iris-setosa (100% precision and recall)
Iris-versicolor is sometimes confused with Iris-virginica (2 misclassifications)
Feature scaling significantly improves KNN performance
A moderate k value (k=9) provides the best balance between bias and variance
Usage Notes Update the dataset_path variable to point to your Iris dataset location
The notebook expects a CSV file containing the Iris dataset with appropriate headers
The first run will install required dependencies automatically