This project demonstrates a complete Supervised Learning Classification workflow using the Iris Dataset and a Decision Tree Classifier. The objective is to classify Iris flowers into three species based on their physical measurements.
| Item | Value |
|---|---|
| Task | Classification |
| Dataset | Iris Dataset |
| Samples | 150 |
| Features | 4 |
| Classes | 3 |
| Model | Decision Tree Classifier |
| Library | Scikit-Learn |
Dataset
↓
Exploratory Data Analysis (EDA)
↓
Feature & Label Selection
↓
Train-Test Split
↓
Model Creation
↓
Model Training
↓
Prediction
↓
Evaluation
↓
Visualization
↓
Conclusion
The Iris dataset contains measurements of Iris flowers from three different species:
- Setosa
- Versicolor
- Virginica
- Sepal Length
- Sepal Width
- Petal Length
- Petal Width
- Species
This dataset was obtained from Kaggle:
- Dataset: Iris Dataset
- Author: Himanshu Nakrani
The dataset was analyzed to understand:
- Dataset shape
- Available columns
- Missing values
- Class distribution
The dataset contains:
- 150 samples
- 4 numerical features
- 3 balanced classes
The dataset was divided into:
- 80% Training Data
- 20% Testing Data
The parameter stratify=y was used to preserve class distribution in both training and testing sets.
The parameter random_state=42 was used to ensure reproducible results.
A Decision Tree Classifier was used for the classification task.
from sklearn.tree import DecisionTreeClassifier
model = DecisionTreeClassifier()The model was evaluated using:
- Accuracy
- Classification Report
- Confusion Matrix
| Metric | Value |
|---|---|
| Accuracy | 96.67% |
| Correct Predictions | 29 / 30 |
| Misclassified Samples | 1 |
Shows the number of correct and incorrect predictions for each class.
Visualizes the relationship between petal length and petal width across different flower species.
Displays how much each feature contributes to the model's decision-making process.
- Decision Tree achieved 96.67% accuracy
- Only 1 sample was misclassified
- Petal Width was identified as the most important feature
- The dataset was perfectly balanced
- Setosa was clearly separable from the other species
Future work may include:
- Random Forest Classifier
- K-Nearest Neighbors (KNN)
- Support Vector Machine (SVM)
- Hyperparameter Tuning
- Model Performance Comparison
- Python
- Pandas
- Matplotlib
- Scikit-Learn
- Jupyter Notebook
iris-classification-sklearn/
│
├── data/
│ └── iris.csv
│
├── notebooks/
│ └── iris.ipynb
│
├── images/
│ ├── confusion_matrix.png
│ ├── scatter_plot.png
│ └── feature_importance.png
│
├── README.md
├── requirements.txt
└── .gitignore
This project was created as a beginner machine learning project to practice the complete supervised learning workflow using the Iris dataset.


