Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

4 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Iris Classification with Scikit-Learn

This project demonstrates a complete Supervised Learning Classification workflow using the Iris Dataset and a Decision Tree Classifier. The objective is to classify Iris flowers into three species based on their physical measurements.

Project Information

Item Value
Task Classification
Dataset Iris Dataset
Samples 150
Features 4
Classes 3
Model Decision Tree Classifier
Library Scikit-Learn

Project Workflow

Dataset
↓
Exploratory Data Analysis (EDA)
↓
Feature & Label Selection
↓
Train-Test Split
↓
Model Creation
↓
Model Training
↓
Prediction
↓
Evaluation
↓
Visualization
↓
Conclusion

Dataset

The Iris dataset contains measurements of Iris flowers from three different species:

  • Setosa
  • Versicolor
  • Virginica

Features

  • Sepal Length
  • Sepal Width
  • Petal Length
  • Petal Width

Target

  • Species

Dataset Source

This dataset was obtained from Kaggle:

Exploratory Data Analysis (EDA)

The dataset was analyzed to understand:

  • Dataset shape
  • Available columns
  • Missing values
  • Class distribution

The dataset contains:

  • 150 samples
  • 4 numerical features
  • 3 balanced classes

Train-Test Split

The dataset was divided into:

  • 80% Training Data
  • 20% Testing Data

The parameter stratify=y was used to preserve class distribution in both training and testing sets.

The parameter random_state=42 was used to ensure reproducible results.

Model

A Decision Tree Classifier was used for the classification task.

from sklearn.tree import DecisionTreeClassifier

model = DecisionTreeClassifier()

Evaluation Metrics

The model was evaluated using:

  • Accuracy
  • Classification Report
  • Confusion Matrix

Results

Metric Value
Accuracy 96.67%
Correct Predictions 29 / 30
Misclassified Samples 1

Visualizations

Confusion Matrix

Shows the number of correct and incorrect predictions for each class.

Confusion Matrix

Scatter Plot

Visualizes the relationship between petal length and petal width across different flower species.

Scatter Plot

Feature Importance

Displays how much each feature contributes to the model's decision-making process.

Feature Importance

Key Findings

  • Decision Tree achieved 96.67% accuracy
  • Only 1 sample was misclassified
  • Petal Width was identified as the most important feature
  • The dataset was perfectly balanced
  • Setosa was clearly separable from the other species

Future Improvements

Future work may include:

  • Random Forest Classifier
  • K-Nearest Neighbors (KNN)
  • Support Vector Machine (SVM)
  • Hyperparameter Tuning
  • Model Performance Comparison

Technologies Used

  • Python
  • Pandas
  • Matplotlib
  • Scikit-Learn
  • Jupyter Notebook

Repository Structure

iris-classification-sklearn/
│
├── data/
│   └── iris.csv
│
├── notebooks/
│   └── iris.ipynb
│
├── images/
│   ├── confusion_matrix.png
│   ├── scatter_plot.png
│   └── feature_importance.png
│
├── README.md
├── requirements.txt
└── .gitignore

Author

This project was created as a beginner machine learning project to practice the complete supervised learning workflow using the Iris dataset.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages