This repository contains beginner-friendly Python data science projects, each with a Jupyter Notebook and a dedicated README**.
└── README.md # ← You are here
└── Statistics.ipynb # Notebook: descriptive statistics & visualisation
└── Regression Analysis.ipynb # Notebook: head and brain prediction with ML
└── Predictive Analysis.ipynb # Notebook: house price prediction with ML
└── Logistic Regression SUV Predictions.ipynb # Notebook: SUV purchase classification with Logistic Regression
└── Titanic Data Analysis.ipynb # Notebook: Titanic survival classification with Logistic Regression
├── Random_Forest_Classification_Penguins.ipynb # Notebook: penguin species classification with Random Forest
File: Statistics.ipynb
An introduction to descriptive statistics using a custom dataset and the classic Iris dataset.
| What you'll learn | Tools used |
|---|---|
| Mean, Median, Mode | statistics (built-in) |
| Variance & Standard Deviation | statistics.pvariance, stdev |
| Summary statistics on real data | pandas .describe() |
| Pairwise feature visualisation | seaborn pairplot |
💡 Key insight: Demonstrates why the median is more reliable than the mean when outliers are present — a right-skewed dataset makes this tangible.
File: Linear Regression.ipynb
A hands-on implementation of Linear Regression from scratch, then validated using Scikit-learn — predicting brain weight from head size measurements.
| What you'll learn | Tools used |
|---|---|
| Implementing regression manually (least squares) | numpy |
| Visualising regression lines & scatter plots | matplotlib |
| Calculating R² (coefficient of determination) | pandas |
| Validating with a ML library | sklearn.linear_model |
💡 Key result: Built Linear Regression from scratch using the least squares formula, then confirmed the result with Scikit-learn — both approaches yielding the same R² score.
File: Predictive Analysis.ipynb
A complete machine learning pipeline that predicts house prices using Linear Regression, achieving a 92% R² score on test data.
| What you'll learn | Tools used |
|---|---|
| Exploratory data analysis (EDA) | pandas |
| Visualising feature relationships | seaborn relplot |
| Train/test splitting | sklearn.model_selection |
| Building & evaluating a model | sklearn.linear_model |
💡 Key result: The model explains 92% of variance in house prices — a strong fit demonstrating the power of Linear Regression on structured data.
File: Logistic Regression SUV Predictions.ipynb
A complete classification pipeline that predicts whether a user will purchase an SUV based on age and salary, using Logistic Regression and achieving 89% accuracy on test data.
| What you'll learn | Tools used |
|---|---|
| Feature selection by index | pandas / numpy |
| Train/test splitting | sklearn.model_selection |
| Feature scaling | sklearn.preprocessing |
| Building & evaluating a classifier | sklearn.linear_model |
💡 Key result: The model correctly classifies 89% of unseen users as buyers or non-buyers — demonstrating Logistic Regression as a practical tool for binary classification problems.
File: Titanic Data Analysis.ipynb
A complete data analysis and classification pipeline that explores the Titanic dataset and predicts passenger survival based on features like age, sex, and class, using Logistic Regression and achieving 79% accuracy on test data.
| What you'll learn | Tools used |
|---|---|
| Exploratory data analysis (EDA) | pandas / numpy |
| Visualising survival patterns | seaborn countplot |
| Data cleaning & feature engineering | pandas |
| Building & evaluating a classifier | sklearn.linear_model |
| Confusion matrix & classification report | sklearn.metrics |
💡 Key result: The model achieves 79% accuracy in predicting survival — uncovering how factors like passenger class, sex, and age influenced the odds of surviving the Titanic disaster.
File: Random Forest Classification_Penguins.ipynb
A complete multi-class classification pipeline that predicts penguin species (Adelie, Chinstrap, Gentoo) from physical measurements using Random Forest, achieving 98% accuracy on test data.
| What you'll learn | Tools used |
|---|---|
| Data cleaning & null value handling | pandas |
| One-hot encoding of categorical features | pandas get_dummies |
| Train/test splitting | sklearn.model_selection |
| Building a Random Forest classifier | sklearn.ensemble |
Tuning number of trees (n_estimators) |
RandomForestClassifier |
💡 Key result: With just 7 decision trees and Gini impurity as the criterion, the model reaches 98% accuracy — demonstrating how ensemble methods outperform single classifiers on structured biological data.
Built with Python and the open-source data science ecosystem: pandas · NumPy · seaborn · scikit-learn · matplotlib