Skip to content

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Repository files navigation

Boston House Prices — Data Analysis & Predictive Modeling

An end-to-end data science project on the Boston House Prices dataset: exploratory data analysis (EDA), visualizations, outlier detection, 3D spatial analysis, and predictive modeling with multiple regression/ensemble models.

Note on the dataset: this project uses a reduced 3-column CSV version of the classic Boston housing dataset (506 rows), containing only: Rooms, Distance, and Value (the target / MEDV).


📦 Dataset

Column Description Type
Rooms Average number of rooms per dwelling (RM) float
Distance Weighted distances to Boston employment centers (DIS) float
Value Median value of owner-occupied homes in $1000s (target / MEDV) float
  • Rows: 506 · Columns: 3
  • No missing, infinite, or duplicate values.
  • Value is capped at $50,000 (right-censored) — homes above the cap were recorded as exactly 50, which creates a visual plateau in the histogram.

🗂️ Project Structure

data_analysis/
├── Boston House Prices.csv        # Raw dataset
├── boston_analysis.py             # End-to-end EDA + modeling pipeline
├── univariate_histograms.py       # Histograms + KDE overlays
├── boxplot_iqr_outliers.py        # Boxplots + IQR outlier detection
├── scatter3d_viridis.py           # 3D scatter (viridis) + cluster analysis
├── mlr_diagnostics.py             # Multiple Linear Regression diagnostics
├── requirements.txt               # Python dependencies
├── README.md                      # This file
└── output/                        # Generated figures + performance summary
    ├── 01_target_distribution.png
    ├── 02_correlation_heatmap.png
    ├── 03_scatterplots_top_features.png
    ├── 04_model_comparison.png
    ├── 05_actual_vs_predicted.png
    ├── 06_feature_importance.png
    ├── 07_univariate_histograms_kde.png
    ├── 08_boxplots_iqr_outliers.png
    ├── 09_scatter3d_viridis.png
    ├── 09b_scatter3d_altview.png
    ├── 09c_scatter3d_clusters.png
    ├── 10_mlr_diagnostics.png
    └── performance_summary.md

🔧 Setup & Installation

Requires Python 3.8+.

# Clone / enter the project directory
cd data_analysis

# (Recommended) create a virtual environment
python -m venv venv
# Windows:
venv\Scripts\activate
# macOS / Linux:
source venv/bin/activate

# Install dependencies
pip install -r requirements.txt

Optional: for the XGBoost model in the main pipeline, also install pip install xgboost.


▶️ How to Run

Each script is standalone and produces figures in the output/ folder:

python boston_analysis.py          # Full EDA + modeling pipeline
python univariate_histograms.py    # Histograms + KDE
python boxplot_iqr_outliers.py     # Boxplots + IQR outliers
python scatter3d_viridis.py        # 3D scatter + clusters
python mlr_diagnostics.py          # MLR diagnostics

All results (figures + performance_summary.md) are saved to output/.

📊 Exploratory Data Analysis

Target Distribution

The target variable (Value) is right-skewed with a visible plateau/cap at 50 — an artifact of the $50k truncation in the classic Boston dataset.

Target Distribution

Correlation Heatmap

Rooms is the strongest correlate of house price; Distance shows a weak positive correlation in this reduced dataset.

Correlation Heatmap

Feature Scatter Plots

Scatter plots with trendlines for the top features correlated with Value.

Top Feature Scatter Plots


📈 Distribution & Outlier Analysis

Univariate Histograms with KDE

Variable Skewness Mean Median
Rooms +0.40 6.28 6.21
Distance +1.01 3.80 3.21
Value +1.10 22.53 21.20

All three are right-skewed; Value is additionally capped at 50 (3.16% of rows sit exactly at the maximum).

Univariate Histograms + KDE

Boxplots & IQR Outliers

Outliers detected via the 1.5 × IQR method:

Column Outlier count Notable indices
Rooms 30 97, 98, 162, 163, 166, 180, …
Distance 5 351–355
Value 40 97, 98, 161, 162, … (many at cap=50)

Boxplots + IQR Outliers


🧭 3D Spatial Analysis

3D scatter of Rooms (X) × Distance (Y) × Value (Z) colored by Value with the viridis colormap:

3D Scatter (viridis)

An alternative viewing angle:

3D Scatter (alt view)

Colored by KMeans cluster (4 clusters reveal a distinct luxury cluster of large-room, high-value homes):

3D Scatter Clusters

🤖 Predictive Modeling

Models trained on an 80/20 train/test split (random_state=42) with StandardScaler feature scaling.

Model Comparison

Model MAE RMSE R²
Random Forest 🔥 3.321 4.926 0.6691
Decision Tree 3.696 5.034 0.6545
Gradient Boosting 3.412 5.348 0.6099
Ridge Regression 4.296 6.660 0.3952
Lasso Regression 4.297 6.661 0.3949
Linear Regression 4.297 6.662 0.3948

(XGBoost included automatically when installed.)

Model Comparison

Best Model — Random Forest

Actual vs. Predicted scatter with the ideal y = x reference line:

Actual vs Predicted

Feature Importance — Rooms dominates (≈ 73%), Distance ≈ 27%:

Feature Importance

Multiple Linear Regression Diagnostics

Value ~ Rooms + Distance — RMSE = 6.662, R² = 0.3948, MAE = 4.297. The Residual vs. Predicted plot shows a roughly homoscedastic spread (no clear funnel pattern).

MLR Diagnostics


🌟 Key Insights

  1. Best model: Random Forest (R² = 0.669) — ensemble averaging reduces variance vs. a single Decision Tree and captures non-linear patterns that linear models miss (R² ≈ 0.39).
  2. Primary driver: Rooms is by far the strongest predictor of house price (r ≈ +0.70). The top-10% value homes have median 7.5 rooms vs the dataset median of 6.2.
  3. Distance has a secondary, weakly-positive effect in this reduced file (unlike the classic dataset where it is negative).
  4. Target capping: Value is truncated at $50k, biasing regression downward for luxury homes.

⚠️ Data Limitations

  • Only 2 features — restricts model expressiveness and caps achievable R².
  • Right-censored target (Value capped at 50) distorts the upper tail.
  • Distance is positively correlated with Value here, an artifact of this reduced dataset.

Suggested Future Improvements

  • Use the full 13-feature Boston (or California) housing dataset.
  • Hyperparameter tuning (GridSearchCV / Optuna) + k-fold cross-validation.
  • Model the $50k cap explicitly (censored/truncated regression).
  • Add engineered features (interactions, room-size ratios, log transforms).

📄 License

For educational purposes only. Data derived from the classic Boston housing dataset (Harrison & Rubinfeld, 1978).

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages