An end-to-end data science project on the Boston House Prices dataset: exploratory data analysis (EDA), visualizations, outlier detection, 3D spatial analysis, and predictive modeling with multiple regression/ensemble models.
Note on the dataset: this project uses a reduced 3-column CSV version of the classic Boston housing dataset (506 rows), containing only:
Rooms,Distance, andValue(the target / MEDV).
| Column | Description | Type |
|---|---|---|
| Rooms | Average number of rooms per dwelling (RM) | float |
| Distance | Weighted distances to Boston employment centers (DIS) | float |
| Value | Median value of owner-occupied homes in $1000s (target / MEDV) | float |
- Rows: 506 · Columns: 3
- No missing, infinite, or duplicate values.
Valueis capped at $50,000 (right-censored) — homes above the cap were recorded as exactly50, which creates a visual plateau in the histogram.
data_analysis/
├── Boston House Prices.csv # Raw dataset
├── boston_analysis.py # End-to-end EDA + modeling pipeline
├── univariate_histograms.py # Histograms + KDE overlays
├── boxplot_iqr_outliers.py # Boxplots + IQR outlier detection
├── scatter3d_viridis.py # 3D scatter (viridis) + cluster analysis
├── mlr_diagnostics.py # Multiple Linear Regression diagnostics
├── requirements.txt # Python dependencies
├── README.md # This file
└── output/ # Generated figures + performance summary
├── 01_target_distribution.png
├── 02_correlation_heatmap.png
├── 03_scatterplots_top_features.png
├── 04_model_comparison.png
├── 05_actual_vs_predicted.png
├── 06_feature_importance.png
├── 07_univariate_histograms_kde.png
├── 08_boxplots_iqr_outliers.png
├── 09_scatter3d_viridis.png
├── 09b_scatter3d_altview.png
├── 09c_scatter3d_clusters.png
├── 10_mlr_diagnostics.png
└── performance_summary.md
Requires Python 3.8+.
# Clone / enter the project directory
cd data_analysis
# (Recommended) create a virtual environment
python -m venv venv
# Windows:
venv\Scripts\activate
# macOS / Linux:
source venv/bin/activate
# Install dependencies
pip install -r requirements.txtOptional: for the XGBoost model in the main pipeline, also install
pip install xgboost.
Each script is standalone and produces figures in the output/ folder:
python boston_analysis.py # Full EDA + modeling pipeline
python univariate_histograms.py # Histograms + KDE
python boxplot_iqr_outliers.py # Boxplots + IQR outliers
python scatter3d_viridis.py # 3D scatter + clusters
python mlr_diagnostics.py # MLR diagnosticsThe target variable (Value) is right-skewed with a visible plateau/cap at 50
— an artifact of the $50k truncation in the classic Boston dataset.
Rooms is the strongest correlate of house price; Distance shows a weak
positive correlation in this reduced dataset.
Scatter plots with trendlines for the top features correlated with Value.
| Variable | Skewness | Mean | Median |
|---|---|---|---|
| Rooms | +0.40 | 6.28 | 6.21 |
| Distance | +1.01 | 3.80 | 3.21 |
| Value | +1.10 | 22.53 | 21.20 |
All three are right-skewed; Value is additionally capped at 50
(3.16% of rows sit exactly at the maximum).
Outliers detected via the 1.5 × IQR method:
| Column | Outlier count | Notable indices |
|---|---|---|
| Rooms | 30 | 97, 98, 162, 163, 166, 180, … |
| Distance | 5 | 351–355 |
| Value | 40 | 97, 98, 161, 162, … (many at cap=50) |
3D scatter of Rooms (X) × Distance (Y) × Value (Z) colored by Value
with the viridis colormap:
An alternative viewing angle:
Colored by KMeans cluster (4 clusters reveal a distinct luxury cluster of large-room, high-value homes):
Models trained on an 80/20 train/test split (random_state=42) with
StandardScaler feature scaling.
| Model | MAE | RMSE | R² |
|---|---|---|---|
| Random Forest 🔥 | 3.321 | 4.926 | 0.6691 |
| Decision Tree | 3.696 | 5.034 | 0.6545 |
| Gradient Boosting | 3.412 | 5.348 | 0.6099 |
| Ridge Regression | 4.296 | 6.660 | 0.3952 |
| Lasso Regression | 4.297 | 6.661 | 0.3949 |
| Linear Regression | 4.297 | 6.662 | 0.3948 |
(XGBoost included automatically when installed.)
Actual vs. Predicted scatter with the ideal y = x reference line:
Feature Importance — Rooms dominates (≈ 73%), Distance ≈ 27%:
Value ~ Rooms + Distance — RMSE = 6.662, R² = 0.3948, MAE = 4.297.
The Residual vs. Predicted plot shows a roughly homoscedastic spread
(no clear funnel pattern).
- Best model: Random Forest (R² = 0.669) — ensemble averaging reduces variance vs. a single Decision Tree and captures non-linear patterns that linear models miss (R² ≈ 0.39).
- Primary driver: Rooms is by far the strongest predictor of house price (r ≈ +0.70). The top-10% value homes have median 7.5 rooms vs the dataset median of 6.2.
- Distance has a secondary, weakly-positive effect in this reduced file (unlike the classic dataset where it is negative).
- Target capping:
Valueis truncated at $50k, biasing regression downward for luxury homes.
- Only 2 features — restricts model expressiveness and caps achievable R².
- Right-censored target (
Valuecapped at 50) distorts the upper tail. Distanceis positively correlated withValuehere, an artifact of this reduced dataset.
- Use the full 13-feature Boston (or California) housing dataset.
- Hyperparameter tuning (GridSearchCV / Optuna) + k-fold cross-validation.
- Model the $50k cap explicitly (censored/truncated regression).
- Add engineered features (interactions, room-size ratios, log transforms).
For educational purposes only. Data derived from the classic Boston housing dataset (Harrison & Rubinfeld, 1978).











