Model Analytics Dashboard
Stack v22 R² = 0.6495 134 Features Spatial CV
← Back to Map

THAMAN v22 — NYC Property Valuation Stack

XGBoost + LightGBM + CatBoost + Ridge meta · 134 features · HPD/DOB/QoL/transit signals · Spatial GroupKFold CV · 157,329 NYC sales

🎯
0.6495
R² Stack Holdout
XGB-A: 0.6445 · LGB: 0.6432 · CAT: 0.6482
📊
20.32%
Median APE
±20.32% confidence interval
💵
$1,047K
Mean Abs Error
v1 was $727K (−46%)
🎲
157K
Training Sales
27,763 held out (newest 15%)
🌳
5,000
Boosting Rounds
max across base learners · v1: 917
⚙️
134
Features
v5: 85 → v22: 134 features

⚖️ Stack v22 — Base Learner Comparison (XGB + LGB + CatBoost)

Holdout R² and MedAPE for each of the 4 diverse base models vs the v22 stack — 10-fold OOF + 5000 rounds + 134 features; Ridge meta-learner, MedAPE 20.32%

🔍 SHAP Feature Importance — Top 20 (v22)

Mean |SHAP| on holdout set — target-encoded bldgclass is now the strongest predictor, replacing raw building size

⚡ Baseline → v1 → v2 → Stack v22

v22 stack on 157K rows achieves R²=0.6495 vs baseline R²=0.238 — 134 features including HPD violations, DOB permits, rodent/heat complaints, MTA transit quality

📈 Train · Val · Holdout · Spatial CV

Spatial CV (0.523±0.18) is the true generalisation estimate — fold variance reveals borough-level difficulty

🗂️ Feature Categories Breakdown (134 Features)

v22 stack: 134 features — transit, QoL, HPD violations, DOB permits, prior sale price, NTA demographics, target-encoded building class

📦 SHAP Importance by Feature Group

Target encoding + urban gravity now dominate — together accounting for >50% of total model explanation

🔲 Price Tier Confusion Matrix (Holdout: 5,256 properties)

Rows = actual price tier · Columns = predicted tier · Diagonal = correct · Green outline = diagonal · Cell color = % of row total
Actual ↓ / Predicted → <$500K $500K–1M $1M–3M $3M–10M $10M+ Row Total
0–15%
15–40%
40–70%
70–100%
Correct prediction (diagonal)

📋 Per-Tier Classification Metrics

Precision, Recall and F1-score for each price tier — $3M–10M is hardest (low recall: 41.4%)

🗺️ MedAPE by Borough (Holdout)

Manhattan is 2× harder to predict than Staten Island — high price heterogeneity within NTAs drives the gap

📈 Model Progression — NYC & Riyadh

MedAPE improvement across training iterations — lower is better
Version Key Addition Hold R² MedAPE Features Δ MedAPE
Riyadh Version Key Addition Hold R² MedAPE Features Δ MedAPE

🎛️ Hyperparameter Profile (v22)

v22 stack uses regularised XGB+LGB+CAT+Ridge — LR=0.03, max_depth=8, subsample=0.7

📉 Error Metric Comparison (v22 vs baseline)

v22 stack MAE vs baseline — $1,047K vs $727K baseline

🗺️ QoL Winsorization Thresholds (p99)

Three Quality-of-Life features are capped at their 99th percentile — outlier NTAs (high crime/noise) are clipped to prevent dominating predictions

💰 ACRIS Missing-Value Imputation

When prior-sale data is absent, v22 training-set medians are used — applied to ~30% of properties with no prior-sale record

🚀 Luxury Cap Experiment (Spatial CV)

Capping sales at $10M reduces MedAPE from 22.4%→20.8% and improves R² — luxury outliers hurt generalisation

🎯 Predicted vs Actual — NYC Holdout (v22)

2,000 randomly sampled holdout transactions · coloured by borough · diagonal = perfect prediction · full-stack R²=0.6495 · MedAPE=20.32%
Loading scatter data…