In a lot of domains — credit, healthcare, hiring — a model that's accurate but unexplainable is worse than useless: it's a liability. If you can't say why a decision was made, you can't defend it to a regulator, debug it when it's wrong, or earn the trust of the person it affects. That's the constraint I designed CreditSense around: every score arrives with the reasons behind it.
SHAP: attributions per prediction
The workhorse there is SHAP. It gives you per-prediction attributions — for this specific applicant, which features pushed the score up, which pulled it down, and by how much. That's a fundamentally different artifact from global feature importance, which only tells you what mattered on average across the dataset. Individuals don't live in the average.
Two things that are easy to get wrong
Two things are easy to get wrong. First, correlated features smear attribution across each other, so a "low" importance can be hiding behind a twin feature — you have to reason about the feature set, not just read the plot. Second, exact SHAP values are expensive; for tree ensembles like XGBoost, TreeSHAP makes it tractable, and that pairing is a big reason gradient-boosted models remain a sweet spot for tabular, high-stakes work.
Log the reason with the score
The operational habit that pays off: log the attributions alongside the prediction, not on demand later. When someone asks why a decision was made six months on, you want the answer stored next to the score — reconstructing it after a model has been retrained is how you end up unable to explain your own system.
import shap
explainer = shap.TreeExplainer(model) # tractable for tree ensembles
shap_values = explainer.shap_values(x_row) # per-applicant attributions
record = {
"score": float(model.predict_proba(x_row)[0, 1]),
"reasons": top_features(shap_values, x_row), # store WHY, next to WHAT
}
audit_log.write(record) # reconstructing this after a retrain is too lateA design requirement, not a slide
Explainability isn't a chart you generate for the slide deck at the end. It's a design requirement that shapes which model you pick and what you store — and it's the difference between a model you can ship and one you can only demo.
Sources & further reading
- Lundberg & Lee, A Unified Approach to Interpreting Model Predictions (2017) — the SHAP framework.
- Lundberg et al., Explainable AI for Trees: From Local Explanations to Global Understanding (2020) — TreeSHAP, exact and fast for tree ensembles.
- SHAP documentation — the maintained library and its explainers.
Editorial note — Draws on my own work on CreditSense AI (explainable credit-risk scoring). The techniques described — SHAP, TreeSHAP, gradient boosting — are established; no performance figures for CreditSense are quoted until they've been measured.


