NSBA Draft Analyticsembargoed · 2026-06-06

93_usage_bias_skeptic.md

F93 — Skeptic audit of the usage / empty-stats bias work (F32, F33)

F93 — Skeptic audit of the usage / empty-stats bias work (F32, F33)

Date: 2026-05-30. Role: adversarial red-team of 32_usage_bias.md and 33_contestation_adjusted_value.md. Mandate (from David): is the measured teammate effect real or just regression-to-mean / collinearity? Is the contestation model identifiable at this n, or driven by a few games? Are the empty-stats flags robust? And the only test that matters for drafting — does adjusting predict OUT-OF-SAMPLE (next-season PPTF on a different team) better than raw PPTF?

Bottom line (lead with this): Both findings reproduce exactly, and their headline conclusion is correct and if anything UNDERSTATED: the usage/empty-stats bias should change the draft board essentially zero. I ran the out-of-sample test neither finding ran — predict next-season PPTF on a different team — and the contestation adjustment does not beat raw PPTF (leak-free Δ = −0.04 Spearman, 95% CI [−0.18, +0.04], P(adj better)=13%). The adjustment buys nothing for drafting; raw PPTF is the right currency. The teammate-invariance of the F33 model is genuine (I re-derived the synthetic recovery independently; corr(error, teammate strength) ≈ 0 even under an adversarial lone-stars-on-weak-teams design). But three softer claims do not survive: (1) F32's clean teammate coefficient is not robust — it flips to null (p≈0.5) under my reasonable spec and is null in every single season; the within-player signal is mostly ordinary regression-to-mean, not a teammate penalty. (2) F33's specific per-name "empty-stats" flags rest on majority/fully nsba3-approximated contesting sets with 7–18 wins — they are the noisiest cells in the dataset and cannot survive F33's own lineup-only robustness refit. (3) F32 and F33 contradict each other on their own headline names: F32 flags ne (#7) and dan.k.memes (#8) as the sharpest empty-stats risks; the better-designed F33 model keeps ne at #7 (zero shift) and dan.k.memes in the top-15.

How much should usage bias change the draft board? Effectively nothing. Keep raw PPTF/VORP. Do NOT apply a team-strength correction. Demote F32's ne/dan.k.memes "carry flags" to noise. The two flags worth one tiebreaker tick are F33's abcisosm5 and akul (lineup-backed "suppressed" cases) — and even those are within-tier.

Verdicts: F33 → TRUST (the invariance proof and the bounded-null are sound; the per-name flag table is the weak part). F32 → WATCHLIST → trim (direction OK, but the coefficient is spec-fragile and the two named flags are contradicted by F33).


What reproduces (full credit)


Test 4 (the decisive one): does adjusting predict OUT-OF-SAMPLE better than raw PPTF?

Neither finding ran this, yet it is the only test of drafting value. F33's Spearman 0.93 is an in-sample re-description; "does the adjustment help you draft" = does it predict a player's PPTF next season on a different roster better than raw PPTF does.

The leakage trap (and why a naive version lies). Both adjusted_ability and teammate_adj_pptf are career-pooled over ALL clean games including season t+1. Regressing the pooled predictor on season t+1 PPTF gives a flattering Spearman +0.74 (adj) / +0.85 (F32 teammate-adj) vs +0.49 (raw) — but this is pure leakage; the predictor literally contains the target's games. Do not believe it.

Leak-free test. I refit the contestation model on train seasons ≤ t only, then predicted each player's season-(t+1) PPTF on their new team. Pairs: 23 (TUF≥30 both seasons), 24 players, all on different teams (the exact confound-relevant set):

Predictor (trained on ≤ t) Spearman → next-season PPTF Pearson
raw PPTF (prior season) +0.523 +0.417
raw PPTF (pooled train) +0.523 +0.418
contestation adj (train-only) +0.471 +0.444

Bootstrap of the difference (adj − raw, 2000 resamples): median −0.042, 95% CI [−0.181, +0.036], P(adj beats raw) = 0.13. The contestation adjustment is, within power, tied-to-slightly-worse than raw PPTF at predicting the next season on a new team.

Implication. This is the real test of David's drafting worry ("draft the empty-stats guy, he regresses on a better team"). If the empty-stats confound were real and large, the adjusted ability would predict the new-team season better than raw PPTF. It does not. The adjustment provides no drafting edge. This both (a) refutes any reading of F32/F33 as "use the adjusted board to draft" and (b) confirms F33's actual conclusion — raw PPTF is the currency and survives the empty-stats objection. n=23 is tiny (caveat below), but it is the honest number and it points clearly against the adjustment.


Test 1: is the teammate effect real, or regression-to-mean / collinearity?

Mostly the latter. Rebuilding M1 from raw (own combine θ + leave-one-out mean teammate θ, standardized PPTF):

Spec teammate-θ coef p n
F32 reported M1 (with SOS) −0.039 0.027 125
My rebuild (own θ + tm θ) −0.051 0.527 106
Team-mean θ incl. self (kills leave-one-out exclusion) −0.065 0.445 114
nsba2 only −0.124 0.293 45
nsba3 only −0.111 0.339 59

The sign is consistently negative but the magnitude is small and the significance is spec-fragile — F32's p=0.027 evidently leans on the SOS control + a specific sample; drop it and the coefficient is null (p≈0.5), and it is null in every single season (matching F32's own load-bearing caveat). Note corr(own θ, teammate θ) = −0.22 — the leave-one-out exclusion mechanically gives the strongest player the weakest "teammate" mean, manufacturing part of the negative slope.

The regression-to-mean confound is real and competes with the teammate story. Among the 26 within-player season transitions: corr(PPTF_t, ΔPPTF) = −0.40 — high scorers decline, low scorers rise, regardless of teammate change. The confound's own signature is weak: corr(team-strength change, ΔPPTF) = −0.18, and controlling for mean-reversion (PPTF_t), the partial coefficient of team-strength-change on ΔPPTF is −0.13, p=0.55 — null. So within-player there is no clean evidence that moving to a stronger team lowers PPTF beyond ordinary mean-reversion. F32's "real but fragile" is the honest read; I'd push it toward "direction suggestive, magnitude indistinguishable from mean-reversion + a leave-one-out artifact."


Test 2: is the contestation model identifiable / teammate-invariant at this n?

The invariance is genuine — this is the strongest part of the whole effort. I re-derived the synthetic recovery from scratch (my own simulator, Luce-generated winners):

So the estimator does not launder team strength into a player's estimate, even when the data-generating process is designed to tempt it. F33's Result 3 holds. (My recovery Spearman 0.98–0.99 is higher than F33's reported 0.93 because my sim has more contests/less ridge bite; the load-bearing number — corr(error, teammate strength)≈0 — replicates cleanly.)

Did the adjustment re-introduce team strength another way? No — and the data confirms the opposite: corr(adjusted, team win%) = +0.111 vs corr(raw PPTF, team win%) = +0.093. The adjustment does not even reduce the (already weak) tie to team strength. The "teammate- invariant" claim is achieved, not laundered.

Identifiability caveat that bites the flags, not the ranking: 29/181 players have contest_n < 50 (heavily shrunk); the ranking is robust (ridge + lineup-only refit), but individual low-contest estimates are noise.


Test 3: are the empty-stats flags robust to sample/approximation? NO.

This is where F33 oversells. The six "empty-stats" names in its flag table rest on majority-or-fully nsba3 buzzed-roster-approximated contesting sets:

F33 "empty-stats" flag approx_frac win_n
Euna Kim 1.00 7
Chris Wang 1.00 8
Aldric Benalan 1.00 13
michael 1.00 18
dan.k.memes 0.77 53
JoshuaW 0.76 56

All six are >50% approximated; four are 100% approximated. F33's own Caveat 2 admits the buzzed-roster approximation "makes the contest look smaller and stronger than reality, biasing nsba3 abilities mildly upward." These flags therefore cannot survive F33's R4 lineup-only robustness refit — drop the approximated events and Euna Kim / Chris Wang / Aldric Benalan / michael have no data left. R4's Spearman 0.95 validates the aggregate ranking, not these specific names. The per-name empty-stats table is the noisiest cell in the analysis and should be treated as illustrative, not actionable. (F33 already hedges "none are early-round picks / use as tiebreaker" — correct, and this is why it stays TRUST.)

The two "suppressed" flags are better grounded: abcisosm5 (approx_frac 0, 179 contests, on a 0.73 team) is a clean lineup-backed case; akul (approx 0.39, 307 contests) is partly approximated but lineup-anchored. These are the only two name-level signals I'd act on, and only as a within-tier nudge.


The cross-finding contradiction (F32 vs F33 on their own headlines)

F32's draft action names ne (#7) and dan.k.memes (#8) as "the two clearest top-10 empty-stats risks… discount toward combine θ when projecting onto a stronger roster."

The F33 contestation model — which is the purpose-built, teammate-invariant test of exactly this — disagrees:

name raw rank F33 adj rank shift approx_frac win_n
ne 7 7 0 0.00 61
dan.k.memes 8 15 −7 0.77 53

ne has clean lineup data (approx 0), 61 head-to-head buzz wins, and zero rank movement under the teammate-invariant model — i.e. ne wins buzzes against the whole room (opponents included), which is precisely not the empty-stats mechanism (vacuuming weak teammates). F33 even keeps dan.k.memes in the top-15 and calls it "an elite, not empty-stats." The two findings should not both be cited as flagging these names. Trust F33's contestation logic over F32's weak-teammate + weak-team correlation here: a player who is the max over the entire contesting set is not an artifact of who their teammates are.


Verdicts

F33 (33_contestation_adjusted_value.md) → TRUST. The model is provably teammate-invariant (independently re-derived), the confound null (+0.03) is a genuine property of the method, the ranking is robust to ridge and to the nsba3 approximation, and the central conclusion — adjustment ≈ raw PPTF, the empty-stats confound is too small to reorder the board — is correct and now reinforced by an out-of-sample test it did not run. The one weak spot is the per-name empty-stats flag table (Test 3): those names are approximation-driven noise; the finding already demotes them to tiebreakers, so no action is mis-stated, but do not quote the specific "Euna Kim 41→89" type rows as findings.

F32 (32_usage_bias.md) → WATCHLIST, and trim the two named flags. Direction (bias real, small, two-sided, nets out at the board level; Spearman 0.97–0.98) is sound and agrees with F33. But: (a) the clean M1 teammate coefficient is spec-fragile (null at p≈0.5 in my rebuild and in every single season); (b) the within-player evidence is largely regression-to-mean, not a teammate penalty; (c) its two headline draft flags (ne, dan.k.memes) are contradicted by the better-designed F33 model. Keep the §0 shared- denominator mechanic and the board-level null; drop the ne/dan.k.memes carry flags.


How much should usage bias change the draft board?

Essentially zero. Concretely: 1. Keep raw pooled PPTF / VORP as the draft currency. It predicts next-season production on a different team as well as (slightly better than) the teammate-adjusted ability out-of-sample. The adjustment buys no drafting edge. 2. Do NOT build any team-strength / usage correction into player projection. Confound test +0.03, OOS advantage −0.04 — both null. A correction adds noise. 3. Drop F32's ne and dan.k.memes empty-stats flags. ne is genuinely elite head-to-head. 4. At most one within-tier tiebreaker tick for F33's two lineup-backed "suppressed" names (abcisosm5, akul) — and even that is optional, n-bound, and not draft-board-moving.


Caveats on THIS audit (load-bearing)

Scripts/artifacts: ran scripts/A33_contestation.py (reproduces). OOS + M1 + synthetic checks were ad-hoc (/tmp/oos_test.py, /tmp/synth.py during the session); core inputs data/processed/{player_season_master,combine_ability,tossups_long,games_meta,canonical_players}.csv, outputs/contestation_movers.csv, data/processed/usage_bias_32.csv.


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub