F93 — Skeptic audit of the usage / empty-stats bias work (F32, F33)
F93 — Skeptic audit of the usage / empty-stats bias work (F32, F33)
Date: 2026-05-30. Role: adversarial red-team of 32_usage_bias.md and
33_contestation_adjusted_value.md. Mandate (from David): is the measured teammate
effect real or just regression-to-mean / collinearity? Is the contestation model identifiable
at this n, or driven by a few games? Are the empty-stats flags robust? And the only test that
matters for drafting — does adjusting predict OUT-OF-SAMPLE (next-season PPTF on a different
team) better than raw PPTF?
Bottom line (lead with this): Both findings reproduce exactly, and their headline
conclusion is correct and if anything UNDERSTATED: the usage/empty-stats bias should change
the draft board essentially zero. I ran the out-of-sample test neither finding ran — predict
next-season PPTF on a different team — and the contestation adjustment does not beat raw
PPTF (leak-free Δ = −0.04 Spearman, 95% CI [−0.18, +0.04], P(adj better)=13%). The adjustment
buys nothing for drafting; raw PPTF is the right currency. The teammate-invariance of the F33
model is genuine (I re-derived the synthetic recovery independently; corr(error, teammate
strength) ≈ 0 even under an adversarial lone-stars-on-weak-teams design). But three softer
claims do not survive: (1) F32's clean teammate coefficient is not robust — it flips to
null (p≈0.5) under my reasonable spec and is null in every single season; the within-player
signal is mostly ordinary regression-to-mean, not a teammate penalty. (2) F33's specific
per-name "empty-stats" flags rest on majority/fully nsba3-approximated contesting sets with
7–18 wins — they are the noisiest cells in the dataset and cannot survive F33's own
lineup-only robustness refit. (3) F32 and F33 contradict each other on their own headline
names: F32 flags ne (#7) and dan.k.memes (#8) as the sharpest empty-stats risks; the
better-designed F33 model keeps ne at #7 (zero shift) and dan.k.memes in the top-15.
How much should usage bias change the draft board? Effectively nothing. Keep raw PPTF/VORP.
Do NOT apply a team-strength correction. Demote F32's ne/dan.k.memes "carry flags" to
noise. The two flags worth one tiebreaker tick are F33's abcisosm5 and akul (lineup-backed
"suppressed" cases) — and even those are within-tier.
Verdicts: F33 → TRUST (the invariance proof and the bounded-null are sound; the per-name flag table is the weak part). F32 → WATCHLIST → trim (direction OK, but the coefficient is spec-fragile and the two named flags are contradicted by F33).
What reproduces (full credit)
- F33
A33_contestation.pyruns clean and reproduces every headline number: Spearman(raw PPTF, adjusted) = 0.929; confound test corr(rank_shift, team win%) = +0.029, p=0.721; ridge sensitivity Spearman ≥ 0.976 for λ∈[0.25,4]; lineup-only refit Spearman 0.953; team-level adjusted +0.283 vs raw +0.394. Stratified rank_shift: weak teams −0.4, strong teams +1.5. All match the writeup to the digit. - F32's artifact (
usage_bias_32.csv) is internally consistent;obs_pptf= career-pooled PPTF (corr 1.00 with my recompute). Note: noA32script is committed — only the artifact — so F32's regressions are not independently re-runnable from source. I rebuilt M1 from raw.
Test 4 (the decisive one): does adjusting predict OUT-OF-SAMPLE better than raw PPTF?
Neither finding ran this, yet it is the only test of drafting value. F33's Spearman 0.93 is an in-sample re-description; "does the adjustment help you draft" = does it predict a player's PPTF next season on a different roster better than raw PPTF does.
The leakage trap (and why a naive version lies). Both adjusted_ability and
teammate_adj_pptf are career-pooled over ALL clean games including season t+1. Regressing
the pooled predictor on season t+1 PPTF gives a flattering Spearman +0.74 (adj) / +0.85 (F32
teammate-adj) vs +0.49 (raw) — but this is pure leakage; the predictor literally contains
the target's games. Do not believe it.
Leak-free test. I refit the contestation model on train seasons ≤ t only, then predicted each player's season-(t+1) PPTF on their new team. Pairs: 23 (TUF≥30 both seasons), 24 players, all on different teams (the exact confound-relevant set):
| Predictor (trained on ≤ t) | Spearman → next-season PPTF | Pearson |
|---|---|---|
| raw PPTF (prior season) | +0.523 | +0.417 |
| raw PPTF (pooled train) | +0.523 | +0.418 |
| contestation adj (train-only) | +0.471 | +0.444 |
Bootstrap of the difference (adj − raw, 2000 resamples): median −0.042, 95% CI [−0.181, +0.036], P(adj beats raw) = 0.13. The contestation adjustment is, within power, tied-to-slightly-worse than raw PPTF at predicting the next season on a new team.
Implication. This is the real test of David's drafting worry ("draft the empty-stats guy, he regresses on a better team"). If the empty-stats confound were real and large, the adjusted ability would predict the new-team season better than raw PPTF. It does not. The adjustment provides no drafting edge. This both (a) refutes any reading of F32/F33 as "use the adjusted board to draft" and (b) confirms F33's actual conclusion — raw PPTF is the currency and survives the empty-stats objection. n=23 is tiny (caveat below), but it is the honest number and it points clearly against the adjustment.
Test 1: is the teammate effect real, or regression-to-mean / collinearity?
Mostly the latter. Rebuilding M1 from raw (own combine θ + leave-one-out mean teammate θ, standardized PPTF):
| Spec | teammate-θ coef | p | n |
|---|---|---|---|
| F32 reported M1 (with SOS) | −0.039 | 0.027 | 125 |
| My rebuild (own θ + tm θ) | −0.051 | 0.527 | 106 |
| Team-mean θ incl. self (kills leave-one-out exclusion) | −0.065 | 0.445 | 114 |
| nsba2 only | −0.124 | 0.293 | 45 |
| nsba3 only | −0.111 | 0.339 | 59 |
The sign is consistently negative but the magnitude is small and the significance is
spec-fragile — F32's p=0.027 evidently leans on the SOS control + a specific sample; drop it
and the coefficient is null (p≈0.5), and it is null in every single season (matching
F32's own load-bearing caveat). Note corr(own θ, teammate θ) = −0.22 — the leave-one-out
exclusion mechanically gives the strongest player the weakest "teammate" mean, manufacturing
part of the negative slope.
The regression-to-mean confound is real and competes with the teammate story. Among the 26
within-player season transitions: corr(PPTF_t, ΔPPTF) = −0.40 — high scorers decline, low
scorers rise, regardless of teammate change. The confound's own signature is weak:
corr(team-strength change, ΔPPTF) = −0.18, and controlling for mean-reversion (PPTF_t), the
partial coefficient of team-strength-change on ΔPPTF is −0.13, p=0.55 — null. So within-player
there is no clean evidence that moving to a stronger team lowers PPTF beyond ordinary
mean-reversion. F32's "real but fragile" is the honest read; I'd push it toward "direction
suggestive, magnitude indistinguishable from mean-reversion + a leave-one-out artifact."
Test 2: is the contestation model identifiable / teammate-invariant at this n?
The invariance is genuine — this is the strongest part of the whole effort. I re-derived the synthetic recovery from scratch (my own simulator, Luce-generated winners):
- Baseline (10 seeds, 60 players, 12 teams): recovery Spearman 0.994, corr(estimation error, teammate strength) mean +0.023, range [−0.090, +0.179].
- Adversarial variant (explicitly assign strong players as lone stars on weak teams, NSBA-scale games): recovery Spearman 0.983, corr(error, teammate strength) mean −0.011, range [−0.209, +0.313] — centered on zero, both signs.
So the estimator does not launder team strength into a player's estimate, even when the data-generating process is designed to tempt it. F33's Result 3 holds. (My recovery Spearman 0.98–0.99 is higher than F33's reported 0.93 because my sim has more contests/less ridge bite; the load-bearing number — corr(error, teammate strength)≈0 — replicates cleanly.)
Did the adjustment re-introduce team strength another way? No — and the data confirms the opposite: corr(adjusted, team win%) = +0.111 vs corr(raw PPTF, team win%) = +0.093. The adjustment does not even reduce the (already weak) tie to team strength. The "teammate- invariant" claim is achieved, not laundered.
Identifiability caveat that bites the flags, not the ranking: 29/181 players have contest_n < 50 (heavily shrunk); the ranking is robust (ridge + lineup-only refit), but individual low-contest estimates are noise.
Test 3: are the empty-stats flags robust to sample/approximation? NO.
This is where F33 oversells. The six "empty-stats" names in its flag table rest on majority-or-fully nsba3 buzzed-roster-approximated contesting sets:
| F33 "empty-stats" flag | approx_frac | win_n |
|---|---|---|
| Euna Kim | 1.00 | 7 |
| Chris Wang | 1.00 | 8 |
| Aldric Benalan | 1.00 | 13 |
| michael | 1.00 | 18 |
| dan.k.memes | 0.77 | 53 |
| JoshuaW | 0.76 | 56 |
All six are >50% approximated; four are 100% approximated. F33's own Caveat 2 admits the buzzed-roster approximation "makes the contest look smaller and stronger than reality, biasing nsba3 abilities mildly upward." These flags therefore cannot survive F33's R4 lineup-only robustness refit — drop the approximated events and Euna Kim / Chris Wang / Aldric Benalan / michael have no data left. R4's Spearman 0.95 validates the aggregate ranking, not these specific names. The per-name empty-stats table is the noisiest cell in the analysis and should be treated as illustrative, not actionable. (F33 already hedges "none are early-round picks / use as tiebreaker" — correct, and this is why it stays TRUST.)
The two "suppressed" flags are better grounded: abcisosm5 (approx_frac 0, 179
contests, on a 0.73 team) is a clean lineup-backed case; akul (approx 0.39, 307 contests) is
partly approximated but lineup-anchored. These are the only two name-level signals I'd act on,
and only as a within-tier nudge.
The cross-finding contradiction (F32 vs F33 on their own headlines)
F32's draft action names ne (#7) and dan.k.memes (#8) as "the two clearest top-10
empty-stats risks… discount toward combine θ when projecting onto a stronger roster."
The F33 contestation model — which is the purpose-built, teammate-invariant test of exactly this — disagrees:
| name | raw rank | F33 adj rank | shift | approx_frac | win_n |
|---|---|---|---|---|---|
| ne | 7 | 7 | 0 | 0.00 | 61 |
| dan.k.memes | 8 | 15 | −7 | 0.77 | 53 |
ne has clean lineup data (approx 0), 61 head-to-head buzz wins, and zero rank movement
under the teammate-invariant model — i.e. ne wins buzzes against the whole room (opponents
included), which is precisely not the empty-stats mechanism (vacuuming weak teammates). F33
even keeps dan.k.memes in the top-15 and calls it "an elite, not empty-stats." The two
findings should not both be cited as flagging these names. Trust F33's contestation logic
over F32's weak-teammate + weak-team correlation here: a player who is the max over the entire
contesting set is not an artifact of who their teammates are.
Verdicts
F33 (33_contestation_adjusted_value.md) → TRUST. The model is provably teammate-invariant
(independently re-derived), the confound null (+0.03) is a genuine property of the method, the
ranking is robust to ridge and to the nsba3 approximation, and the central conclusion —
adjustment ≈ raw PPTF, the empty-stats confound is too small to reorder the board — is
correct and now reinforced by an out-of-sample test it did not run. The one weak spot is
the per-name empty-stats flag table (Test 3): those names are approximation-driven noise; the
finding already demotes them to tiebreakers, so no action is mis-stated, but do not quote the
specific "Euna Kim 41→89" type rows as findings.
F32 (32_usage_bias.md) → WATCHLIST, and trim the two named flags. Direction (bias real,
small, two-sided, nets out at the board level; Spearman 0.97–0.98) is sound and agrees with
F33. But: (a) the clean M1 teammate coefficient is spec-fragile (null at p≈0.5 in my
rebuild and in every single season); (b) the within-player evidence is largely
regression-to-mean, not a teammate penalty; (c) its two headline draft flags (ne,
dan.k.memes) are contradicted by the better-designed F33 model. Keep the §0 shared-
denominator mechanic and the board-level null; drop the ne/dan.k.memes carry flags.
How much should usage bias change the draft board?
Essentially zero. Concretely:
1. Keep raw pooled PPTF / VORP as the draft currency. It predicts next-season production on
a different team as well as (slightly better than) the teammate-adjusted ability
out-of-sample. The adjustment buys no drafting edge.
2. Do NOT build any team-strength / usage correction into player projection. Confound test
+0.03, OOS advantage −0.04 — both null. A correction adds noise.
3. Drop F32's ne and dan.k.memes empty-stats flags. ne is genuinely elite head-to-head.
4. At most one within-tier tiebreaker tick for F33's two lineup-backed "suppressed" names
(abcisosm5, akul) — and even that is optional, n-bound, and not draft-board-moving.
Caveats on THIS audit (load-bearing)
- The OOS test is n=23 (24 players, all team-changers), and 19 of 23 are nsba1→nsba2; the leak-free contestation refit on a single train season is itself noisy. The −0.04 point estimate is not precise — but its CI excludes a meaningful adjustment advantage, and the direction (no edge) is unambiguous. This is the binding limit: a true OOS adjustment edge smaller than ~0.15 Spearman is invisible here.
- My M1 rebuild omits F32's SOS control and uses n=106 vs its 125 (different combine-link filtering); the sign agrees, only the significance differs. The point stands that the coefficient is not robust to reasonable spec choices, which is the relevant claim.
- The synthetic recovery vindicates the method's invariance; it cannot prove the empirical abilities at NSBA's n are noise-free. Low-contest (<50) estimates remain unreliable.
- I did not re-audit the aggregate team-level PPTF→win validation (+0.62, F2). The usage confound is a player-attribution issue and washes out at the team level by construction (David's own framing) — out of scope and untouched.
Scripts/artifacts: ran scripts/A33_contestation.py (reproduces). OOS + M1 + synthetic
checks were ad-hoc (/tmp/oos_test.py, /tmp/synth.py during the session); core inputs
data/processed/{player_season_master,combine_ability,tossups_long,games_meta,canonical_players}.csv,
outputs/contestation_movers.csv, data/processed/usage_bias_32.csv.