NSBA Draft Analyticsembargoed · 2026-06-06

90_methodology_redteam.md

90 — Methodology Red-Team (adversarial review of findings 10–20)

90 — Methodology Red-Team (adversarial review of findings 10–20)

Date: 2026-05-30. Reviewer role: methodology adversary. Method. Read every findings file in docs/findings/ plus the data dictionary and the underlying artifacts. Re-ran spot-check computations against the processed tables (.venv/bin/python, clean-game guardrail reconciles==True & not(0-0) = 180 games). Each issue cites the file + claim, says why it is wrong/overstated, and gives the fix.

Overall posture. This is a careful body of work. The guardrails (clean-game filter, two-scoring-world separation, identity joins) are respected almost everywhere, the foundational counts (180 games / 41 team-seasons / 207 player-seasons / 35-TUH stabilization) all reproduce exactly, and the authors self-disclose most small-n limits. The serious problems are concentrated in one finding (17, CS value) where a headline magnitude does not reproduce, plus a scatter of artifact-vs-text mismatches and direction-stated-as-strength overreach. No evidence of leakage into the draftable metrics, and the worst small-n claims are already flagged by the authors.


Severity-tagged issue list

HIGH

H1 — Finding 17 (CS value): the "~3.5×" / "t=4.55" head-to-head CS coefficient does not reproduce, and it silently mixes bonus points into a "per-buzz" value claim. - Claim: "CS points are ~3.5× as valuable as generic points… CS coef 0.0282 ± 0.0062 (t=4.55) vs non-CS 0.0081 ± 0.0017 (t=4.69)"; "mean pts/game cs = 9.1"; "robust to jackknife." - Why wrong/overstated: Re-running the stated regression (WLS win% ~ cs_ppg + noncs_ppg, weighted by games, 29 nsba2/3 team-seasons) on toss-up points only yields cs=0.0318 t=1.70 (not significant), non-cs=0.0181 t=2.64, ratio 1.76× — not 3.5×. The only way to get the finding's cs mean ppg = 9.1 is to add bonus points to CS (toss-up-only CS ppg = 3.33; toss-up+bonus = 8.73). But bonuses are team-answered and only reachable after a CS toss-up — folding them into "CS points" both (a) double-counts a team skill the finding 15 north-star explicitly refuses to attribute per-player, and (b) inflates CS's apparent leverage. Even with bonuses included I get cs=0.0148 t=2.25, ratio 1.67× — still nowhere near 3.5×/t=4.55. The reported CS t-stat (4.55) could not be reproduced under either specification. - Severity rationale: HIGH because "reach earlier than raw point totals suggest" is an actionable draft directive resting on a magnitude (3.5×) and significance (t=4.55) that my recomputation contradicts (true ratio ≈1.7×, CS coef not significant at .05). The direction (CS worth somewhat more) survives; the size and confidence do not. - Fix: Re-derive the regression from the artifact and publish the exact spec. Report toss-up-only CS value (the per-player-attributable quantity), state ratio ≈1.7× with a wide CI and a non-significant CS t, and explicitly exclude or separately label bonus points. Soften "reach earlier" to "modest, imprecise CS premium; don't punt CS." The conversion-gap result (H1-adjacent) is more defensible — see M1.

MEDIUM

M1 — Finding 17 §1: the CS-main vs generalist conversion table numbers don't reproduce and the χ² is overstated as p=0.011 on a fragile n=8 group. - Claim: CS-mains 8 players / 129 buzzes / 0.620 conv; generalists 385 buzzes / 0.486; χ²=6.47, p=0.011. - Why: My recomputation (CS buzzes in clean nsba2/3 games, main subject = argmax of per-subject pts) gives 106 CS-main buzzes (conv 0.670) vs 426 generalist (conv 0.469), χ²=13.63, p=0.0002 — more significant but with different counts, meaning the finding's exact buzz/player partition (it adds a ≥5-correct "rotation player" filter and an "attempted ≥1 CS buzz" gate) is not transparently specified. The qualitative result (CS-mains convert higher) is robust to my respec, so this is MEDIUM not HIGH. - Fix: State the exact cohort filter; the conclusion stands but the table is irreproducible as written. Note the "CS-main" group is 8 players — too thin to lean on; the pooled-buzz χ² (which the finding correctly calls "the more stable version") should be the headline, not the player-level split.

M2 — Finding 19 (buy-low): the advertised "de-biked" price is actually the non-de-biked theta; the r=0.74 anchor is the raw-IRT correlation. - Claim: "z(price) = z of their nsba4 combine theta (de-biked IRT, finding 10)" and "The two correlate r=0.74 across these 18." - Why: For the 18 returning nsba4 entrants, corr(career PPTF, theta_overall) = 0.74 but corr(career PPTF, debiked_overall) = 0.669. So the 0.74 the finding cites — and presumably the residual line every mispricing is measured against — is the raw theta, contradicting the finding's repeated insistence that the price is the de-biked signal (the whole D7 rationale). The mispricings are computed off whichever column the script actually used; the text and the number disagree. - Fix: State which theta column the residual uses and make it consistent. If de-biked is intended, the anchor r is 0.67 and every residual shifts; if raw is intended, drop the "de-biked" framing for Track A. Either way the Track-A board is n=18 with 1–4-game reads near noise (the finding does flag this).

M3 — Finding 13 (player value): "Only 5 of the top-25 carry an nsba4 combine row" is contradicted by the artifact — it is 3. - Claim: stated twice ("Only 5 of the top-25…", "5 of the top-25 have an nsba4 combine row"). - Why: outputs/player_value_table.csv top-25 has has_nsba4_combine.sum() == 3 (Akhil Batchu, Rohan G, Kian Dhawan). - Severity: MEDIUM because this is the central caveat of the whole value board ("the board is historical talent, not the nsba4 pool"). Understating the pool-thinness by stating 5 when it is 3 makes the live draft pool look less thin than it is — and it actually strengthens the author's own caveat. A factual slip in the load-bearing limitation. - Fix: Change 5 → 3 in both places.

M4 — Multiple-comparisons / "park factors" presented as composable point estimates (Finding 20). - Claim: season factors 1.089/1.100, playoff 1.111, six subject factors, "roughly composable (season × playoff × subject)." - Why: That is 3 + 6 = 9 difficulty cells estimated and then multiplied, with no joint CI and no correction for having scanned subjects for "hardest." The playoff factor rests on nsba3-only, n=265 playoff questions, and the finding itself concedes conversion "conflates question difficulty with field strength" — i.e. the playoff factor may be stronger opponents, not harder packets. Presenting a single multiplicative product invites users to over-trust a quantity that is one season of confounded data. - Why only MEDIUM: the season/subject decomposition is a clean logit on 4,155 questions and the direction (nsba3 easiest, CS hardest, playoffs harder) is well-supported; my recomputation matched the conversions (reg 0.838 / playoff 0.755 / season 0.783/0.738/0.819). The authors do caveat the field-strength confound. - Fix: Attach CIs to each factor; flag that the product compounds uncertainty; restate playoff factor as "nsba3-only, possibly field-strength."

M5 — Finding 11/13: PPTF reliability (M≈35, ρ=0.81 at n=150) is real but the per-category M's lean on a self-admitted ±25% proxy, and several findings then use those category numbers without re-flagging the proxy. - Claim: per-category M's (CS≈69, Energy≈72, etc.) used to justify "shrink CS/Energy almost entirely to the mean," then per-category VORP in finding 13 and per-category combine→game in finding 12 build on category rates. - Why: finding 11 honestly says category within-variance is a proxy (category correct/neg splits absent from the master) good to ±25% ranking-grade. That is fine in finding 11, but downstream per-category VORP (finding 13 §4) and per-category transfer (finding 12 §3) inherit that imprecision without re-stating it. Overall M=35 reproduces tightly (my pooled ≈35.6; nsba1/2 ≈27–31, nsba3 inflated ≈64 — matches the finding's "nsba3 TUH-estimation inflates it" note). - Fix: Carry the ±25% category-proxy caveat forward into 12 §3 and 13 §4; treat per-category VORP as ordinal, not cardinal.

LOW

L1 — Finding 14: "% improving" for combine deltas reported as 57% but my recompute gives 73%. - The headline t=3.36, p=0.002, +0.43 z/edition reproduce exactly (n=37). Only the secondary "% improving" differs (likely a different sub-denominator), so LOW — but it should be reconciled. The PPTF-growth side (+0.086, p=0.10) is correctly labeled non-significant.

L2 — Finding 14 / 15 / 16: "breadth wins" is repeatedly shown to be entangled with talent and the authors say so — but the TL;DRs still lead with breadth as a lever. - corr(player toss_points, #subjects) = +0.76 and the pairing bonus "dissolves" under an HHI control; HHI's win-correlation (−0.46) doesn't survive once production rate is in (partial −0.07 in 15). The bodies are honest; the bullet-point TL;DRs ("build for breadth") slightly oversell a second-order, collinear effect. LOW because the nuance is present one paragraph down. Fix: lead the TL;DR with "breadth is mostly a proxy for talent; the only clean breadth effect is the don't-leave-a-hole threshold (finding 16)."

L3 — Finding 16 punt effect: clean and reproduced (corr −0.537, p=0.0006), but it rests on n=5 punting teams all in nsba3. - The finding discloses this ("punt variation is entirely in nsba3" at the main threshold). LOW only because the headline win-rate contrast (52% vs 28%) is driven by 5 teams in one season; the higher-threshold robustness (which spreads punts across seasons) is the better evidence and should be foregrounded over the 5-team split.

L4 — Finding 13 RAPM and Finding 15 point_margin: correctly labeled non-draftable/null. No issue — flagging as a positive control: the authors explicitly refuse to use point_margin (mechanical) as a draft input and ship RAPM (corr 0.19) only as a diagnostic null. This is exactly right and should be preserved.

L5 — Finding 18: 95.1% / 100%@≥8g / 81% natural all reproduce exactly. No issue; the games-played confound is honestly the headline caveat. Positive control.


Cross-cutting methodological notes

  1. Small n is pervasive and mostly disclosed. 41 team-seasons, 207 player-seasons, 18-player buy-low Track A, 8-player CS-main group, n=2 suspected-tanker group (finding 12), 5 punting teams (16). The authors flag these almost everywhere. The danger is the TL;DR/headline layer, which repeatedly states directionally-true findings with more confidence than the bodies support (H1, M1, L2). Treat every bolded magnitude as ±a lot.
  2. No leakage into draftable metrics found. PPTF, combine theta, coverage are all built from own-player buzzes / pre-game tryout; point_margin and points_against are explicitly quarantined as outcome-only (15). Good.
  3. Scoring worlds kept separate — combine (pyramidal) never enters value stats; combine appears only as a predictor (10, 12, 13, 19). The one violation is H1's bonus-into-CS-value, which mixes a team-attributed quantity into a per-player-value claim.
  4. nsba3 estimated-TUH proxy correctly threads through every PPTF-based finding as a caveat (11, 12, 13, 15, 17, 19, 20). It is the single biggest structural weakness (nsba3 carries the playoff test, most CS data, and a large share of value), and the authors are consistent about it.
  5. Identity discipline is good. The 65 needs_review rows are kept raw, not force-matched; game joins are 0-loss within season; the sparse tanking_risk flag is labeled nsba4-era and qualitative. No finding leans on the unresolved 65. Finding 12's tanker test (n=2) and finding 19's Track B (intel-only, conf 0.40) correctly refuse to dress qualitative chatter as measurement.

Trust verdict per analysis

finding verdict one-line reason
10 combine IRT Trust validations reproduce; de-biking gain correctly localized to nsba1; honest about partial de-biking.
11 reliability Trust M≈35 and per-season pattern reproduce; two methods agree; ±25% category proxy disclosed.
12 combine→game Trust w/ caveats core r≈0.62 / CV R²≈0.35 solid; stress tests (n=2 tankers, 12 playoff games) are directional only — already labeled.
13 player value Trust w/ caveats board reproduces exactly; fix the "5→3 top-25 nsba4" slip (M3); compressed VORP honestly flagged.
14 growth/aging Trust +0.43 z/edition (p=0.002) and mean-reversion reproduce; Gideon claims correctly reported as unsupported; reconcile the 57%/73% slip (L1).
15 north-star Trust PPTF univariate reproduces exactly; outcome metrics correctly quarantined; LOSO R² point estimates honestly called noisy at n=41.
16 generalist/specialist Trust w/ caveats punt corr reproduces (−0.54); but the headline split is 5 nsba3 teams — foreground the threshold-robustness, not the 5-team contrast (L3).
17 CS value DO NOT TRUST the magnitudes the 3.5× / t=4.55 headline does not reproduce (≈1.7×, CS t not significant) and silently mixes team bonus points into a per-player value claim (H1). Direction (CS worth somewhat more, specialist not redundant) survives; size/significance do not. Re-derive before acting.
18 natural vs draft Trust 95.1% / 100%@≥8g / 81% all reproduce; games-played confound is the honest headline caveat.
19 tanking/buy-low Trust as a watchlist, not a measurement reconcile the de-biked-vs-raw anchor (M2, r=0.74 is raw not de-biked); Track A n=18 with 1–4-game noise reads; Track B is intel priors (conf 0.40) — all disclosed. Use for shortlisting, not point estimates.
20 difficulty Trust the direction, not the composed factors season/subject logit clean and reproduced; playoff factor is nsba3-only/n=265 and possibly field-strength; attach CIs and stop multiplying factors without joint uncertainty (M4).

Bottom line. 9 of 11 findings are trustworthy with their stated caveats. Finding 17 (CS value) needs its central magnitude re-derived — the actionable "reach earlier for CS" rests on an unreproducible 3.5×/t=4.55 that drops to ≈1.7× and non-significance under the documented spec. Two findings (13, 19) carry small factual/labeling slips (top-25 count; de-biked-vs-raw anchor) that should be corrected. The recurring soft failure is TL;DR overreach: directionally-correct results stated with more certainty than the small-n bodies support — fix by demoting magnitudes to ranges and leading with the entanglement/threshold caveats the bodies already contain.


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub