NSBA Draft Analyticsembargoed · 2026-06-06

KEY_FINDINGS.md

Key Findings (canonical)

KEY FINDINGS — NSBA Draft Analytics (canonical, A30 synthesis)

Date: 2026-05-30. Lead analyst synthesis (now incl. second wave F27–F31 + skeptic audit F92). This file supersedes the running stub. Each finding gives: the claim (with honest magnitude/range), the evidence + source file, a confidence (High / Med / Low), and the load-bearing caveat. Where a revision (14b/15b/16b/17b) or red-team (90/91/92) changed a number, the revised value is canonical and the original is noted as superseded. PART C (F13–F17) folds in the bonus, duo-synergy, archetype, undervaluation, and field-depth waves, each carrying its F92 skeptic verdict.

Hard sample caps that bound every magnitude below. 41 team-seasons (nsba1=12, nsba2=13, nsba3=16); 207 game player-seasons; 180 clean games; 3 game seasons (nsba1/2/3) + 4 combine seasons; 216 draft picks. CS exists in games only in nsba2/nsba3. nsba3 PPTF uses an estimated-TUH proxy. Read every bolded magnitude as ±a lot; trust directions and rankings, not point estimates.

Two scoring worlds are never mixed: the combine (pyramidal, gameable) is a PREDICTOR only; game value (standard SB, PPTF) is the outcome currency.


PART A — General science-bowl roster-construction findings (embargoed writeup)

F1 — In natural (school) rosters, breadth + speed win decisively

F2 — Production rate (PPTF) is the north-star value currency

F3 — The combine is gameable; raw combine predicts game value at r≈0.6, no better

F4 — Coverage is table-stakes the draft auto-provides; "don't leave a hole" is the

only clean breadth effect, and even it does NOT survive a design rebuild - Claim: Drafted NSBA teams reach near-universal 6-subject coverage (95.1% cover all 6; 100% vs 100% at ≥8 games vs natural teams), including CS at 96.6%. Coverage is therefore a floor the draft already clears, not a differentiator. The realized "don't punt a subject" penalty (corr −0.54) does not reproduce when coverage is measured from roster DESIGN (drafted players' combine ability): design-punt → win% corr = −0.108 (p=0.56), null under every threshold and every talent control. - Evidence: Realized coverage corr(win%, #punts) −0.575 on the bridged sample; design coverage −0.108. The two measures barely agree (corr −0.119); only 1 of ~7 punters overlaps. All 5 realized punters are nsba3; SOS-exclusion shows ALL breadth/coverage variance is nsba3-only (nsba1/nsba2 teams all cover 6; subject_hhi→win% = −0.01 without nsba3). - Source: 18_natural_vs_draft.md, 16_generalist_specialist.md (superseded), 16b_coverage_design_revised.md (canonical), 15b_sos_robustness_revised.md, 23_draft_to_outcome.md; team_design_coverage_16b.csv. - Confidence: Med-High that the design effect is null (robust across 8 thresholds, 3 controls, all season slices); the realized penalty is a symptom of weak/short- schedule teams, not a roster-construction lever. - Caveat: Effective test = 24 nsba2/nsba3 team-seasons via an imperfect draft→game bridge. Actionable residue: don't actively punt a subject (the snake draft hands you coverage anyway); breadth is otherwise mostly a proxy for talent (red-team L2).

F5 — CS is worth somewhat more than a generic point, via confound-robust legs —

NOT 3.5×, and the scarcity premium is already arbitraged away - Claim: The original "CS ~3.5× / t=4.55" headline does not survive — it folded team bonus points into a per-player value claim. On toss-up points only, CS≈2.0× generic (t=2.08, p=0.047, bootstrap 95% CI [−0.28, 7.14] — cannot rule out CS being less valuable), collapsing to 1.4× under SOS and 1.2× under team fixed effects. What does hold are player-level legs: CS-mains convert CS 62.0% vs generalists 48.6% (χ²=7.00, p=0.008), and combine theta_cs → in-game CS r=0.60. CS is covered by 96.6% of teams, so coverage is table-stakes; depth/speed is a modest, uncertain tiebreaker. - Source: 17_cs_value.md (superseded), 17b_cs_value_revised.md (canonical), 15b_sos_robustness_revised.md, red-team 90 (H1/M1), 91 (T3). - Confidence: Med on direction (CS worth more, specialist not redundant, combine identifies CS via theta_cs); Low on any magnitude. - Caveat: 29 team-seasons, 2 seasons, 101 combine-linked players, 8-player CS-main group (lean on the 514-buzz pooled χ², not the split). The "premium erodes as adopted" decay test is underpowered because the field already covers CS. Do NOT reach.

F6 — Development is positive but small NET of mean-reversion, and does not transfer to

new entrants - Claim: Returners improve, but the naive "+0.43 z/edition" is partly regression-to- mean of a negatively-selected returning cohort (returners start ~0.18 z below the field on their first combine). Net of reversion: debiked combine theta +0.41 z (p=0.006, 95% CI [+0.13, +0.68]); game PPTF net +0.22 (p=0.028). Reversion inflated the pooled combine number only modestly (raw +0.46 → net +0.41) — it was not mostly reversion. - Source: 14_growth_aging.md (development half superseded), 14b_growth_net_revised.md (canonical), red-team 91 (T1). - Confidence: Low-Med (22–41 transitions; theta is within-season z, so "growth" is relative-rank climb; 24/30 combine transitions are into nsba4 which has no games). - Caveat: Two draft implications are load-bearing: (1) do NOT add a growth bump on top of reversion-shrink in the buy-low board — that double-counts the same climb; use one model proj = prior + β1·prior + β0 (β1≈−0.43, β0≈+0.41). (2) Growth does NOT transfer to a brand-new nsba4 entrant with no prior — price them off shrunk combine theta only.

F7 — Question difficulty varies by season/subject; playoff "difficulty" is confounded


PART B — NSBA4-specific draft implications

F8 — The field drafts the visible raw combine almost mechanically; ~13 recoverable

toss-points/slot by drafting on projected game value instead - Claim: Draft pick order tracks the raw combine tightly (|ρ| 0.67–0.75 every season — managers anchor on the visible score, not any bias correction). That combine only weakly predicts realized value (ρ ≈ 0.36–0.43), leaving a defensible ~13 toss-points/season/slot (~65 over a 5-round draft) recoverable by drafting on projected game value; oracle ceiling ~35. - Source: 21_adp_market.md; outputs/adp_table.csv. - Confidence: Low-Moderate. The directional results (field anchors on raw combine; combine is a noisy value signal; value-greedy beats combine-greedy) are consistent and significant across the two usable seasons. - Caveat: The honest ex-ante edge number is nsba2 alone (60 picks); nsba3's equal number is circular (same-season leakage). nsba1 nearly blind (11 links). Survivorship makes the combine look more predictive than it is, so the true edge is likely larger.

F9 — Snake-pick value is steeply convex (Jimmy-Johnson shape); the real loser's curse

is variance, not undervalued early picks - Claim: Round 1 retains only 46% of its value into Round 2; rounds 5–6 are nearly flat. No top-end flattening — the curve is convex (a₂=+3.5), so paying linear prices undervalues early picks. The genuine late-pick problem is variance: R1 bust rate 15% vs R6 73% — late picks are lottery tickets. Tradeable chart (pick 1 = 100): R1 100, R2 42, R3 22, R4 15, R5 12, R6 11. - Source: 22_pick_value.md; outputs/pick_value_chart.csv. - Confidence: Low-Moderate. The shape (steep-top, flat-tail) is robust across all three cuts (nsba2-gold, observed-only, toss_points) and is trustworthy. - Caveat: Per-slot values are small-n point estimates with wide bootstrap bands (pick 1: 1.5–3.7 WS). Chart prices slots; subject-coverage/CS fit and the 4–6 picks- per-team rule sit outside it.

F10 — Availability busts are real and common in the most recent format

F11 — The draft "matters" because talent wins, but the combine-seeded draft is a weak

handle on that talent - Claim: Realized roster scoring rate → win% is strong (r ≈ 0.68, clean subset, CI [+0.37, +0.85] excludes 0; star/rate-driven, not depth- or coverage-driven). But the arrow that makes the draft work — combine θ of who you drafted → realized roster PPTF — is weak (r ≈ 0.1–0.3, CIs include 0). Coverage/CS design did NOT separate winners (saturation). The design_has_cs "−0.34" is an n=4 fluke, not evidence CS hurts. - Source: 23_draft_to_outcome.md; draft_outcome_features.csv, draft_team_bridge.csv. - Confidence: Low-Moderate — direction trustworthy, magnitudes not. - Caveat: Clean test rests on nsba2 (n=12) + Jaccard≥0.5 subset (n=19); nsba1 effectively unlinkable (median roster Jaccard 0.11). Bootstrap CIs ~±0.3 wide.

F12 — Buy-low / fade board: a watchlist, not a measurement


PART C — Second-wave findings (F27–F31; skeptic-audited in F92)

F13 — Bonuses are 55–61% of the scoreboard, but carry NO exploitable per-player signal

F14 — Duo / pair synergy is a clean (bounded) NULL → team-building is ADDITIVE

F15 — Subject-archetypes are real but SOFT; only the broad elite tier reliably wins; CS is not a type

F16 — "Value-over-field" undervaluation: the new validation is CIRCULAR; no new edge

F17 — Field depth under playoff difficulty: lean broad elites (LOW confidence, nsba3-only)


Reconciliation ledger (what changed, and why this file trusts the revision)

Topic Original Canonical (revised) Why
CS value F17: 3.5×, t=4.55 F5 / 17b: ~2× raw → 1.2–1.4× net, p=0.047 Original folded team bonus into per-player claim; non-reproducible (90 H1, 91 T3).
Coverage rule F16: −0.54, ~7 win-pts/punt F4 / 16b: design corr −0.11, null Realized coverage is tautological with being good; design rebuild kills it (91 T2).
Growth F14: +0.43 z/edition F6 / 14b: net +0.41 (combine) / +0.22 (PPTF) Returners negatively selected; growth partly reversion; must not double-count (91 T1).
Top-25 nsba4 pool "5 of 25" 3 of 25 (Akhil, Rohan G, Kian) Artifact says 3 (90 M3).
Buy-low anchor "de-biked r=0.67" raw theta_overall r=0.74 The script used raw theta (90 M2).
Playoff factor 1.11 packet difficulty selection + difficulty, nsba3-only Confounded by surviving-field strength (91 T4).

Bottom line. The thesis spine — draft for projected production RATE (PPTF/VORP), treat coverage incl. CS as cheap insurance the draft auto-provides, and exploit the field's mechanical raw-combine anchoring — survives every red-team and the second wave firms it up: team-building is additive (no duo-synergy lever, F14/F92-TRUST), bonuses add no per-player signal (F13 — rate already captures them), the broad-elite type is the only one that wins (F15), the "new" undervaluation edge was circular (F16/F92-DISCARD, leaving F8's ~13 pts/slot as the only real edge), and playoff-robustness is a low-confidence lean toward broad elites (F17). The two most aggressive original actions (reach hard for CS; punt-avoidance as a lever) remain demoted to tiebreaker / table-stakes, and three new tempting levers (chase fit, value slow-but-smart bodies for bonuses, a second undervaluation edge) are rejected. Trust directions; distrust magnitudes; do not reach. Reproducibility gaps to close: the F31 and F27 scripts are not committed to scripts/.


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub