NSBA Draft Analyticsembargoed · 2026-06-06

23_draft_to_outcome.md

23 — Draft → Team Outcome (D4): does the draft *matter*?

23 — Draft → Team Outcome (D4): does the draft matter?

Question. This is the end-to-end validation of the whole thesis. Link drafted rosters (draft_rosters.csv) to realized standings (team_season.csv win_pct) and ask: did teams that drafted more projected value / better coverage / earlier-combine players actually win more? Is the draft → wins chain signal, or noise? And which roster-construction features (top-end talent, depth, coverage, CS presence) best separated winners from losers historically?

TL;DR — three claims, in descending order of how much the data supports them.

  1. Talent wins games. A roster's realized toss-up production rate is the single best predictor of win% — corr(win%, mean drafted-player PPTF) = +0.68 on the clean-linkage subset (90% bootstrap CI [+0.37, +0.85], excludes 0), and the star/top-player rate separates winners from losers (t=2.65, p=0.017). So the roster's talent → wins arrow is real and reasonably strong, even at this n.

  2. But the combine-seeded draft only loosely controls that talent. The arrow that is supposed to make the draft workcombine ability of who you drafted → how good your roster actually was — is weak and not distinguishable from zero: corr(top-2 drafted combine θ, realized roster PPTF) = +0.13 to +0.27 (CIs include 0). The draft's input signal (a gameable combine) is a noisy handle on the output (wins). Top-end combine θ does point at win% (nsba2 r=+0.53) but with a CI [−0.10, +0.85] that crosses zero — directionally right, statistically unproven.

  3. Coverage / CS design did NOT separate winners from losers — because every drafted team already had it. Design coverage (r=+0.07 with win%), and the apparent "no-CS-answerer → wins" blip, are saturation artifacts, not levers. This confirms finding 18: the snake draft makes broad coverage table stakes, so coverage cannot be a differentiator. The differentiator is how much production rate you concentrate, not breadth.

The honest verdict: the draft "matters" in that rosters with more scoring talent win — but the historical record can only weakly attribute that to the combine-driven draft choices, partly because the cleanest test exists in one season (nsba2, n=12) and the combine→realized link is loose. Lead with caveats; this is a 12-to-34 team-season result built on top of a fragile draft↔game team-name bridge.


The binding constraint: drafted rosters barely link to game rosters (except nsba2)

Before any outcome test, the draft team (a captain handle, e.g. Coby, polaris, karthik) must be mapped to the in-game team name (e.g. mount ellis, okc). They are not the same strings. The only reliable bridge is canonical_id overlap: match each game team to the drafted roster it shares the most players with (best Jaccard). Doing that exposes how recoverable each season's draft→outcome link actually is:

season drafted players who appear in any clean game game-teams mapped median Jaccard(draft, game roster)
nsba1_2022 11 / 59 (19%) 8 / 12 0.11
nsba2_2023 59 / 59 (100%) 12 / 13 1.00
nsba3_2025 53 / 84 (63%) 14 / 16 0.50

Everything below is therefore reported four ways: all 34 mapped team-seasons, the clean-linkage subset (Jaccard ≥ 0.5, n=19), nsba2-only (n=12, gold), and per-season. Trust the nsba2 and Jaccard≥0.5 columns; treat the pooled-34 as contaminated by the nsba1 near-random mapping.


Method


Results

1. Realized roster talent → wins (the strong, robust arrow)

subset n corr(win%, mean drafted PPTF) corr(win%, top drafted PPTF) corr(win%, total pts)
nsba2 (gold) 12 +0.530 (p=0.08) +0.598 (p=0.04) +0.598 (p=0.04)
Jaccard ≥ 0.5 19 +0.678 (p=0.001) +0.594 (p=0.007) +0.549 (p=0.015)
nsba3 14 +0.657 (p=0.011) +0.507 (p=0.064) +0.628 (p=0.016)
all mapped 34 +0.405 (p=0.017) +0.355 (p=0.039) +0.276 (p=0.11)

Read: rosters that actually scored faster won more. This is the part of the thesis that holds up.

2. Combine-design → realized talent (the weak arrow that should make the draft "work")

This is the crux of "does the draft matter vs noise." The draft's job is: pick high-θ players → get a high-production roster → win. Arrow 2 (talent→wins) is strong (§1). But arrow 1 — combine θ of who you drafted → how good your roster actually was — is weak:

subset n arrow 1: corr(top-2 drafted θ, realized roster PPTF) arrow 2: corr(realized PPTF, win%)
all mapped 34 −0.06 (p=0.76) +0.41 (p=0.02)
Jaccard ≥ 0.5 19 +0.13 (p=0.60) +0.68 (p=0.00)
nsba2 (gold) 12 +0.27 (p=0.39) +0.53 (p=0.08)

The combine→realized-PPTF correlation (+0.13 to +0.27, all CIs spanning 0) is much weaker than finding 12's player-level combine→game r≈0.62 — expected, because (a) this is at the roster-aggregate level (averaging washes out, and a team is only as linked as its noisiest member), and (b) the combine is gameable (the guardrail). So the historical record shows that drafting high-θ players gave only a loose grip on realized roster quality.

Does top-end combine θ point at wins at all? Directionally yes, but unproven:

subset n corr(win%, design_top1_θ) bootstrap 90% CI
nsba2 (gold) 12 +0.533 [−0.10, +0.85]
Jaccard ≥ 0.5 19 +0.312 [−0.17, +0.61]
nsba3 14 −0.154
all 34 +0.246

Top-end combine θ → win% is +0.53 in the gold-standard season but flips slightly negative in nsba3, and every CI crosses zero. So "draft the highest-θ players and you'll win" is not refuted, but not demonstrated in the historical sample — it is exactly as strong as the combine→game link is loose (§ arrow 1).

3. Coverage / CS design did NOT separate winners (saturation, per finding 18)

feature corr with win% (all, n=34) winner mean loser mean
design_n_cov (# subjects w/ credible answerer) +0.07 (p=0.68) 4.25 3.96 (t=0.66, p=0.52)
design_has_cs −0.34 (p=0.05) — see below
realized n_subjects_scored +0.56 (p=0.001)

4. Per-season summary of "what separated winners"

signal nsba1 (n=8, unlinked) nsba2 (n=12, gold) nsba3 (n=14, partial)
realized mean PPTF → win +0.27 (ns) +0.53 +0.66
top combine θ → win +0.40 (ns) +0.53 −0.15
design coverage → win +0.51 (ns) +0.14 −0.14
realized HHI → win −0.17 +0.09 −0.63

The only signal positive and material in both trustworthy seasons (nsba2, nsba3) is realized scoring rate. Top combine θ works in nsba2 but not nsba3; the HHI/breadth effect is an nsba3-only phenomenon (matching findings 16/18 and red-team T2's "the coverage result is essentially nsba3").


Interpretation — does the draft matter?

Yes, conditionally, and weakly attributable to the combine. Decompose the thesis chain:

   combine θ of picks  --(weak, r≈0.1–0.3)-->  realized roster scoring rate  --(strong, r≈0.5–0.7)-->  wins
        [arrow 1: the draft's predictive grip]                [arrow 2: talent wins]

Actionable for nsba4: prioritize projected production rate (PPTF / VORP from finding 13) over combine rank, over depth, and over coverage. The draft's job is to maximize roster scoring rate; combine θ helps rank-order candidates but is a weak predictor of realized output, so blend it with game-history PPTF where available (finding 12/13) and do not over-pay for it. Coverage and a CS body are cheap insurance the draft provides almost automatically — don't reach for them at the expense of rate.


Limitations & caveats (lead with these)

Artifacts


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub