Synthesis — Master Narrative
SYNTHESIS — NSBA Draft Analytics
Lead analyst, A30 synthesis pass. Date: 2026-05-30.
Written for a skeptical reader who will try to poke holes. Every magnitude here is a
small-n point estimate; the contribution is a ranked, confound-audited set of
directions, not precise numbers. Two adversarial red-teams (90_methodology_redteam.md,
91_causal_redteam.md) were run against the first modeling pass; their flags drove four
revisions (14b, 15b, 16b, 17b), and this synthesis down-weights everything they
flagged and reports final confidence honestly. Where an original finding and its
revision conflict, the revision is canonical.
1. The question and the data
NSBA is a science-bowl league run like the NBA: players take a 24-question combine (6 categories × 4 pyramidal questions, scoring +4/+3/+2 for earlier correct buzzes, −1 for a wrong buzz), then 14 teams build rosters via a snake draft. We have:
- Game scoresheets for 3 seasons (nsba1_2022, nsba2_2023, nsba3_2025), reconciled to
180 clean games (
reconciles==True & not 0-0). This yields 41 team-seasons (12/13/16) and 207 player-seasons. - Combine data for 4 seasons including nsba4_2026 (the draft target, which has no games yet).
- 216 draft picks across 3 drafts (
draft_picks.csv), plus eligibility, rosters, and a natural Science-Bowl baseline (scibowl_*, 84 teams / 13,632 buzzes, Energy not CS).
The binding constraints, stated up front (these bound everything that follows):
- n is tiny everywhere. 41 team-seasons, 3 game seasons, 18-player buy-low track, 8-player CS-main group. Treat magnitudes as ±a lot.
- CS exists in games only in nsba2/nsba3. Every CS conclusion is a 2-season result.
- nsba3 PPTF uses an estimated-TUH proxy (the paired scoresheet records no per-player tossups-heard). nsba3 is simultaneously the noisiest-measured season, the only season carrying coverage/punt and playoff variation, and the season most like the nsba4 target (red-team T8). Where possible we use buzz-count-denominated rates (conversion, points/ game) that don't depend on the TUH denominator.
- Two scoring worlds, never mixed. The combine is pyramidal and gameable ("biking" = late-MC guessing is +EV); it is used only as a predictor. Game value uses standard SB scoring. The guardrail held everywhere except the one place the red-team caught (CS-bonus-into-CS-value, now fixed).
- Identity discipline. 65 needs_review identities are kept raw, not force-matched; the draft→game team-name bridge is by canonical-id roster overlap and is imperfect (nsba1 median Jaccard 0.11 — effectively unlinkable; nsba2 1.00; nsba3 0.50).
Methods. PPTF (points per toss-up faced) is the rate stat everything regresses toward
(the Marcel → stabilization → regress-to-mean → aging spine from sports analytics, A6).
PPTF stabilizes at ≈35 faced toss-ups (reliability 0.5), ~1.5 games; projections are
empirical-Bayes shrunk toward the grand mean with that padding. Combine ability is a
graded-response IRT model (girth.grm_mml_eap) giving per-player EAP θ with posterior-SD
shrinkage, fit only on combine data and validated out-of-sample against game PPTF.
2. What we found
2.1 Value = production rate (PPTF), not roster shape
The single best draftable predictor of standings is the team's toss-up production rate,
PPTF: corr(PPTF, win%) = +0.624, positive every season (nsba1 +0.535 / nsba2 +0.561 /
nsba3 +0.752), and it survives both an SOS adjustment (+0.576) and dropping nsba3 entirely
(+0.50). point_margin correlates higher (+0.80) but is a mechanical outcome and is
quarantined as an illegal draft input. PPG_total edges PPTF (+0.73) only by folding in
team-answered bonus conversion, which cannot be pushed down to individuals — a
structural reason to prefer PPTF as the per-player currency. Player-concentration metrics
(top-2 share, player-HHI, depth) are ~0 with win%: how points are spread across the
roster barely matters; total production rate dominates. Win-shares decompose this exactly
(player shares sum to team wins, error 0.0). [15_north_star.md, 15b]
2.2 The combine is a gameable, r≈0.6 seed
Combine ability moderately predicts game PPTF — pooled r ≈ 0.62, ridge CV R² ≈ 0.35,
so ~⅔ of game-value variance is unexplained by the tryout. Raw combine is the best
single linear predictor; the IRT θ and de-biked θ are close to it. De-biking earns its keep
only on the heavily-biked nsba1/Energy era (debiked r=0.56 vs raw 0.24, nearly tripling
it) and slightly hurts on clean per-question seasons. Per-category signal genuinely
transfers for all six subjects (r 0.38–0.54; math/ESS best, physics weakest), so the
combine is not a single g-factor. [10_combine_irt.md, 12_combine_to_game.md]
The selection caveat the red-team forced (T5): the linked cohort is players who took the combine and showed up to ≥2 games. High-combine no-shows and tankers — the exact players who break the combine→value link — are missing, so r≈0.6 is optimistic for the unproven pool you actually draft from. Prefer the de-biked combine for unproven players.
2.3 Coverage is table-stakes the draft already provides
Drafted NSBA teams reach the same near-universal 6-subject coverage that school teams reach
organically: 95.1% cover all 6 (vs 81% natural), and 100% vs 100% once schedule is
held fixed (≥8 games). CS specifically is covered by 96.6% of nsba2/nsba3 teams. The
snake draft over a pool with ≥1 capable answerer per subject mechanically fills coverage
without anyone optimizing for it. [18_natural_vs_draft.md]
So "don't punt a subject" cannot be a differentiator — it is a floor everyone clears. The
original F16 claimed a realized punt penalty (corr(win%, #punts) = −0.54, ~7 win-pts/punt).
The red-team (T2) was right that this is tautological with being good, and the revision
confirms it: when coverage is rebuilt from roster design (drafted players' combine
ability) instead of realized in-game production, the correlation collapses to −0.108
(p=0.56), is null under every threshold and every talent control, and the design and
realized punt measures barely agree (corr −0.119; only 1 of ~7 punters overlaps). All 5
realized punters are nsba3 short-schedule/weak teams; the SOS cut shows ALL breadth
variation is nsba3-only (nsba1/nsba2 teams every cover 6 subjects). The only honest
breadth rule that survives is: don't actively punt a subject — and the draft makes even
that nearly automatic. [16_generalist_specialist.md superseded → 16b, 15b, 23]
2.4 CS: a modest, confound-robust premium — not 3.5×
This is the finding that changed most. The original "CS points are ~3.5× as valuable,
t=4.55" was non-reproducible and confounded (red-team H1/T3). The error: it fed the
regression CS toss-up points + team bonus points (~9.1 ppg), but bonuses are
team-answered and not per-player attributable. Re-derived on toss-up points only, the
CS/generic ratio is ≈2.0× (CS t=2.08, p=0.047, bootstrap 95% CI [−0.28, 7.14] —
4.2% of bootstrap draws put the CS coefficient ≤ 0, so we cannot even rule out CS being
less valuable). And most of even that 2× is "we were the better team": adding an opponent-
strength (SOS) control drops it to 1.38×, team fixed effects to 1.22×.
corr(share_cs, win%) = +0.42 is the reverse-causation flag — fast/strong teams win the
buzzer race in every subject, most in CS (the hardest, most-dead subject). [17b, 15b]
What does hold are two player-level legs not subject to the team buzzer-race confound: - CS-mains convert CS better: 62.0% vs generalists 48.6% (pooled χ²=7.00, p=0.008, 514 buzzes). A generalist roster does not convert CS for free. - The combine finds CS honestly: combine theta_cs → in-game CS correct r=0.60 (vs 0.44 for theta_overall). This is the single most operationally useful CS result — you can trust theta_cs to identify a CS answerer pre-draft.
And the scarcity premium is already arbitraged away: the field covers CS 96.6% of the time, opponent CS supply ranges only 1.8–6.0 ppg, and the "premium erodes as adopted" decay test is underpowered precisely because everyone already covers CS. The reconciliation of F17 vs F18 the red-team demanded (T7): coverage is table-stakes the draft auto-provides; the only place leveraged CS value could live is depth/speed beyond bare coverage, and within 2 seasons that component is small, imprecise, and not separable from team strength.
2.5 Growth is positive but small net of mean-reversion
Returners improve, but the naive "+0.43 z/edition" is partly regression-to-mean of a
negatively-selected cohort: returners start ~0.18 z below the field on their first
combine (red-team T1). The right decomposition regresses the within-player change on the
prior level in one model; the intercept at prior=0 is the net development: debiked
combine θ +0.41 z (p=0.006, CI [+0.13, +0.68]); game PPTF net +0.22 (p=0.028).
Reversion inflated the pooled combine number only modestly (raw +0.46 → net +0.41) — it was
not mostly reversion. Two load-bearing draft implications: (1) the buy-low board must
not add a growth bump on top of a reversion shrink — that double-counts the same
climb; use one model proj = prior + β1·prior + β0 (β1≈−0.43, β0≈+0.41). (2) Growth does
not transfer to a brand-new nsba4 entrant with no prior to revert. [14b]
2.6 The market: the field anchors on the visible raw combine
Draft pick order tracks the raw combine tightly (|ρ| 0.67–0.75 every season) — managers
anchor on the visible score, not on any bias correction. That raw combine only weakly
predicts realized value (ρ ≈ 0.36–0.43). The exploitable gap: drafting on projected game
value instead of raw combine recovers a defensible ~13 toss-points/season/slot
(~65 over a 5-round draft), with an oracle ceiling ~35. The honest ex-ante number is nsba2
alone; nsba3's equal figure is circular (same-season leakage). Survivorship makes the
combine look more predictive than it is, so the true edge is likely larger. [21]
2.7 Pick value and availability
Snake-pick value is steeply convex (classic Jimmy-Johnson): Round 1 retains only 46%
of its value into Round 2; the curve is convex (a₂=+3.5), so paying linear prices
undervalues early picks. The real loser's curse is variance: R1 bust rate 15% vs R6
73% — late picks are lottery tickets. Tradeable chart (pick 1 = 100): R1 100 / R2 42 /
R3 22 / R4 15 / R5 12 / R6 11. [22]
And availability is a real, recent-format risk: 37% of nsba3 drafted players never played
a clean game (vs 0% in nsba2); even 3 of 14 nsba3 captains never fielded. Earlier picks
have better availability (R1≈85% vs R6≈23%, pooled spearman −0.36). The discord
availability_flag tracks realized availability monotonically (none 0.86 / some 0.77 /
concern 0.49, n=16). CS-mains show NO availability penalty. [24]
2.8 Does the draft matter? Talent wins; the combine is a blurry instrument for it
End-to-end (red-team-style validation): the chain is
combine θ of picks --(weak, r≈0.1–0.3)--> realized roster scoring rate --(strong, r≈0.5–0.7)--> wins.
Arrow 2 is solid (clean-subset r=0.68, CI [+0.37, +0.85] excludes 0; the league is
star/rate-driven, not depth- or coverage-driven). Arrow 1 is loose — the gameable
combine is a noisy handle on realized roster quality. So talent matters; the
combine-seeded draft is a blurry instrument for acquiring it, which is exactly why
drafting on game-history PPTF where available beats drafting on combine. [23]
3. The draft strategy that follows
- Draft for projected production RATE (PPTF / VORP, finding 13), not combine rank. Blend combine θ with game-history PPTF where a player has tape; for unproven players use shrunk (and de-biked) combine θ. The field will draft the visible raw combine board (|ρ|≈0.7) — fade combine-inflated early names (Vishnu, Kian-as-#1), target combine- underrated mid/late values. Concentrate the ~13-pt/slot edge in rounds 3–5, where the field's anchoring leaks most; round 1 is efficient.
- Treat coverage and a CS body as cheap insurance the draft hands you almost automatically — don't reach for them. Don't actively punt a subject; secure ONE credible CS answerer (use theta_cs, the r=0.60 honest signal, to pick a good one); then stop — a 2nd/3rd CS specialist is redundant. Do not pay a 3.5× premium or "reach earlier than raw totals suggest." CS depth/speed is a tiebreaker, not a thesis.
- Prioritize aggregate scoring rate over depth and breadth. Wins are star/rate-driven; the drafted 3rd-man and design coverage are ~0 with win%. Use the pick-value chart for quantity of value (early picks buy reliability, not just ceiling) and the VORP/ buy-low boards for which player.
- Discount for availability and double-count discipline. Multiply projected per-season
contribution by
max(reliability_prior, flag_prior); for new entrants use the round- conditional base rate (R1≈0.85 … R6≈0.23). On the buy-low board, use ONE reversion+growth model — never stack a separate growth bump on a reversion shrink — and never credit a first-time entrant with returner "development."
4. What is SOLID vs SHAKY
SOLID (trust the direction; ranking robust): - PPTF is the right value currency (survives SOS, nsba3-exclusion, positive every season). [F2] - The combine is a moderate, gameable r≈0.6 seed; raw is best; de-biking is nsba1-local. [F3] - The draft mechanically provides 6-subject coverage including CS (96.6%). [F4/F11] - The field anchors on raw combine; value-greedy beats combine-greedy. [F8] - Pick-value shape is steeply convex; late picks are high-variance lottery tickets. [F9] - nsba3 availability busts are real (37% never played). [F10] - Talent (realized rate) → wins is the one arrow with a CI that excludes 0. [F11]
SHAKY (down-weighted; magnitudes distrusted — per the red-teams): - The CS "3.5×" is dead (90 H1, 91 T3). Real ≈2× raw → 1.2–1.4× net of SOS, p=0.047, CI spans a factor of several. Lean on the player-level conversion gap + theta_cs only. [F5] - The "don't punt → ~7 win-pts" rule is dead as a design lever (91 T2). Realized coverage is tautological with being good; the design rebuild is null. [F4] - Growth shrinks and is reversion-entangled (91 T1); does not transfer to new entrants. [F6] - Playoff difficulty factor is confounded with surviving-field strength (91 T4); nsba3-only. - The whole breadth/coverage story is one season (nsba3), ~5 teams (T2/T8). The SOS cut shows nsba1/nsba2 have zero coverage variance. - The buy-low board is a watchlist, not a measurement (90 M2 — anchor is raw θ r=0.74, not de-biked; Track B is intel-only conf 0.40). [F12] - Combine→value validity is optimistic because no-shows/tankers are missing (T5). - Win% is schedule-confounded; SOS shaves 0.03–0.06 off most correlations (small) but ~25% off the CS team correlation specifically (T6).
Factual corrections already applied: top-25 nsba4-combine pool is 3, not 5 (Akhil, Rohan G, Kian); buy-low anchor is raw theta_overall (0.74), not de-biked.
5. Open questions for David (the holes a skeptic should press on)
- CS magnitude is unfalsifiable at this n. The team CS ratio's 95% CI is [−0.28, 7.14]. We recommend "secure one CS body, use theta_cs, don't reach." Does David's lived NSBA experience corroborate that CS is modest depth/speed value, not a 3.5× lever? (His read on whether strong generalists pick up early-undergrad CS — OPEN_Q #8 — is the tiebreaker.)
- The draft→game bridge. nsba1 is effectively unlinkable (Jaccard 0.11); much of the draft→outcome and availability story is nsba2+nsba3. Can David confirm the nsba1 nickname↔real-name identities, or should we permanently treat nsba1 as games-only?
- Who actually re-registered for nsba4. Only 3 of the top-25 value-board players carry an nsba4 combine row. The board is historical talent; the live pool is thin. We need the confirmed nsba4 eligibility + captain list (esp. Ziang — captain or draftable?) before the buy-low board is actionable, and David's draft slot (OPEN_Q #9) to tune VONA.
- nsba3 estimated-TUH. The noisiest-measured season carries the coverage, playoff, and CS signals and is most like nsba4. Where possible we used buzz-count rates; David's read on whether nsba3 packets/format are comparable enough to pool (OPEN_Q #6) bounds how hard to weight it.
- Captaincy and the "remove a pair" home rule (OPEN_Q #1) still shape the optimizer and the buy-low board (captains can't be bought low).
Reproduce any number above with the named scripts/*.py under the venv at
/home/david/code/nsba/.venv/bin/python. Artifacts are in data/processed/ and outputs/.