NSBA Draft Analyticsembargoed · 2026-06-06

DECISIONS.md

Decisions Log

Decisions Log

Settled decisions and their rationale. Append-only; date everything. (Today = 2026-05-30.)

# Decision Rationale Date
D1 14 teams in NSBA4 (snake draft, 5 picks default, 4–6 tradeable; 4–7 registered players; 4 play at once; all 6 categories cycle per match). Confirmed by David. Sets replacement level / board depth. 05-30
D2 North-star value metric is a research OUTPUT, not an assumption. Working candidate: per-player win shares derived from "points per tossup faced" (PPTF) that best predicts standings (wins, then total points per the tiebreak rule). David: "this should be an outcome of your research." 05-30
D3 Availability demoted to a soft signal. No reliable availability data exists; we mine Discord (A5) for flakiness/engagement and flag it, but do NOT model it as if precise. David: "there is no good way to get availability." 05-30
D13 PARTIAL availability recovery via draft data. With actual draft results (who was drafted) + game appearances (who played), we can measure realized availability = games-played / games-possible for each historically drafted player. This is a real, backward-looking reliability signal (a drafted player who barely played = availability bust) — upgrades D3 from "soft only" to "measurable for returning players." NSBA1/2/3 draft files provided. 05-30
D4 Team complementarity is central. Core research question = generalist vs. specialist; category coverage modeled as a covering + portfolio problem. David's explicit priority. 05-30
D5 Two separate scoring systems. Combine = pyramidal/gameable tryout (+4/+3/+2, −1 MC). Games = standard Science Bowl (4 / −4 / 10, not pyramidal). Celerity/early-buzz value is combine-only. David correction — earlier I wrongly assumed games were pyramidal. 05-30
D6 No heavy ML (n is dozens). Core = hierarchical-Bayes IRT (GPCM/MIRT) on combine+game response matrices; Marcel-style projection as the baseline the engine must beat; empirical-Bayes shrinkage for rate stats. A7 + A6 research; small-sample overfitting risk. 05-30
D7 Trust game data over the combine. Combine is a censored signal: high scores informative, low scores ambiguous (possible tank). Lean on combine mainly for rookies with no game history. A6/A7; tanking bias. 05-30
D8 CS is a scarce-specialist category. Value via VONA (value over next available) + tier-counting, not raw VORP; reach before the cliff — but FIRST falsify the thesis (do generalists already convert early-undergrad CS?). A9; CS replaced Energy and is rare in the pool. 05-30
D9 Embargo. No public release until after the SSB draft, to keep it interesting. Findings are for science-bowl roster construction in general, not only NSBA. David. 05-30
D10 Orchestrator model. Claude-as-orchestrator dispatches subagents/workflows for the heavy lifting and keeps context for the big picture; subagents parse, model, audit, synthesize. David: "you are the orchestrator." 05-30
D12 The combine is actively GAMED ("biking"). Under "−1 only for wrong multiple-choice," waiting until the last seconds to guess MC is +EV, and short-answer buzzes carry no penalty — so raw combine score rewards aggressive guessing strategy, not just knowledge. Honest skippers are under-scored. → Build a de-biked combine signal (down-weight late-MC-guess points, value early/short-answer/"unbikable" buzzes), treat raw combine as a gamed ceiling, and lean on game data (D7). Confirmed cheater santhosh_b. → exclude/discount. Maintain a staff-tier non-draftable list (organizers/readers incl. David xpoes). A5 Discord intel. 05-30
D11 Subject coverage is a headline deliverable: quantify (a) the optimal coverage profile and (b) how common solid coverage of each subject is in general. Use scibowl.live SSB-2026 (Stanford) stats for celerity (speed) + natural-roster coverage baselines (real SB uses Energy, not CS, so CS scarcity stays internal). David. 05-30

David's answers → decisions (2026-05-30, batch 2)

Season ↔ year ↔ format mapping


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub