Decisions Log
Decisions Log
Settled decisions and their rationale. Append-only; date everything. (Today = 2026-05-30.)
| # | Decision | Rationale | Date |
|---|---|---|---|
| D1 | 14 teams in NSBA4 (snake draft, 5 picks default, 4–6 tradeable; 4–7 registered players; 4 play at once; all 6 categories cycle per match). | Confirmed by David. Sets replacement level / board depth. | 05-30 |
| D2 | North-star value metric is a research OUTPUT, not an assumption. Working candidate: per-player win shares derived from "points per tossup faced" (PPTF) that best predicts standings (wins, then total points per the tiebreak rule). | David: "this should be an outcome of your research." | 05-30 |
| D3 | Availability demoted to a soft signal. No reliable availability data exists; we mine Discord (A5) for flakiness/engagement and flag it, but do NOT model it as if precise. | David: "there is no good way to get availability." | 05-30 |
| D13 | PARTIAL availability recovery via draft data. With actual draft results (who was drafted) + game appearances (who played), we can measure realized availability = games-played / games-possible for each historically drafted player. This is a real, backward-looking reliability signal (a drafted player who barely played = availability bust) — upgrades D3 from "soft only" to "measurable for returning players." | NSBA1/2/3 draft files provided. | 05-30 |
| D4 | Team complementarity is central. Core research question = generalist vs. specialist; category coverage modeled as a covering + portfolio problem. | David's explicit priority. | 05-30 |
| D5 | Two separate scoring systems. Combine = pyramidal/gameable tryout (+4/+3/+2, −1 MC). Games = standard Science Bowl (4 / −4 / 10, not pyramidal). Celerity/early-buzz value is combine-only. | David correction — earlier I wrongly assumed games were pyramidal. | 05-30 |
| D6 | No heavy ML (n is dozens). Core = hierarchical-Bayes IRT (GPCM/MIRT) on combine+game response matrices; Marcel-style projection as the baseline the engine must beat; empirical-Bayes shrinkage for rate stats. | A7 + A6 research; small-sample overfitting risk. | 05-30 |
| D7 | Trust game data over the combine. Combine is a censored signal: high scores informative, low scores ambiguous (possible tank). Lean on combine mainly for rookies with no game history. | A6/A7; tanking bias. | 05-30 |
| D8 | CS is a scarce-specialist category. Value via VONA (value over next available) + tier-counting, not raw VORP; reach before the cliff — but FIRST falsify the thesis (do generalists already convert early-undergrad CS?). | A9; CS replaced Energy and is rare in the pool. | 05-30 |
| D9 | Embargo. No public release until after the SSB draft, to keep it interesting. Findings are for science-bowl roster construction in general, not only NSBA. | David. | 05-30 |
| D10 | Orchestrator model. Claude-as-orchestrator dispatches subagents/workflows for the heavy lifting and keeps context for the big picture; subagents parse, model, audit, synthesize. | David: "you are the orchestrator." | 05-30 |
| D12 | The combine is actively GAMED ("biking"). Under "−1 only for wrong multiple-choice," waiting until the last seconds to guess MC is +EV, and short-answer buzzes carry no penalty — so raw combine score rewards aggressive guessing strategy, not just knowledge. Honest skippers are under-scored. → Build a de-biked combine signal (down-weight late-MC-guess points, value early/short-answer/"unbikable" buzzes), treat raw combine as a gamed ceiling, and lean on game data (D7). Confirmed cheater santhosh_b. → exclude/discount. Maintain a staff-tier non-draftable list (organizers/readers incl. David xpoes). |
A5 Discord intel. | 05-30 |
| D11 | Subject coverage is a headline deliverable: quantify (a) the optimal coverage profile and (b) how common solid coverage of each subject is in general. Use scibowl.live SSB-2026 (Stanford) stats for celerity (speed) + natural-roster coverage baselines (real SB uses Energy, not CS, so CS scarcity stays internal). | David. | 05-30 |
David's answers → decisions (2026-05-30, batch 2)
- D14 — Home advantage: persists; equal home games per team; the home team removes the opposing team's best subject (a category). Implication: a one-subject team is vulnerable in playoffs (its best subject gets stripped), so balance has genuine playoff/seeding value via this rule — a narrow, real rehabilitation of breadth that the regular-season data couldn't see.
- D15 — Tanking: DEMOTE. It happens (a NC group coordinated-tanked to fly under the radar) but David says don't weight combine-throwing hard. Keep buy-low as a game-backed signal, not heavy combine-tank modeling.
- D16 — Lineup/subs data is unreliable ("not measured properly") → do NOT trust lineup-based metrics. Vindicates relying on box-score PPTF over RAPM/contestation per-name flags.
- D17 — Bonuses: no bouncebacks (confirms the parse; bonus = 10, one per correct toss-up).
- D18 — No captain tier: captains are not different from regular players → captaincy is irrelevant to player valuation; drop the "captain can't be bought low" concern (Ziang etc.).
- D19 — Season weighting: weight nsba3 more when projecting nsba4 (both are packet-submission, similar difficulty). Overrides the earlier "down-weight nsba3 for noise" lean for projection purposes.
- D20 — nsba1 categories: "Energy" actually included ~1 CS question per packet — not pure Energy. Tiny nsba1 CS signal exists; immaterial to conclusions.
- D21 — Why CS is hard: the field doesn't study CS (no established canon), unlike the other 5 subjects → CS scarcity thesis validated (David: "cs analysis seems legit"); but CS may be noisier / less predictable from combine (less canon to anchor on).
- D22 — No grade/age data at all → growth proxied only by years-in-league.
- D23 — Draft→game identity link gap is a BUG (David: "shouldn't be that big of a gap"), not reality. Fix the reconciliation and produce a worklist for David to confirm names (answers his "what's the best way to figure out which names to reconcile").
- D24 — Anurag's cleanup-get idea: a get that comes after the opposing team negged = cleaning up a question, not out-buzzing them → should be credited differently from a clean first-buzz win. To investigate.
- D25 — Draft slot still unknown (~06-06); combine still in progress, some may not finish → fall back to historic results for non-finishers (answers Q10/Q17).
Season ↔ year ↔ format mapping
- nsba1_2022, nsba2_2023, nsba3_2025, nsba4_2026 (current; combine in progress, finalizes ~06-06).
- Game scoresheets exist for nsba1/2/3 only (we're drafting nsba4).
- Question source confound: NSBA1/2 alumni-written; NSBA3/4 team packet-submissions (alumni edit). NSBA3 is the most format-comparable to NSBA4.