NSBA Draft Analyticsembargoed · 2026-06-06

OPEN_QUESTIONS.md

Open Questions (for David)

Open Questions for David

David has several years of science-bowl knowledge — use it. Questions are grouped by how blocking they are. I'll keep appending as analysis surfaces more. Answer inline (e.g. > A: ...).

✅ RESOLVED 2026-05-30 (David answered — see DECISIONS D14–D25)

🔴 Blocking-ish (shape the models)

  1. Home-advantage "remove a pair": the rules say the home team may "remove one pair." Does "pair" mean a toss-up/bonus pair, a category, or a pair of players? This changes the roster optimizer materially. What exactly gets removed?
  2. Tanking — who and why? Which specific players are known/suspected to sandbag the combine, and what's the motive (avoid captaincy? slide to a friend's team? collusion)? Knowing the mechanism lets us predict who and correct their combine score.
  3. Lineup/subs in scoresheets: do the game sheets reliably encode who was on the floor each toss-up (subs allowed), or only who buzzed? This gates whether we can do RAPM/plus-minus vs. box-score value only. (Parser A2 will report what's recoverable — please confirm.)
  4. Bonus mechanics: standard SB bonus = 10, one per correct toss-up, no bouncebacks? Confirm so the parser's score reconstruction is right.

🔴 New, from Discord intel (A5)

  1. "Biking" de-bias: confirm the combine-gaming mechanism (late-MC guessing +EV; short-answer no penalty). Do you want a de-biked combine score as a primary input, and is "CS is unbikable" right (i.e., should we trust CS combine numbers more than other categories)?
  2. Eligibility calls: confirm santhosh_b. is excluded (confirmed cheat), and confirm the staff-tier non-draftable list (Anurag, Jonathan, Mingle, Naveen-23, Cryo/Bennett, eagle_student, you). Also: is Ziang a captain (drafts) or a draftable player (potential buy-low)? Can't be both.

🟡 Directional (sharpen strategy)

  1. North-star metric: once A15 proposes the team quantity that best predicts standings, I'll ask you to gut-check whether it matches how winning actually feels in NSBA (e.g., is it raw scoring, category coverage, top-end stars, or depth?).
  2. Alumni-written (NSBA1/2) vs packet (NSBA3/4): do you think the format change shifted difficulty/category balance enough that we should down-weight NSBA1/2 when projecting NSBA4? Or are they comparable enough to pool fully?
  3. Category set per season: confirm CS was present as a category in NSBA2 and NSBA3 games (you said so), and that NSBA1 (2022) used Energy (not CS). Affects cross-season CS analysis.
  4. CS thesis sanity check: is early-undergrad CS actually hard for the field, or do strong generalists pick it up? Your read here tells us whether the CS-scarcity edge is real before we over-commit draft capital to it.

🟢 Logistics (when convenient)

  1. Your draft slot: you'll know your pick position ~06-06 — please drop it in then so the live draft tool can be tuned to it.
  2. Combine completeness: the NSBA4 combine is in progress. Are there known players still to take it, or expected no-shows, that we should hold board slots for?
  3. Grade/age data: any roster/signup with actual grades? Otherwise we proxy development by years-in-league + whatever A5 extracts from Discord.
  4. Identity reconciliation: A3 will produce data/processed/identity_needs_review.csv with ambiguous name matches across seasons/sheets — a few minutes from you to confirm those will materially improve every returning-player projection.

🔴 New, from the A30 synthesis + red-teams (post-revision, blocking for the live board)

  1. CS magnitude sanity check (supersedes #8, now urgent). The "CS ~3.5× a generic point" headline did not survive review — re-derived it is ~2× on toss-ups only (p=0.047) and collapses to ~1.2–1.4× once we control for team strength (95% CI [−0.28, 7.14], i.e. we can't even rule out CS being less valuable). The only robust CS legs are player-level: CS specialists convert CS at 62% vs 49% for generalists, and the combine's theta_cs honestly identifies CS ability (r=0.60). Recommendation: secure one credible CS body (pick a good one via theta_cs), don't actively punt CS, and don't reach beyond that. Does this match your lived NSBA read — is CS modest depth/speed value rather than a lever worth a big reach?
  2. nsba1 identities — permanently games-only? The draft→game roster bridge is effectively null for nsba1 (median Jaccard 0.11; only 19% of drafted players ever appear in a clean game, because nsba1 game sheets use nicknames vs real names on the draft sheet). This guts nsba1 from the draft→outcome, ADP, and availability analyses. Can you reconcile the nsba1 nickname↔real- name map, or should we lock nsba1 as games-only and stop trying to use its draft?
  3. Live nsba4 pool is thin — confirm eligibility + captains. Only 3 of the top-25 value- board players carry an nsba4 combine row (Akhil Batchu, Rohan G, Kian Dhawan). Most historical value is locked in players who may not return. We need: (a) the confirmed nsba4 draft-eligible list, (b) the captain list (captains can't be "bought low" — esp. Ziang, flagged as both), (c) your draft slot ("#9" is a PLACEHOLDER the synthesis assumed — NOT confirmed; you learn the real slot ~06-06) to compute VONA / tune reach.
  4. Buy-low intel confirmation. The marquee intel-only buys (Riyan N / nocombomomento, Ziang / deasert_willow) rest on Discord chatter ("genuinely sandbagging," "CS is unbikable") with NO game tape (conf 0.40). The game-backed buys (Mihir K conf 0.71, Rohan G 0.79) and fades (Vishnu M −1.20 conf 1.00, Kian D) are firmer. Can you confirm the sandbagging claims, and that the de-biked-vs-raw combine anchor choice (we use raw theta_overall, r=0.74) is acceptable?
  5. Don't double-count development. For the buy-low board we now use ONE reversion+growth model (a low combine from a real returner already implies most of their projected rise; we don't add a separate "+0.4 z growth" on top). And we credit growth to returners only — a brand-new nsba4 entrant gets priced off shrunk combine theta with no development bump. Confirm this is the intended treatment of young/rising players (we have almost no grade/age data — see #11).
  6. nsba3 weighting. nsba3 is the noisiest-measured season (estimated TUH) yet carries the coverage, playoff, and most CS signal, and is the most nsba4-like. We default to pooling with nsba3 down-weightable; the value-board top-15 is stable to dropping it. Your read on nsba1/2-vs- nsba3/4 difficulty/format comparability (#6) sets how hard we lean on nsba3 for the live draft.

NSBA Draft Analytics · embargoed until after the SSB draft · ← hub