Open Questions (for David)
Open Questions for David
David has several years of science-bowl knowledge — use it. Questions are grouped by how
blocking they are. I'll keep appending as analysis surfaces more. Answer inline (e.g. > A: ...).
✅ RESOLVED 2026-05-30 (David answered — see DECISIONS D14–D25)
- Home "remove a pair" → removes the opponent's best subject (category); balance matters in playoffs (D14).
- Tanking → real but DON'T over-weight (D15). Lineup/subs → unreliable, don't trust (D16). Bonuses → no bouncebacks (D17).
- Captains → not a distinct tier; irrelevant to valuation (D18). CS analysis → David confirms legit; CS hard because nobody studies it (D21).
- Season weighting → weight nsba3 more for nsba4 (D19). nsba1 "Energy" had ~1 CS Q/packet (D20). No grade/age data → years-in-league only (D22).
- Draft slot → still unknown (D25). Non-finishers → fall back to historic results (D25).
- ACTION ITEMS: the draft→game link gap is a bug to fix + build a reconciliation worklist for David (D23 → finding 37); Anurag's cleanup-get refinement (D24 → finding 38).
- STILL OPEN: final eligible list + your draft slot (need ~06-06); plus the reconciliation worklist (I'll generate it for you to confirm).
🔴 Blocking-ish (shape the models)
- Home-advantage "remove a pair": the rules say the home team may "remove one pair." Does "pair" mean a toss-up/bonus pair, a category, or a pair of players? This changes the roster optimizer materially. What exactly gets removed?
- Tanking — who and why? Which specific players are known/suspected to sandbag the combine, and what's the motive (avoid captaincy? slide to a friend's team? collusion)? Knowing the mechanism lets us predict who and correct their combine score.
- Lineup/subs in scoresheets: do the game sheets reliably encode who was on the floor each toss-up (subs allowed), or only who buzzed? This gates whether we can do RAPM/plus-minus vs. box-score value only. (Parser A2 will report what's recoverable — please confirm.)
- Bonus mechanics: standard SB bonus = 10, one per correct toss-up, no bouncebacks? Confirm so the parser's score reconstruction is right.
🔴 New, from Discord intel (A5)
- "Biking" de-bias: confirm the combine-gaming mechanism (late-MC guessing +EV; short-answer no penalty). Do you want a de-biked combine score as a primary input, and is "CS is unbikable" right (i.e., should we trust CS combine numbers more than other categories)?
- Eligibility calls: confirm
santhosh_b.is excluded (confirmed cheat), and confirm the staff-tier non-draftable list (Anurag, Jonathan, Mingle, Naveen-23, Cryo/Bennett, eagle_student, you). Also: is Ziang a captain (drafts) or a draftable player (potential buy-low)? Can't be both.
🟡 Directional (sharpen strategy)
- North-star metric: once A15 proposes the team quantity that best predicts standings, I'll ask you to gut-check whether it matches how winning actually feels in NSBA (e.g., is it raw scoring, category coverage, top-end stars, or depth?).
- Alumni-written (NSBA1/2) vs packet (NSBA3/4): do you think the format change shifted difficulty/category balance enough that we should down-weight NSBA1/2 when projecting NSBA4? Or are they comparable enough to pool fully?
- Category set per season: confirm CS was present as a category in NSBA2 and NSBA3 games (you said so), and that NSBA1 (2022) used Energy (not CS). Affects cross-season CS analysis.
- CS thesis sanity check: is early-undergrad CS actually hard for the field, or do strong generalists pick it up? Your read here tells us whether the CS-scarcity edge is real before we over-commit draft capital to it.
🟢 Logistics (when convenient)
- Your draft slot: you'll know your pick position ~06-06 — please drop it in then so the live draft tool can be tuned to it.
- Combine completeness: the NSBA4 combine is in progress. Are there known players still to take it, or expected no-shows, that we should hold board slots for?
- Grade/age data: any roster/signup with actual grades? Otherwise we proxy development by years-in-league + whatever A5 extracts from Discord.
- Identity reconciliation: A3 will produce
data/processed/identity_needs_review.csvwith ambiguous name matches across seasons/sheets — a few minutes from you to confirm those will materially improve every returning-player projection.
🔴 New, from the A30 synthesis + red-teams (post-revision, blocking for the live board)
- CS magnitude sanity check (supersedes #8, now urgent). The "CS ~3.5× a generic point"
headline did not survive review — re-derived it is ~2× on toss-ups only (p=0.047) and
collapses to ~1.2–1.4× once we control for team strength (95% CI [−0.28, 7.14], i.e. we
can't even rule out CS being less valuable). The only robust CS legs are player-level: CS
specialists convert CS at 62% vs 49% for generalists, and the combine's
theta_cshonestly identifies CS ability (r=0.60). Recommendation: secure one credible CS body (pick a good one via theta_cs), don't actively punt CS, and don't reach beyond that. Does this match your lived NSBA read — is CS modest depth/speed value rather than a lever worth a big reach? - nsba1 identities — permanently games-only? The draft→game roster bridge is effectively null for nsba1 (median Jaccard 0.11; only 19% of drafted players ever appear in a clean game, because nsba1 game sheets use nicknames vs real names on the draft sheet). This guts nsba1 from the draft→outcome, ADP, and availability analyses. Can you reconcile the nsba1 nickname↔real- name map, or should we lock nsba1 as games-only and stop trying to use its draft?
- Live nsba4 pool is thin — confirm eligibility + captains. Only 3 of the top-25 value- board players carry an nsba4 combine row (Akhil Batchu, Rohan G, Kian Dhawan). Most historical value is locked in players who may not return. We need: (a) the confirmed nsba4 draft-eligible list, (b) the captain list (captains can't be "bought low" — esp. Ziang, flagged as both), (c) your draft slot ("#9" is a PLACEHOLDER the synthesis assumed — NOT confirmed; you learn the real slot ~06-06) to compute VONA / tune reach.
- Buy-low intel confirmation. The marquee intel-only buys (Riyan N / nocombomomento, Ziang / deasert_willow) rest on Discord chatter ("genuinely sandbagging," "CS is unbikable") with NO game tape (conf 0.40). The game-backed buys (Mihir K conf 0.71, Rohan G 0.79) and fades (Vishnu M −1.20 conf 1.00, Kian D) are firmer. Can you confirm the sandbagging claims, and that the de-biked-vs-raw combine anchor choice (we use raw theta_overall, r=0.74) is acceptable?
- Don't double-count development. For the buy-low board we now use ONE reversion+growth model (a low combine from a real returner already implies most of their projected rise; we don't add a separate "+0.4 z growth" on top). And we credit growth to returners only — a brand-new nsba4 entrant gets priced off shrunk combine theta with no development bump. Confirm this is the intended treatment of young/rising players (we have almost no grade/age data — see #11).
- nsba3 weighting. nsba3 is the noisiest-measured season (estimated TUH) yet carries the coverage, playoff, and most CS signal, and is the most nsba4-like. We default to pooling with nsba3 down-weightable; the value-board top-15 is stable to dropping it. Your read on nsba1/2-vs- nsba3/4 difficulty/format comparability (#6) sets how hard we lean on nsba3 for the live draft.