Cool Findings — the highlights reel
Cool Findings — the highlights reel
A curated tour of the most surprising results, for people who actually know quiz bowl / science bowl. All of it is real code on 3 seasons of game data + 4 combines + 216 draft picks. Small samples (41 team-seasons), so every number is ±a lot — trust the directions, not the decimals. Full detail behind each link.
🏆 The headline: one humble stat beats everything we threw at it
We tried five clever ways to build a "better" player rating — and all five collapsed back to a dead-simple production rate (PPTF = net toss-up points per toss-up faced):
- buzzpoints/celerity, buzz-rank, difficulty-weighting, teammate/usage adjustment, and "cleanup gets" — each should help, each is a statistical wash out-of-sample.
- The reason: "who converts toss-ups against real opponents, net of negs" already contains nearly all the signal. Everything fancy reduces to it. → Key Findings
🔬 The buzzpoints debate — settled on real pyramidal data (for Radius)
- We tested whether buzzpoints make a markedly better ranking, on real ACF pyramidal buzz curves (91% of buzzes), not a proxy.
- Plot twist: dan's own Plackett-Luce model uses buzz position not at all for ranking — the load-bearing ingredient is opponent-adjustment, a model, not buzzpoints. Adding buzzpoints moved the ranking by ρ=0.978 (≈ nothing).
- thedoge's "buzz-rank > celerity" bet loses — and badly: rank throws away conversion volume, so it buries elite high-volume players (one player: 112 gets, rank 9 → 319). Volume isn't inflation — it's the signal. → 35 Plackett-Luce · 36 buzz-rank vs difficulty
💸 The draft market is genuinely beatable
- The field drafts the raw combine almost mechanically (ρ≈0.7), but the combine only predicts real value at ρ≈0.4 — leaving ~13 toss-points per slot of edge for anyone who drafts on projected value instead.
- And the combine is gamed: under "−1 only for a wrong multiple-choice," late-MC guessing ("biking") is +EV — so combine rank partly measures guessing strategy, not knowledge. → 21 ADP / market
- (Sorry Anurag — drafting on vibe is precisely the inefficiency a value model eats. 🙂)
🧩 Roster-construction myths, busted
- Coverage is table-stakes: drafted teams already cover all 6 subjects ~96% of the time — "don't punt a subject" doesn't survive a proper roster-design test.
- Team-building is additive: no duo/chemistry synergy survives (0 of 8 tests) — so you draft the best player, not the best fit. There is no real "if you have A and B, take C."
- CS is worth somewhat more — ~2×, NOT the 3.5× a first pass claimed (our own red-team caught it sneaking team bonus points into a per-player stat). CS is scarce mostly because nobody studies it — no canon.
- Bonuses are 55–61% of all points — yet carry no per-player signal, because you must win the toss-up to even reach the bonus. → 16b coverage · 29 duo synergy · 17b CS · 27 bonus
🧪 We put Gideon's roster theory to the test
Gideon's grade-by-role meta — "ESS main = soph, captain = senior phys/math, chem main = junior quick-with-numbers, bio with some astro" — we tested it against the data. Verdict: largely unsupported at this sample size — the specific grade-by-role pairings don't beat just taking the best available players (which tracks with the "additive, no synergy" result above). Still a great falsifiable hypothesis — more than most "metas" offer. → 14b growth/aging · 28 archetypes
📉 The pick-value curve + an on-brand self-own
- Pick value is steeply convex (Jimmy-Johnson shaped): a Round-1 pick is worth far more than two 2nds — never trade your R1 down for volume.
- Exhibit A, from our own GM: in NSBA2, David traded his R1 (slot 6) for two 2nds to load up on specialists — gave up Daniel Sun (a star), netted −2.73 win-shares, finished 9th… while the team that took Daniel Sun with that pick finished 2nd. The convex curve, in one trade. → 22 pick value · 25 trade strategy
🕵️ Steals the field slept on
- thedoge — 2nd-round pick (16th overall) → rank 3 of 66 in NSBA2.
- Vishwa "Vish" Akkati — the single biggest NSBA1 steal (pick 33 → #3 overall in win-shares).
- The steals consistently lived in rounds 2–5, not round 1. → 26 NSBA2 retro · 34 NSBA1
🙏 The honest part (why you can trust the rest)
Small n — 41 team-seasons — so everything is ±a lot. We ran adversarial red-teams against our own findings and killed several (the original CS "3.5×", a couple of buy-low flags, a circular undervaluation metric). The buy-low/fade boards are watchlists, not gospel. Trust directions and rankings, not point estimates.
Full project: the hub · embargoed until after the SSB draft.