Science Bowl Coverage Baseline (SSB-2026)
Science Bowl Coverage Baseline (SSB-2026)
Natural-roster (school-formed team) baseline for subject coverage, breadth, and celerity, built from real Science Bowl buzz-level data. Feeds:
- A16 — generalist vs specialist / subject-coverage optimization
- A18 — natural (school) vs drafted (NSBA) rosters
Source & method
- Origin:
scibowl-org/stats/ssb-2026-combined(the scibowl.live "SSB 2026 — All Tournaments" dataset,schema_version 1, buzzpoints generated 2026-05-15). Also viewable at https://www.scibowl.live/tournaments/ssb-2026-combined. - Combined from 3 real tournaments: Johns Hopkins Invitational, Stanford College+ Science Bowl, Stanford Science Bowl.
- Grain: raw data is one row per buzz (13,632 buzzes over 831 questions, 286 scored games). We aggregate to stable per-team and per-player season totals.
- Categories: real Science Bowl uses 6 — Biology, Chemistry, Physics, Math, Earth/Space (ESS), Energy. There is no Computer Science (NSBA-only). 5 categories overlap NSBA + Energy + celerity.
- Scoring (canonical, matches scibowl.live
classifyBuzz): tossup correct +4, tossup neg (incorrect interrupt) −4, end-of-read wrong 0, bonus correct +10. - Celerity (canonical): fraction of the stem unread at a correct buzz =
(word_count − word_index)/word_count, clamped [0,1]. Higher = earlier/faster. Only meaningful for correct tossups.
Data quality: excellent. Every team's reconstructed buzz points exactly equal its
actual recorded game points (mean |diff| = 0.0 across all 84 teams; 100% reconcile).
The one caveat we had to fix: raw player_id/team_id are per-game identifiers,
not stable entities (one person → many player_ids). We aggregate on
(name, team, tournament) and (team, tournament) to recover true season totals.
After aggregation: 84 stable teams, 366 stable players. All teams played ≥4 games
(median 5; the deep-bracket teams reach 9–12). One more nuance: bonuses are
team-answered (their buzz rows have a blank player_id), so the player table is a
pure tossup profile and bonus points live only in the team table — exactly as
Science Bowl scores them. All team-level findings below are unaffected.
A. How common is solid coverage of each subject?
Across all 84 natural teams (covered = ≥2 correct tossups in the subject; punted = 0):
| Subject | Covered (≥2 TU) | Punted (0 TU) | Mean pts share |
|---|---|---|---|
| Chemistry | 97.6% | 1.2% | 16.9% |
| Earth/Space | 96.4% | 2.4% | 16.7% |
| Physics | 94.0% | 3.6% | 18.7% |
| Math | 92.9% | 1.2% | 17.8% |
| Energy | 92.9% | 2.4% | 13.5% |
| Biology | 90.5% | 0.0% | 16.4% |
Takeaway: in natural rosters, every subject is covered by the large majority of teams — there is no "commonly punted" subject. Biology is essentially never fully punted (0% with zero correct TU). Energy carries the smallest point share (~13.5%), reflecting that it is a smaller slice of the question distribution, not that teams neglect it. Coverage of all six is the norm, not the exception — the natural baseline is broad. Where teams differ is in depth/celerity, not breadth.
B. Breadth vs concentration — which wins?
| Win-pct quartile | n_subjects_covered | breadth_HHI (lower=broader) | avg_celerity | tu_conv% |
|---|---|---|---|---|
| Q1 (worst) | 5.14 | 0.261 | 0.048 | 0.451 |
| Q2 | 5.75 | 0.225 | 0.071 | 0.550 |
| Q3 | 5.90 | 0.195 | 0.116 | 0.633 |
| Q4 (best) | 6.00 | 0.180 | 0.155 | 0.707 |
corr(win_pct, n_subjects_covered) = +0.52corr(win_pct, breadth_HHI) = −0.57(more concentrated → worse)- Even controlling for total scoring, breadth still helps:
partial
corr(win_pct, HHI | total_pts) = −0.30, partialcorr(win_pct, n_subjects_covered | total_pts) = +0.29.
Takeaway: breadth beats concentration. Balanced, generalist teams win more. A specialist stack (high HHI) is negatively associated with success even after adjusting for raw point output. In the natural game, you cannot win by punting a subject and over-loading another — the question distribution is even, so coverage gaps are directly exploited.
C. Celerity (buzz speed) vs success and coverage
corr(win_pct, avg_celerity) = +0.78— the single strongest style predictor, on par with tossup conversion (+0.79) and just under raw points (+0.85).corr(total_buzz_points, avg_celerity) = +0.78.- Celerity is the most separating metric across quartiles: Q4 teams buzz at 0.155 mean celerity vs 0.048 for Q1 — >3× earlier.
corr(avg_celerity, n_subjects_covered) = +0.46— faster teams also tend to be broader. Speed and breadth travel together; they are not a trade-off.- Even controlling for total points,
partial corr(win_pct, celerity | total_pts) = +0.35— buzzing earlier carries information about winning beyond just scoring more.
Takeaway: celerity is a first-class success signal. Elite natural teams win by buzzing earlier across a broad base, not by out-grinding on a narrow specialty.
D. What the TOP teams actually look like
Top-10 teams by win-pct vs the rest:
| Group | mean n_subjects_covered | mean breadth_HHI | mean celerity |
|---|---|---|---|
| Top 10 | 6.00 | 0.179 | 0.172 |
| Rest | 5.59 | 0.224 | 0.084 |
Every single top-10 team covers all 6 subjects, runs a near-balanced profile (HHI ≈ 0.18, vs the 0.167 floor of a perfectly even team), and buzzes ~2× faster than the field. Representative top-8 subject-share profiles (% of positive points):
Montgomery Blair A W%1.00 cel.148 bio18 che19 phy15 mat20 ess16 ene12 (textbook balance)
Mission San Jose A W%0.92 cel.233 bio18 che17 phy17 mat16 ess18 ene14 (most balanced + fastest)
MIT Team #2 W%0.91 cel.162 bio23 che25 phy10 mat23 ess10 ene 9 (the rare tilt: bio/chem/math heavy)
Davidson A W%0.90 cel.154 bio 8 che16 phy21 mat19 ess18 ene17 (light bio, deep everywhere else)
Even the "tilted" top teams still cover every subject (≥2 TU) — they just weight toward strengths. The dominant pattern is balance + speed.
Implications for NSBA (A16 / A18)
- Optimal coverage = all subjects covered, balanced. The natural baseline says broad beats narrow. A drafted roster that punts a subject to stack another is fighting the evidence: HHI concentration correlates with losing.
- Celerity is a draftable edge. Buzz speed is nearly as predictive of winning as conversion and points. NSBA should treat a player's speed profile, not just accuracy, as a core asset.
- Natural rosters set a high coverage bar (≈90–98% per subject). For A18, the right comparison is: do NSBA drafted rosters achieve the same near-universal six-subject coverage that school teams reach organically? If drafting produces more coverage gaps than natural formation, that is a draft-process inefficiency worth flagging. (Note: CS scarcity is NSBA-internal and not in this baseline.)