18 — Natural vs Drafted Rosters: Does Drafting Create Subject-Coverage Gaps?
18 — Natural vs Drafted Rosters: Does Drafting Create Subject-Coverage Gaps?
Question. Real Science Bowl school teams form organically and reach near-universal 6-subject coverage (90–98% per subject; ~81% cover all six at a depth threshold). NSBA teams are drafted via a snake draft seeded by a gameable combine. Does the draft reproduce that broad coverage, or does it create gaps — especially in Computer Science, the NSBA-only subject flagged as scarce/unbikable? And what does a gap cost in win%?
TL;DR. Drafting does not create a coverage deficit. NSBA drafted teams cover all six subjects at least as broadly as natural school teams — 95.1% cover all 6 (depth ≥2 correct TU) vs 81.0% of natural teams, and 100% vs 100% once you restrict to teams with a comparable schedule (≥8 games). CS is covered by 96.6% of nsba2/nsba3 teams — it is not a draft hole at the team level. The within-NSBA breadth→winning signal is the same as the natural game (corr(win%, HHI) = −0.47, corr(win%,
subjects) = +0.53, vs scibowl −0.57 / +0.52). Roster construction, on the
coverage dimension, is effectively solved by the draft — gaps appear only on tiny-sample teams that barely played.
Method
- Data.
team_season.csv(41 folded NSBA team-seasons: nsba1=12, nsba2=13, nsba3=16),tossups_long.csv(clean-game buzzes),scibowl_team_stats.csv(84 natural teams from JHU + 2 Stanford tournaments). Env:.venv/bin/python. - Clean game set (guardrail).
games_metafiltered toreconciles==TrueAND not(score_a==0 & score_b==0) → 180 clean games (nsba1=66, nsba2=61, nsba3=53). Buzzes inner-joined to that set. - Team-name folding. Re-applied the
build_masteralphanumeric-key fold so split spellings (Mount Ellis/mount ellis,Chonk/chonk,no like bluevariants,dha wei/dhawei,Zootopia/zootopia,Fish noobs/fishnoobs) collapse to the 41 canonical team-seasons inteam_season.csv. 0 correct-buzz rows unmapped. - Apples-to-apples coverage metric. A subject is covered if the team has ≥2
correct toss-ups in it — the exact threshold the scibowl baseline uses
(
n_subjects_covered). Each season has a clean 6-subject universe: nsba1 = {bio, chem, ess, math, phys, Energy}; nsba2/nsba3 = {bio, chem, ess, math, phys, CS}. Energy↔CS swap means every NSBA season still spans exactly 6, matching scibowl's 6 (which include Energy, not CS). CS is NSBA-internal and absent from the natural baseline, so CS is reported separately. - Caveat baked into the metric. This is in-game realized coverage (did the team actually convert ≥2 TU in the subject), not roster-design coverage. It is partly mechanical: more games → more chances to clear the threshold. We control for this by also reporting the ≥8-game subsample.
- Artifact:
data/processed/team_subject_coverage_nsba.csv(per-team correct-TU counts + coverage flags + win%/HHI).
Results
1. Drafted rosters match — even exceed — natural coverage
| Population | n teams | cover all 6 (≥2 TU) | cover ≥5 | mean #subjects |
|---|---|---|---|---|
| NSBA drafted (all) | 41 | 95.1% | 97.6% | 5.85 |
| Natural (scibowl, all) | 84 | 81.0% | 89.3% | 5.64 |
| NSBA, ≥8 games | 29 | 100.0% | 100% | 6.00 |
| Natural, ≥8 games | 31 | 100.0% | 100% | 6.00 |
NSBA looks broader than natural in the raw all-teams comparison, but that is the games-played confound: NSBA teams play more games (median 10 vs 5), so they clear the ≥2 threshold more easily. Hold schedule fixed (≥8 games) and both populations hit 100% all-six coverage — they are indistinguishable. The honest read is parity, not NSBA superiority.
Per-subject covered rates are essentially identical to the natural baseline:
| Subject | NSBA covered (≥2 TU) | Natural (scibowl) |
|---|---|---|
| Biology | 97.6% | 90.5% |
| Chemistry | 95.1% | 97.6% |
| Earth/Space | 97.6% | 96.4% |
| Math | 100.0% | 92.9% |
| Physics | 97.6% | 94.0% |
| Energy (nsba1 only, n=12) | 100.0% | 92.9% |
| CS (nsba2/3 only, n=29) | 96.6% | — (NSBA-only) |
2. CS is not a draft hole
The strongest prior worry — that drafting under-supplies the scarce, "unbikable" CS
subject — does not show up at the team level. Across the 29 nsba2/nsba3 teams: median
9 correct CS toss-ups, only 1 team (3.4%) has 0, and that team
(2yellow4brown) played a single game. CS coverage (96.6%) sits right in the pack
with the other five subjects. The draft reliably lands at least one CS-capable answerer
on essentially every team.
3. The only NSBA teams missing a subject are tiny-sample teams
| team | season | games | win% | #subjects (≥2) |
|---|---|---|---|---|
| 2yellow4brown | nsba3 | 1 | .000 | 1 |
| okc | nsba3 | 5 | .200 | 5 |
Restricting to teams with ≥4 games, 97.5% cover all six. There is no real drafted team with a genuine, schedule-driven coverage gap.
4. Breadth still beats concentration inside NSBA (same as the natural game)
| metric | NSBA (n=39 w/ win%) | Natural (scibowl, n=84) |
|---|---|---|
| corr(win%, subject HHI) — lower HHI = broader | −0.47 | −0.57 |
| corr(win%, #subjects scored) | +0.53 | +0.52 |
| mean / median breadth HHI | 0.216 / 0.184 | 0.219 / 0.195 |
The concentration→losing relationship and the breadth→winning relationship reproduce in NSBA at nearly the same magnitude, and the HHI distributions overlap almost perfectly (NSBA's lone HHI=1.0 is the 1-game team). Drafted teams are not building lopsided specialist stacks; they look like balanced natural teams.
5. What a coverage gap would cost (priced from the natural data)
Coverage gaps barely exist in NSBA, so the price of one is estimated where gaps are common — the natural baseline:
- scibowl OLS:
win% = −0.49 + 0.164 × #subjects_covered→ each covered subject ≈ +16 win-percentage points. - Cover-all-6 teams win 51.0%; teams missing ≥1 subject win 12.7% → a 38-point win-rate deficit.
So a coverage gap is expensive (~16 pts/subject), which is exactly why it matters that the NSBA draft almost never produces one. The cost is real; drafted teams simply don't pay it.
Interpretation
Is roster construction (coverage) solved by the draft? Yes, on this axis. A 6-round snake draft over a player pool that has at least one capable answerer per subject is enough to guarantee broad coverage without anyone optimizing for it — the same near-universal six-subject coverage that schools reach organically. Drafting does not differ from organic formation in coverage breadth, and it does not specifically starve CS.
The remaining edge therefore is not "cover all six" (that's table stakes both in the draft and in nature) but depth and speed within the covered six — consistent with the natural baseline, where the separating variables among full-coverage teams are celerity (corr +0.78) and conversion, not breadth. For the nsba4 draft target, the actionable implication: a single-subject "punt-and-stack" strategy is not available as an edge (the pool fills coverage automatically and the data punishes concentration); the exploitable margin is drafting faster, deeper answerers across an already-broad base.
Limitations
- Realized, not designed, coverage. Metric = ≥2 correct TU in games, which is partly mechanical in games-played. We mitigate with the ≥8-game subsample (100%/100%) and the ≥4-game cut (97.5%), but a pure roster-design coverage measure (combine ability per drafted player) would be a cleaner test of the draft mechanism itself and is left to a combine-side analysis.
- Small n. 41 NSBA team-seasons (12/13/16). The within-NSBA win% correlations rest on 39 teams with a win%; treat ±0.1 around the point estimates as plausible.
- Energy↔CS asymmetry. The "6th subject" differs (Energy in nsba1's natural-baseline question set; CS in nsba2/3). The 5 shared subjects are directly comparable; the 6th is matched by position, not identity. CS is reported separately precisely because it has no natural analog.
- Win% as result proxy. No official NSBA final standings; win% over the clean-game set is the placement proxy, and several teams have short/uneven schedules.
- nsba3 lineup blindness. The paired layout records buzzers but not on-floor lineups; coverage here is buzz-attributed, which is the right grain for who scored in a subject but cannot speak to bench depth that never buzzed.
Artifacts. data/processed/team_subject_coverage_nsba.csv.