31 — Field Depth Under Difficulty (Task D)
31 — Field Depth Under Difficulty (Task D)
Date: 2026-05-30. Env: /home/david/code/nsba/.venv/bin/python.
Inputs (GAME data only — standard SB scoring, never combine as outcome):
tossups_long.csv, games_meta.csv, canonical_players.csv,
player_archetypes.csv, outputs/combine_to_game_linked.csv.
Artifact: data/processed/playoff_robustness_31.csv (per-player reg vs playoff
PPTF + archetype, nsba3).
LEAD WITH THE CAVEATS (read these first)
-
Clean playoff games exist in nsba3 only. All 9 nsba1 playoff games are non-reconciling (corrupted sheets, dropped by the guardrail); nsba2 had no playoffs. Everything below is one season: 12 clean nsba3 playoff games, 265 playoff questions, 36–43 players. This is the same nsba3-only limitation that bounds F7/finding-20. Read magnitudes as ±a lot; trust the ranking of which subjects/archetypes fade.
-
"Difficulty" is confounded with surviving-field strength. Playoff conversion is lower partly because packets are harder and partly because only the strong teams are left, so a weak buzzer faces a stronger opponent racing them to the buzzer. The two cannot be separated in n=1 season. The within-player retention (each player as their own control) is the cleanest cut but still cannot tell "the question got harder" from "my opponent got faster."
-
Depth counts (#players above a bar) are NOT comparable across phases because the playoff pool is pre-filtered (85 reg players → 61 playoff players; the weak tail is already eliminated). The fraction above a bar mechanically rises in playoffs purely from this selection — it is not evidence the field got deeper. Treat (2) as a measurement caution, not a finding.
-
nsba3 PPTF uses estimated/derived team-tossups-faced denominators (paired layout). PPTF here = player toss-up points ÷ team tossups faced in that phase.
(1) Per-subject conversion drops sharply — but only for 4 of 6 subjects
Question-level conversion (any team converts), nsba3 regular vs playoff, FDR-BH across the 6 subjects:
| subject | reg conv | playoff conv | rel. drop | Fisher OR | raw p | FDR p |
|---|---|---|---|---|---|---|
| ess (Earth/Space) | 0.888 | 0.689 | −22% | 3.56 | 0.002 | 0.012 ✓ |
| m (Math) | 0.830 | 0.674 | −19% | 2.36 | 0.036 | 0.108 |
| cs (Comp Sci) | 0.725 | 0.571 | −21% | 1.98 | 0.098 | 0.147 |
| p (Physics) | 0.859 | 0.745 | −13% | 2.09 | 0.076 | 0.147 |
| ch (Chem) | 0.795 | 0.830 | +4% | 0.80 | 0.681 | 0.681 |
| b (Bio) | 0.905 | 0.978 | +8% | 0.22 | 0.205 | 0.246 |
Only ESS survives FDR strictly (n=45 playoff questions). Math, CS, Physics are suggestive (OR ~2, raw p<0.10) but don't clear multiple-testing — the per-subject playoff n is 35–47 questions, genuinely underpowered. Bio and Chem do not drop at all (point estimates even rise; not significant either direction).
The signal: difficulty bites the quantitative/Earth-Space cluster (ESS, Math, CS, Physics) and spares the knowledge-recall cluster (Bio, Chem). This is consistent with finding-20's prior that CS/Math are the hardest subjects — under playoff pressure the hard subjects get harder while easy-recall subjects hold. (The aggregate playoff conversion drop 0.838→0.755, OR=0.59, reproduces F7.)
(2) "Depth" counts are not interpretable (selection) — reported for honesty
Players above a PPTF bar, by phase:
| bar | regular (n=85) | playoff (n=61) |
|---|---|---|
| ≥0.3 | 35 (41%) | 32 (52%) |
| ≥0.5 | 20 (24%) | 21 (34%) |
| ≥0.7 | 12 (14%) | 11 (18%) |
The fraction above each bar rises in playoffs. This is entirely a selection artifact — the playoff field is the surviving strong teams, the weak tail is gone. It does NOT mean the field is deeper when it's hard; if anything the opposite (see (3): individual production falls). Do not cite these as a depth result. The honest statement: the talent that reaches playoffs is pre-filtered, so "how deep is the bench under difficulty" cannot be answered from this data without a within-player design — which is (3).
(3) Difficulty-robust vs front-runner: by ARCHETYPE (the real finding)
Within-player retention = pooled playoff PPTF ÷ pooled regular PPTF, for the same players (each is their own control), aggregated by archetype with player-bootstrap 95% CIs. Overall retention is ~flat (0.98, CI [0.69, 1.27]) — the average surviving player roughly holds, but that average hides a strong split:
| archetype | n | retention | 95% CI | reads as |
|---|---|---|---|---|
| elite multi-science (chem+phys core) | 5 | 1.47 | [1.16, 1.74] | rises ✓ |
| math+phys specialist | 12 | 1.07 | [0.75, 1.48] | holds |
| bio specialist | 7 | 0.96 | [0.24, 2.33] | holds (noisy) |
| low-output / replacement | 11 | 0.50 | [−0.16, 1.53] | fades (noisy) |
| ess specialist | 8 | 0.44 | [0.22, 0.63] | collapses ✓ |
Two CIs exclude 1.0: - ESS specialists collapse to ~44% of their regular output in playoffs. This is mechanically downstream of (1): ESS conversion craters 22%. ESS specialists have nowhere to hide — their one subject is the one that gets hardest. - Chem+phys-core elites rise to ~147%. As a subject they sit in the cluster (chem) that holds, plus they have the breadth to pick up the points the faders leave on the table. These are the difficulty-robust drafts.
Front-runners who faded (top regular producers, low retention): Akhil Batchu (ESS, 1.00→0.53), Vishnu M (ESS, 0.79→0.53), Peter B (ESS, 0.55→0.00), Theenash Sengupta (bio, 0.73→−0.08 — the one bio exception, drags the bio CI), Harry Gao (math+phys, 0.51→0.00), Eli Mrug (math+phys, 0.69→0.29).
Difficulty-robust (held or rose from a high base): Rohan G (chem+phys elite, 1.03→1.67), Anurag S (0.85→0.97), Kian Dhawan (0.70→0.87), Anish A (0.58→0.69), Edwin He (bio, 0.66→0.73), Lockheed Martin (math+phys, 0.80→1.32). The buy-low board's top game-backed buys (Rohan G, Mihir class) and its top fade (Vishnu M) line up with robustness — Rohan G is robust, Vishnu M is a fader, independently corroborating F12.
Caveat: bio (7) and low-output (11) CIs are wide and span 1.0 — bio "holds" rests heavily on Edwin/Arjun against Theenash's collapse; don't over-read the bio row. The two load-bearing rows (ESS down, chem+phys-elite up) are the only ones whose CIs exclude 1.
(4) Combine → game: does combine predict PLAYOFF performance better? NO.
The hypothesis (deep/de-biked theta should pick out playoff-robust players) is not supported. On the 36 nsba3 players with both a combine and ≥1 playoff game, combine predicts regular-season PPTF better than playoff PPTF, for every theta variant:
| combine signal | → reg PPTF (spearman) | → playoff PPTF (spearman) |
|---|---|---|
| raw_overall | 0.71 | 0.58 |
| theta_overall | 0.71 | 0.55 |
| debiked_overall | 0.60 | 0.42 |
The reg-minus-playoff correlation gap is +0.12 (raw) / +0.15 (theta); bootstrap 85–90% of resamples positive but CI includes 0 (n=36) — so the decay is directional, not significant. And combine does not predict retention at all (theta_overall → retention spearman 0.17, p=0.32; debiked 0.10, p=0.56).
De-biking does not help here — debiked theta is the weakest playoff predictor (0.42), consistent with F3 (de-biking only earns its keep on the biked nsba1/Energy era, and slightly hurts on clean seasons; nsba3 is clean).
Interpretation: the combine is a regular-season talent signal that decays, not sharpens, under playoff difficulty. It is not a hidden playoff-robustness detector. Whatever makes a player hold up when it's hard (breadth into the holding subjects, composure vs a fast surviving field) is not captured by combine theta.
Implications for drafting playoff-robust players
-
The combine cannot find your playoff performers beyond finding good regular-season players generally — and it does that worse for playoffs. Do not pay a premium for a high combine expecting playoff upside; price combine as a regular-season signal (F3) and shrink it further for the games that decide titles.
-
Archetype is a better playoff-robustness handle than combine score. The one repeatable structural signal: breadth (chem+phys-core multi-science) holds or rises; single-subject ESS specialists collapse because their subject is the one difficulty hits hardest. This is a refinement of F1/F4: coverage is table-stakes for the regular season, but depth-into-the-robust-subjects (Bio/Chem recall + breadth) is what survives playoff difficulty. A roster that is "covered" on paper but leans on an ESS or pure-quant specialist for its points is a regular-season team that fades in March.
-
Tiebreaker, weakly held. With n=1 season and CIs that mostly span 1.0 outside the two flagged rows, this is a tiebreaker between similar players, not a re-ordering lever: prefer the broader chem/phys/bio-flavored producer over the equivalent-PPTF ESS/quant specialist for playoff equity, all else equal. Do not reach past the value board (PPTF/VORP, F2/F8) on this alone.
Limitations summary
- n=1 season (nsba3), 12 playoff games, 265 questions, 36–43 players. Not reproducible across seasons; no out-of-sample check possible.
- Difficulty vs surviving-field-strength is unidentified (caveat 2).
- Per-subject playoff n=35–47 questions; only ESS clears FDR.
- Archetype retention CIs are wide; only ESS-down and chem+phys-elite-up exclude 1.
- Combine→playoff decay is directional (85–90% bootstrap) but not significant at n=36; retention is unpredicted by combine (this null is reported honestly).
- 281/1781 nsba3 buzz rows unmapped to canonical_id (name-join misses); the per-player table covers the mapped majority and pools by team-tossups-faced.