NSBA Draft Analyticsembargoed · 2026-06-06

31_field_depth.md

31 — Field Depth Under Difficulty (Task D)

31 — Field Depth Under Difficulty (Task D)

Date: 2026-05-30. Env: /home/david/code/nsba/.venv/bin/python. Inputs (GAME data only — standard SB scoring, never combine as outcome): tossups_long.csv, games_meta.csv, canonical_players.csv, player_archetypes.csv, outputs/combine_to_game_linked.csv. Artifact: data/processed/playoff_robustness_31.csv (per-player reg vs playoff PPTF + archetype, nsba3).


LEAD WITH THE CAVEATS (read these first)

  1. Clean playoff games exist in nsba3 only. All 9 nsba1 playoff games are non-reconciling (corrupted sheets, dropped by the guardrail); nsba2 had no playoffs. Everything below is one season: 12 clean nsba3 playoff games, 265 playoff questions, 36–43 players. This is the same nsba3-only limitation that bounds F7/finding-20. Read magnitudes as ±a lot; trust the ranking of which subjects/archetypes fade.

  2. "Difficulty" is confounded with surviving-field strength. Playoff conversion is lower partly because packets are harder and partly because only the strong teams are left, so a weak buzzer faces a stronger opponent racing them to the buzzer. The two cannot be separated in n=1 season. The within-player retention (each player as their own control) is the cleanest cut but still cannot tell "the question got harder" from "my opponent got faster."

  3. Depth counts (#players above a bar) are NOT comparable across phases because the playoff pool is pre-filtered (85 reg players → 61 playoff players; the weak tail is already eliminated). The fraction above a bar mechanically rises in playoffs purely from this selection — it is not evidence the field got deeper. Treat (2) as a measurement caution, not a finding.

  4. nsba3 PPTF uses estimated/derived team-tossups-faced denominators (paired layout). PPTF here = player toss-up points ÷ team tossups faced in that phase.


(1) Per-subject conversion drops sharply — but only for 4 of 6 subjects

Question-level conversion (any team converts), nsba3 regular vs playoff, FDR-BH across the 6 subjects:

subject reg conv playoff conv rel. drop Fisher OR raw p FDR p
ess (Earth/Space) 0.888 0.689 −22% 3.56 0.002 0.012
m (Math) 0.830 0.674 −19% 2.36 0.036 0.108
cs (Comp Sci) 0.725 0.571 −21% 1.98 0.098 0.147
p (Physics) 0.859 0.745 −13% 2.09 0.076 0.147
ch (Chem) 0.795 0.830 +4% 0.80 0.681 0.681
b (Bio) 0.905 0.978 +8% 0.22 0.205 0.246

Only ESS survives FDR strictly (n=45 playoff questions). Math, CS, Physics are suggestive (OR ~2, raw p<0.10) but don't clear multiple-testing — the per-subject playoff n is 35–47 questions, genuinely underpowered. Bio and Chem do not drop at all (point estimates even rise; not significant either direction).

The signal: difficulty bites the quantitative/Earth-Space cluster (ESS, Math, CS, Physics) and spares the knowledge-recall cluster (Bio, Chem). This is consistent with finding-20's prior that CS/Math are the hardest subjects — under playoff pressure the hard subjects get harder while easy-recall subjects hold. (The aggregate playoff conversion drop 0.838→0.755, OR=0.59, reproduces F7.)


(2) "Depth" counts are not interpretable (selection) — reported for honesty

Players above a PPTF bar, by phase:

bar regular (n=85) playoff (n=61)
≥0.3 35 (41%) 32 (52%)
≥0.5 20 (24%) 21 (34%)
≥0.7 12 (14%) 11 (18%)

The fraction above each bar rises in playoffs. This is entirely a selection artifact — the playoff field is the surviving strong teams, the weak tail is gone. It does NOT mean the field is deeper when it's hard; if anything the opposite (see (3): individual production falls). Do not cite these as a depth result. The honest statement: the talent that reaches playoffs is pre-filtered, so "how deep is the bench under difficulty" cannot be answered from this data without a within-player design — which is (3).


(3) Difficulty-robust vs front-runner: by ARCHETYPE (the real finding)

Within-player retention = pooled playoff PPTF ÷ pooled regular PPTF, for the same players (each is their own control), aggregated by archetype with player-bootstrap 95% CIs. Overall retention is ~flat (0.98, CI [0.69, 1.27]) — the average surviving player roughly holds, but that average hides a strong split:

archetype n retention 95% CI reads as
elite multi-science (chem+phys core) 5 1.47 [1.16, 1.74] rises
math+phys specialist 12 1.07 [0.75, 1.48] holds
bio specialist 7 0.96 [0.24, 2.33] holds (noisy)
low-output / replacement 11 0.50 [−0.16, 1.53] fades (noisy)
ess specialist 8 0.44 [0.22, 0.63] collapses

Two CIs exclude 1.0: - ESS specialists collapse to ~44% of their regular output in playoffs. This is mechanically downstream of (1): ESS conversion craters 22%. ESS specialists have nowhere to hide — their one subject is the one that gets hardest. - Chem+phys-core elites rise to ~147%. As a subject they sit in the cluster (chem) that holds, plus they have the breadth to pick up the points the faders leave on the table. These are the difficulty-robust drafts.

Front-runners who faded (top regular producers, low retention): Akhil Batchu (ESS, 1.00→0.53), Vishnu M (ESS, 0.79→0.53), Peter B (ESS, 0.55→0.00), Theenash Sengupta (bio, 0.73→−0.08 — the one bio exception, drags the bio CI), Harry Gao (math+phys, 0.51→0.00), Eli Mrug (math+phys, 0.69→0.29).

Difficulty-robust (held or rose from a high base): Rohan G (chem+phys elite, 1.03→1.67), Anurag S (0.85→0.97), Kian Dhawan (0.70→0.87), Anish A (0.58→0.69), Edwin He (bio, 0.66→0.73), Lockheed Martin (math+phys, 0.80→1.32). The buy-low board's top game-backed buys (Rohan G, Mihir class) and its top fade (Vishnu M) line up with robustness — Rohan G is robust, Vishnu M is a fader, independently corroborating F12.

Caveat: bio (7) and low-output (11) CIs are wide and span 1.0 — bio "holds" rests heavily on Edwin/Arjun against Theenash's collapse; don't over-read the bio row. The two load-bearing rows (ESS down, chem+phys-elite up) are the only ones whose CIs exclude 1.


(4) Combine → game: does combine predict PLAYOFF performance better? NO.

The hypothesis (deep/de-biked theta should pick out playoff-robust players) is not supported. On the 36 nsba3 players with both a combine and ≥1 playoff game, combine predicts regular-season PPTF better than playoff PPTF, for every theta variant:

combine signal → reg PPTF (spearman) → playoff PPTF (spearman)
raw_overall 0.71 0.58
theta_overall 0.71 0.55
debiked_overall 0.60 0.42

The reg-minus-playoff correlation gap is +0.12 (raw) / +0.15 (theta); bootstrap 85–90% of resamples positive but CI includes 0 (n=36) — so the decay is directional, not significant. And combine does not predict retention at all (theta_overall → retention spearman 0.17, p=0.32; debiked 0.10, p=0.56).

De-biking does not help here — debiked theta is the weakest playoff predictor (0.42), consistent with F3 (de-biking only earns its keep on the biked nsba1/Energy era, and slightly hurts on clean seasons; nsba3 is clean).

Interpretation: the combine is a regular-season talent signal that decays, not sharpens, under playoff difficulty. It is not a hidden playoff-robustness detector. Whatever makes a player hold up when it's hard (breadth into the holding subjects, composure vs a fast surviving field) is not captured by combine theta.


Implications for drafting playoff-robust players

  1. The combine cannot find your playoff performers beyond finding good regular-season players generally — and it does that worse for playoffs. Do not pay a premium for a high combine expecting playoff upside; price combine as a regular-season signal (F3) and shrink it further for the games that decide titles.

  2. Archetype is a better playoff-robustness handle than combine score. The one repeatable structural signal: breadth (chem+phys-core multi-science) holds or rises; single-subject ESS specialists collapse because their subject is the one difficulty hits hardest. This is a refinement of F1/F4: coverage is table-stakes for the regular season, but depth-into-the-robust-subjects (Bio/Chem recall + breadth) is what survives playoff difficulty. A roster that is "covered" on paper but leans on an ESS or pure-quant specialist for its points is a regular-season team that fades in March.

  3. Tiebreaker, weakly held. With n=1 season and CIs that mostly span 1.0 outside the two flagged rows, this is a tiebreaker between similar players, not a re-ordering lever: prefer the broader chem/phys/bio-flavored producer over the equivalent-PPTF ESS/quant specialist for playoff equity, all else equal. Do not reach past the value board (PPTF/VORP, F2/F8) on this alone.


Limitations summary


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub