NSBA Draft Analyticsembargoed · 2026-06-06

91_causal_redteam.md

91 — Causal / Selection-Bias Red-Team

91 — Causal / Selection-Bias Red-Team

Role. Adversarial review of findings 10–20 for causal interpretation and selection effects. I do not re-litigate the metric math; I attack whether the correlations license the draft decisions and whether the samples that produced them are the samples we think they are. Where cheap, I ran the confound check myself against data/processed/ (env /home/david/code/nsba/.venv/bin/python).

Verdict in one line. The findings are unusually honest about small-n, and most already flag their own confounds. But four threats are under-weighted relative to how load-bearing the conclusion is, and three "robustness" claims rest on a single season or 1–2 teams. None of this overturns the core draft strategy (production rate + breadth + reach-for-CS); it should widen the error bars on the magnitudes and change how hard you lean on the nsba3-only results.

Numbers below are real (recomputed this session) unless labeled as a proposed test.


Empirical checks I ran (so these are not hypotheticals)

Check Result Bears on
Game-PPTF returners (≥2 game seasons) vs one-and-done, first-season PPTF (gp≥3) 0.288 vs 0.285, t-test p=0.96 — no positive selection F14 growth
Combine returners vs one-and-done, first-season debiked theta −0.179 vs +0.038 (returners start below avg, p=0.17) F14 growth (reverses the naive bias)
Players with multi-season game data 25 of 181 (24 have 2, 1 has 3) F14, F11, F13
Players with multi-season combine data 33 of 214 (29 have 2, 4 have 3) F14, F10
Punting teams (FLOOR=8, games≥5) by season 5/5 are nsba3; nsba1/nsba2 have zero F16, F18
corr(punts, win%) dropping the two 0%-win teams (usmic, okc) −0.54 → −0.32 (n=35) F16
corr(punts, win%) within nsba3 only −0.76 (n=13), holds F16
CS collinearity: corr(cs_ppg, non-cs_ppg) across nsba2/3 only +0.19 (mild — better than the finding's caveat implies) F17
corr(share_cs, win%) +0.42 (raises a reverse-causation flag — see T3) F17
Games per team: nsba3 min=1, median 7; 12/41 team-seasons have <8 games F15, F16, F18

Prioritized threats

T1 — [HIGH] The growth/development curve is a survivorship + regression-to-mean composite, and the bias runs the opposite way the finding assumes

Threatened conclusion (F14): "Returning players get better, +0.43 z/edition (p=0.002)" → used to justify modest year-over-year upgrades and to interpret the −0.52 mean-reversion as a buy-low signal.

Confound. Two coupled selection effects: 1. Survivorship. Only players who chose to return appear in any Δ. Returning is endogenous to satisfaction/engagement, which correlates with having done well — classic aging-curve survivorship. The finding flags this for the +growth direction. 2. The under-acknowledged twist. I checked who actually returns: combine returners started at theta −0.179 vs +0.038 for one-and-done. Returners are negatively selected on their first combine. With a −0.52 mean-reversion coefficient, a cohort that starts below the mean will mechanically drift upward regardless of any true development. So +0.43 z/edition is partly regression-to-the-mean of a below-average-but-returning cohort, not skill growth. The two stories (growth, reversion) are not independent — they are the same selected cohort seen twice.

Why it matters for the draft. The buy-low board (F19) and "expect modest growth from returners" both inherit this. If the "growth" is mostly reversion, then (a) it should not be added on top of a reversion-shrink (that double-counts), and (b) it does not generalize to a new nsba4 entrant with no prior — you only get the reversion bump if they posted an extreme prior.

Tests / corrections. - Re-estimate growth with a mean-reversion control: regress Δtheta on prior-theta and an edition trend; the trend coefficient net of the −0.49 reversion slope is the real development estimate. (Finding 14 reports them separately but never nets them.) - Two-season-only floor: restrict to players present in consecutive editions and report the within-player fixed-effect trend, not pooled Δ. - Matched non-returner counterfactual: you cannot observe non-returners' next season, but you can bound the bias by reweighting returners to the full first-season theta distribution (inverse-propensity on return).


T2 — [HIGH] The coverage / "don't punt a subject" headline is a 5-team, single-season result wearing a 37-team coat

Threatened conclusion (F16, echoed F18): "0-punt teams 52% vs punting 28%, corr(win%, #punts) = −0.54, survives a star-talent control" → the headline generalist/specialist takeaway and a primary draft rule.

Confound / fragility. At the main FLOOR=8 threshold, every punting team (5 of them) is nsba3; nsba1 and nsba2 have zero punting teams at games≥5. So: - The cross-season "n=37" framing overstates independent evidence: the contrast (punt vs no-punt) is entirely within nsba3, n=13. - It is leverage-sensitive. Dropping the two 0%-win teams (usmic 4-punt/0%, okc 2-punt/20%) takes corr(punts,win%) from −0.54 to −0.32 (n=35). usmic alone is doing a lot of work — and a 4-punt team is also just a bad/short-schedule team, so "punting" and "weak roster" are not separable in that point. - The finding's own "robustness" (raising the floor strengthens the effect to −0.60) is not independent — it is the same nsba3 teams crossing a moving threshold, plus nsba1/nsba2 teams only entering as you lower the credible-answerer bar enough to manufacture punts that didn't exist in-game.

Mitigant (credit where due). Within nsba3 the effect is strong and survives dropping usmic (−0.76 → −0.68), and the partial-correlation-over-talent (−0.37) is a real control. So the direction is sound. The problem is treating the magnitude ("≈7 win-pts per punt") and the cross-season generality as established.

Tests / corrections. - Report the coverage result explicitly as nsba3-only with n=13, and show the leverage diagnostic (drop-one-team DFBETA) in the finding. - Design-coverage instead of realized-coverage. The current "covered = ≥2 net-correct in games" is mechanically tied to games-played and to being good. Rebuild coverage from combine ability per drafted player (a roster-design measure available pre-game) and re-test — this breaks the "good teams convert ≥2 everywhere" tautology. F18 explicitly leaves this as future work; it is the single most valuable robustness check in the set because F16/F18 both depend on it.


T3 — [HIGH] CS "points are worth 3.5× wins" risks reverse causation: winning the buzzer race produces CS conversions

Threatened conclusion (F17): "CS points worth ~3.5× generic points toward winning; reach for a CS answerer" → an aggressive, specific draft action (reach earlier than raw value).

Confound. Two distinct reverse-causation channels the WLS controls do not fully close: 1. Buzzer-race endogeneity. A team that is faster/stronger overall wins more toss-ups in every subject, including CS. CS is the hardest subject (conversion 70%, F20), so it is precisely where the strong, fast team's edge over a weak opponent is largest (more dead questions for the weak team to fail to convert). So "more CS points" can be a symptom of being the better team, and "CS → wins" is partly "good → CS." I find corr(share_cs, win%) = +0.42, consistent with this. 2. Scarcity-coverage tautology. CS wins are leveraged "because opponents punt CS" (the finding's own mechanism). But that means the CS premium is conditional on the field continuing to under-supply CS. If the nsba4 field reads the same analysis and everyone drafts a CS answerer, the 3.5× premium compresses — this is a strategy that erodes as it is adopted (a market-timing, not a structural, edge).

Mitigant. I did find the collinearity the finding worried about is mild (cs_ppg vs non-cs_ppg r=0.19), so the partial coefficient is cleaner than the caveat suggested. The conversion-gap falsifier (CS-mains 62% vs generalists 49%, p=0.011) is a player-level result not subject to the team buzzer-race confound, and is the strongest leg — lead with that, not the team 3.5×.

Tests / corrections. - Re-estimate the CS win coefficient with opponent-strength / SOS on the RHS (or as team fixed effects across the two seasons) to absorb "we were just the better team." - Re-price CS value conditional on field CS-supply (interact CS pts/game with a season-level field-CS-coverage index) to show how fast the premium decays if everyone covers CS — directly relevant since F18 shows the draft already covers CS 96.6% of the time. - Lean the draft recommendation on the player-level conversion gap + combine theta_cs r=0.60 (both confound-robust), and treat the team 3.5× as an upper bound.


T4 — [MED-HIGH] The playoff difficulty factor (OR=0.59) conflates harder packets with stronger surviving fields

Threatened conclusion (F20): "Playoffs significantly harder, OR=0.59, p=0.002; multiply playoff rates by 1.11." Used to up-adjust playoff production and feeds the combine→game playoff degradation in F12.

Confound. Playoff games are played by survivors — the teams good enough to qualify. Lower conversion in playoffs is jointly (a) harder packets and (b) two strong teams facing each other so neither gets the easy buzzes a strong team gets vs a weak one. The factor cannot separate these, and it is nsba3-only (all nsba1 playoff sheets are corrupted/excluded, nsba2 has none). The finding states this caveat but still ships a numeric 1.11 factor that downstream analyses (F12 playoff degradation) treat as packet difficulty.

Tests / corrections. - Hold field strength fixed: restrict to the same strong teams' regular-season vs playoff conversion (they exist in both), or add team-pair fixed effects. If the gap survives within strong-team matchups, it is packet difficulty; if it shrinks, it was field selection. - Until then, label 1.11 as a playoff-context factor (selection + difficulty combined), not a pure park factor, and do not use it to "rescue" combine→playoff degradation as evidence of gaming (F12 §4) — the degradation could be field selection.


T5 — [MED] Combine→game validation (r≈0.62) is selected on players who kept playing, biasing the predictor toward looking honest

Threatened conclusion (F12): combine moderately predicts game PPTF; raw combine is the best single predictor (CV R²≈0.35).

Confound. The 102 linked player-seasons are exactly the players who (a) took the combine and (b) showed up to ≥2 clean games. A biker who games the combine and then doesn't play / quietly tanks games is missing from this correlation. So the linked cohort is enriched for "combine score → real play" honest types, which inflates the apparent combine→game validity for the population you actually draft from (which includes no-shows and tankers). This is a who-shows-up selection that cuts against the very thing the analysis is for (predicting players you haven't seen play).

Mitigant. The finding's tanker stress test (suspected tankers under-deliver) points the right way — but n=2. The selection bias and the tiny tanker n are the same gap: the people who break the combine→game link are mostly absent from the data.

Tests / corrections. - Quantify the missingness: how many nsba4 combine entrants have zero game history, and how does their combine distribution compare to the linked cohort? If high-combine no-game players exist, the validity estimate is optimistic for them specifically. (F13 already notes only 18/181 board players carry an nsba4 combine row — the inverse of this gap.) - Report combine→game r with a selection caveat band, and prefer the de-biked combine for unproven players (the population most subject to this bias).


T6 — [MED] Win% as the universal outcome is schedule-confounded; round-robins are not balanced and nsba3 schedules are short/uneven

Threatened conclusion: all team-level results (F15 north-star, F16 coverage, F17 CS, F18 breadth) regress on win%, treated as a clean talent outcome.

Confound. (a) Schedules are uneven — nsba3 ranges 1–10 games (median 7); 12/41 team-seasons have <8 games, and short schedules make win% high-variance and non-comparable across teams. (b) In any non-complete round-robin, win% encodes whom you played — "winners get easier remaining schedules" and unbalanced opponent draws are unmodeled. The findings acknowledge "schedule strength not modeled" but then quote win% correlations as if SOS-neutral. points_against std balloons from 87 (nsba1) to 200 (nsba3), i.e. nsba3 opponents faced are wildly uneven.

Tests / corrections. - Add a minimal SOS adjustment (Bradley-Terry / iterated win% of opponents) and re-run the headline win% correlations; or weight team-seasons by games and report whether F15/F16/F17 magnitudes move. - For the metric ranking in F15 this is minor (point_margin already dominates and is excluded as illegal). For F16/F17 magnitudes it compounds T2/T3.


T7 — [MED] The scibowl(Energy) baseline → NSBA(CS) transfer is asserted on 5 shared subjects but used to price the 6th

Threatened conclusion (F18, F16): NSBA breadth→win replicates the scibowl F1 baseline (corr +0.52→+0.54), and the cost of a coverage gap (~16 win-pts/subject) is priced from scibowl because NSBA has almost no gaps.

Confound. The natural baseline's 6th subject is Energy; NSBA's is CS. Energy (nsba1-era) has no combine analog and is a different skill market than CS (scarce, unbikable, hardest-to-convert). Pricing an NSBA CS gap using a scibowl OLS whose 6th slot is Energy assumes the marginal subject is interchangeable — but F17's whole thesis is that CS is not interchangeable (3.5× value). So F18's "a gap costs ~16 pts" is internally in tension with F17's "a CS gap costs far more than a generic subject." The two findings should not both be true at face value.

Tests / corrections. - Price the coverage-gap cost on the 5 shared subjects only, and report the CS gap cost separately from F17 (don't borrow the Energy-inclusive scibowl slope for CS). - Reconcile F17 vs F18 explicitly: F18 says "draft solves coverage incl. CS (96.6%)"; F17 says "reach for CS." Both can hold only if the message is "coverage is table stakes the draft already provides, but CS depth/speed beyond bare coverage is where the leveraged value is" — state that joint reading.


T8 — [LOW-MED] nsba3 estimated-TUH contaminates every PPTF-based result, and nsba3 is also where the coverage/CS/difficulty signals live

Threatened conclusion: PPTF reliability (F11), value board (F13), north-star (F15) all use nsba3 PPTF with a season-mean TUH proxy (paired layout records no per-player heard count).

Confound. This is flagged everywhere individually, but the aggregate point isn't made: nsba3 is simultaneously (a) the season with the proxy denominator, (b) the only season carrying the playoff and coverage-punt variation, and (c) the season most comparable to the nsba4 draft target (F20). So the most decision-relevant season is also the noisiest-measured one, and several "robust across seasons" claims are really "robust where the data is exact (nsba1/2) + suggestive where it matters (nsba3)."

Tests / corrections. - Sensitivity: re-run F13/F15 with nsba3 down-weighted (or excluded) and report how the value board top-N and north-star ranking move. If the draftable conclusions are stable without nsba3, good; if they hinge on nsba3, the TUH proxy uncertainty must propagate into VORP CIs. - Where possible, replace estimated TUH with buzz-count-denominated rates (conversion, points-per-game) that F11's split-half showed agree (0.82) and are denominator-free.


What is not a serious threat (to avoid over-correcting)


Recommended robustness checks, ranked by value-of-information for the draft

  1. Design-coverage rebuild (T2): replace realized-in-game coverage with combine-ability-per-roster coverage and re-test F16/F18. Breaks the biggest tautology; both headline coverage findings depend on it.
  2. Net growth of reversion (T1): report the development trend net of the −0.49 reversion slope, on consecutive-edition within-player fixed effects. Decides whether "expect growth" is real or double-counts reversion in the buy-low board.
  3. SOS-adjusted win% (T6) re-run of F16/F17 magnitudes. Cheap; recalibrates the two most aggressive draft actions.
  4. Field-CS-supply interaction (T3): show how fast the CS premium decays as the field covers CS — turns "reach for CS" into a calibrated "how far to reach."
  5. nsba3 down-weight sensitivity (T8) on the value board top-N and north-star ranking.
  6. Within-strong-team playoff conversion (T4) to separate packet difficulty from field selection before trusting the 1.11 factor.
  7. Combine→game missingness audit (T5): combine distribution of zero-game nsba4 entrants vs the linked cohort.

Caveats on this review


NSBA Draft Analytics · embargoed until after the SSB draft · ← hub