14b — Growth NET of mean-reversion (R3 revision of F14, red-team T1)

14b — Growth NET of mean-reversion (R3 revision of F14, red-team T1)
Date: 2026-05-30
Supersedes the development half of 14_growth_aging.md (the Gideon roster
hypotheses there are unchanged). Addresses causal red-team T1
(91_causal_redteam.md): the headline "+0.43 z/edition growth" is partly
regression-to-the-mean of a negatively-selected returning cohort (returners
start below the field mean), so growth and reversion are the same selected
cohort seen twice and must not be double-counted.
Inputs: data/processed/combine_ability.csv (255 combine rows, theta is
z-scored within season), data/processed/player_season_master.csv (207
game player-seasons).
Script: scripts/growth_net_revised.py.
Artifact: data/processed/growth_net_revised.csv (41 debiked-theta transitions).
LEAD CAVEATS (read first)
- Tiny, survivorship-selected samples. Development can only be measured on players who returned: 33 of 214 combine players have multi-season combine data (29×2 seasons, 4×3); 25 of 181 players have multi-season game data (24×2, 1×3). Everything below is 22–41 transitions. CIs are wide.
- Theta is standardized WITHIN season (each season re-centered to mean ≈ 0,
std ≈ 1 — confirmed: nsba1..4 means −0.00/0.00/+0.06/+0.15). So "+0.43 z/edition"
is a relative-rank climb within that year's field, not an absolute skill
gain on a fixed scale. There is no cross-season level trend to regress on — a
per-edition trend term (
edn_from) is ill-posed here and was the wrong test. - The development metric is combine-heavy and nsba4-heavy. 24 of 30 consecutive combine transitions are nsba3→nsba4, and nsba4 has no game data — so most of the "growth" is a combine→combine relative climb into the draft-target season, not demonstrated on-floor improvement.
- The combine is gameable (F10). A combine-theta "rise" can be a player tanking less, not getting better.
The T1 charge, tested
Returners are negatively selected on their first combine — confirmed:
| metric (first season) | returners (multi-season) | one-and-done | Welch t | p |
|---|---|---|---|---|
| combine debiked theta | −0.179 (n=33) | +0.038 (n=181) | −1.39 | 0.17 |
| combine raw theta | −0.136 (n=33) | +0.010 (n=181) | −1.13 | 0.27 |
| game PPTF (gp≥3) | +0.290 (n=22) | +0.285 (n=133) | 0.12 | 0.91 |
So on the combine side the red-team is right: returners start ~0.18 z below the field. With a reversion slope near −0.45, a below-mean cohort drifts upward mechanically. (On the game side there is no such selection — game returners start exactly at the field mean, p=0.91 — so the game-PPTF growth was never a reversion artifact. This matches the red-team's own note that PPTF survivorship is not a serious threat.)
The fix: net development = reversion-purged intercept
The right decomposition is not "trend net of a separately-reported reversion slope" (the two are entangled and theta has no cross-season level), but a single regression of the consecutive-edition within-player change on the prior level:
Δtheta = β0 + β1 · (prior theta)
- β1 is the mean-reversion slope (negative).
- β0 is the net development — the expected gain for a returner who started at the field mean (prior = 0), i.e. with the reversion bump removed.
Results (consecutive-edition pairs, within-player)
| metric | net development β0 (prior=0) | p | reversion slope β1 | naive mean Δ | n pairs |
|---|---|---|---|---|---|
| combine debiked theta | +0.41 z | 0.006 | −0.43 | +0.46 | 30 |
| combine raw theta | +0.25 z | 0.062 | −0.58 | +0.28 | 30 |
| game PPTF (gp≥3) | +0.22 | 0.028 | −0.44 | +0.09* | 22 |
*The naive game-PPTF mean is small because the game cohort isn't negatively selected (no reversion lift to subtract); its net is larger than its naive pooled-Δ because the pooled-Δ mixed in long 2-edition gaps.
Cluster-bootstrap 95% CI on net development (debiked theta, resample players): [+0.13, +0.68] z, mean +0.40.
Reconciliation with the "100% reversion" decomposition
Evaluating the reversion-only model at the cohort's mean starting prior (−0.16) recovers ~100% of the naive mean Δ by construction — that is just the regression passing through the centroid and says nothing about development. The decision-relevant quantity is the intercept at prior = 0, which is +0.41 and survives. So: reversion inflates the pooled number modestly (raw +0.46 → net +0.41 for debiked theta), but a real, positive net development of ≈ +0.4 z (combine) / +0.2 PPTF (game) per edition remains after purging reversion.
Within-player fixed-effects cross-check
Demeaning theta and edition within each player and regressing gives a positive trend (debiked +0.38/edition, t=4.88; raw +0.25, t=3.59; PPTF +0.07, t=1.98, p=0.054). FE agrees with the intercept method in sign and rough magnitude, confirming the net signal is not a pooling artifact. (FE on a within-season-z metric measures the same "climb the field" quantity, so do not over-read its tiny p-value — it shares the 33-player sample.)
Where the signal lives (heterogeneity, all thin)
- Game-era transitions only (nsba1/2 starts, n=6): net +0.95 (p=0.06).
- nsba3→nsba4 only (n=24): net +0.29 (p=0.06), reversion slope −0.46. Both subsets positive; neither alone clears p<0.05. The headline rests on the pooled 30.
What changed vs F14
| F14 (original) | 14b (this revision) | |
|---|---|---|
| Combine growth | +0.43 z/edition (naive pooled Δ) | net of reversion +0.41 z (intercept, p=0.006); raw-theta net +0.25 (p=0.06) |
| Reversion | reported separately (−0.52 corr / −0.49 slope) | −0.43 slope, now subtracted from growth, not reported alongside it |
| Interpretation | "returners get better" | returners get better on a relative-rank scale, partly but not mostly reversion; ~10% of the pooled combine number was reversion lift |
| Game PPTF | +0.086 (p=0.10), looked weak | net +0.22 (p=0.028) once long-gap pairs dropped — the cleaner, confound-robust leg |
The qualitative bottom line ("expect modest growth from real returners; don't overpay") survives — the net effect is real and positive — but the magnitude shrinks and the mechanism is now correctly separated from reversion.
Draft implications (the load-bearing part)
- Do NOT add a growth bump on top of a reversion shrink in the buy-low board
(F19) — that double-counts. The buy-low board already prices a returner's
low combine as a buy by leaning on the −0.43..−0.49 reversion pull. The growth
intercept (+0.41) is the gain for a returner starting at the mean; for a
returner who posted a low prior, the reversion term already delivers most
of their projected rise. Applying reversion-shrinkage AND a separate "+0.4 z
growth bump" to the same low-combine returner counts the same climb twice. Use
one model:
projected = prior + β1·prior + β0with β1≈−0.43, β0≈+0.41 — the reversion and growth terms come out of the same regression and are already net of each other. - Growth does NOT transfer to a brand-new nsba4 entrant with no prior. β0 is a returner effect estimated on players seen ≥2 editions. A first-time nsba4 player has no prior to revert and no within-player trajectory — you cannot credit them +0.4 z of "development." Price them off their combine theta (shrunk toward the pool mean per F10's EAP), full stop.
- The reversion lever is the robust, actionable one (it survives in both the selection check and the slope): shrink the biggest combine scores hardest; treat a real returning player's low combine as a buy. The growth term is a smaller, second-order add that is already inside the reversion regression.
- Trust the game-PPTF net (+0.22, p=0.028) over the combine net where a player has game tape — it is not reversion-driven (no returner selection) and is on the real-scoring scale. The combine net is on a gameable, within-season-relative metric and is nsba4-combine-heavy.
Limitations
- n = 22–41 transitions; 33/214 combine and 25/181 game multi-season players. Survivorship is severe and endogenous to returning — we cannot observe the development of the 181 players who did not return, and reweighting to the full first-season distribution (inverse-propensity on return) is not done here (the return-propensity model would itself be fit on 33 positives).
- Within-season standardization makes all combine "growth" a relative-rank quantity; if the whole field improved year-over-year, that common gain is invisible by construction. The net intercept measures relative climb only.
- nsba4 has no games, so 24/30 consecutive combine transitions cannot be validated against on-floor production; the game-PPTF leg (nsba1↔nsba2↔nsba3) is the only one anchored to real play.
- The combine is gameable; a theta rise can be reduced tanking, not skill.
- Net development is positive but its CI ([+0.13, +0.68] z) is wide enough that the practical instruction is "small positive, don't lean hard," not a precise number.
Reproduce
/home/david/code/nsba/.venv/bin/python scripts/growth_net_revised.py
Prints the survivorship counts, selection table, naive/reversion/net models for
debiked theta, raw theta, and game PPTF, the within-player FE cross-checks, the
decomposition, and writes data/processed/growth_net_revised.csv.