92 — Skeptic Audit: Archetypes / Duo Synergy / Undervaluation / Field Depth (F28–F31)
92 — Skeptic Audit: Archetypes / Duo Synergy / Undervaluation / Field Depth (F28–F31)
Date: 2026-05-30. Role: methodology adversary. Scope: re-read F28–F31 and their artifacts, re-ran spot checks. Bottom line up front: three of the four findings are honestly caveated and survive adversarial probing as directional claims; F30's headline validation is largely circular and is overstated — ~90% of its flagship ρ=0.60 is a mechanical shared-PPTF artifact, not validation of the combine→value mechanism.
All n's are tiny (344 player-seasons, 38 team-seasons, 1 clean playoff season). Reproduced
F28, F29 (tables), F31 (artifact recompute) exactly. F31's script A31*.py is missing
from scripts/ — only the artifact playoff_robustness_31.csv exists, so F31 is
partially non-reproducible (I recomputed from the artifact and it matches the doc).
VERDICTS
| Finding | Verdict | One-line reason |
|---|---|---|
| F28 archetypes | watchlist | Reproduces; honest about silhouette 0.23. But stability is overstated by the wrong metric — feature-perturbation ARI drops to ~0.52, and the load-bearing elite cluster has only median 0.64 membership recovery. |
| F29 duo synergy | trust | Best-built of the batch. Pseudo-replication correctly handled (team-level n=38), FDR honest, the one nominal hit explained as additive inheritance with a clean interaction test. Bounded null stands. |
| F30 undervaluation | discard the validation; watchlist the board | The flagship ρ=0.60 is ~90% a shared-+PPTF artifact; it survives randomizing the combine signal (ρ=0.54) and combine's own independent link to slot-residual is 0.007, p=0.95. "Validates the F8 mechanism per-player" is not supported. |
| F31 field depth | watchlist | nsba3-only (n=1 season). The two flagged rows are robust to leave-one-out within that season, but the elite "rise to 147%" rests on n=5 and zero cross-season replication. Correctly caveated; don't elevate above tiebreaker. |
F28 — Archetypes: WATCHLIST
Reproduced exactly. Silhouette 0.234 @ k=5, seed ARI 0.986, bootstrap ARI 0.844, cluster n's and profiles match the doc to the decimal.
Where the doc oversells stability. The doc leans on seed-ARI (0.99) and bootstrap-ARI (0.84). Seed-ARI is near-tautological (identical data + features, different init). I ran the test the doc didn't: leave-one-subject-out re-clustering:
drop bio: ARI=0.53 drop ess: ARI=0.52 drop math: 0.66
drop chem: 0.71 drop phys: 0.85 drop cs: 0.94
Dropping bio or ESS reshuffles ~half the cluster assignments. So the partition is materially dependent on which features are in the vector — exactly the analyst-degrees-of-freedom surface the task flagged. The "five archetypes" are a defensible description of this 6-dim choice, not a subject-invariant structure.
The load-bearing cluster is the shakiest in membership. A2 "elite multi-science" is the only archetype claimed to reliably win and carries essentially all the VORP. Tracking A2 membership across 200 bootstraps (identifying the elite pod by max-chem centroid each time): median elite-recovery rate 0.64, min 0.14, only 39% of A2 members land elite in >70% of resamples. The direction (a broad-and-deep top tier exists and wins) is safe; the roster of who is "elite" churns more than bootstrap-ARI 0.84 implies. The doc's "trust the directions, not the labels of boundary players" is correct but undersells how many A2 members are effectively boundary.
Confidence flag: the doc's Low-Medium is right. The CS-has-no-standalone-archetype claim is fine (it's an absence claim, robust to the soft boundaries). Treat archetype labels attached to individual players (which propagate into F30's board and F31's retention split) as soft inputs, not facts.
F29 — Duo synergy: TRUST
Best-engineered finding in the batch. The central trap (474 pairs inside 38 team-seasons,
pairs sharing one outcome) is identified and defused: all clean inference is at the team level
(n=38) or via team-season cluster bootstrap. Verified duo_pairs.csv = 474 pairs / 38
team-seasons; the archetype-pair table reproduces.
The one nominal hit, elite+repl q=0.037, is correctly diagnosed as additive inheritance
(resid_sum is additive by construction, so any pair containing an under-predicted elite
inherits a positive residual) rather than interaction. The decisive interaction test — does a
non-elite's own residual rise with an elite teammate? — is +0.135 ws, p=0.155, ns. That
is the right test and it kills the synergy reading. FDR across the 8-test synergy family: 0/8
survive. The wrong-signed disjoint-pair "effect" is honestly flagged as a talent-density
confound, not a coverage signal.
Caveat that holds: this is a bounded null — a true pair synergy below ~±0.3 win-shares is invisible at n=38. Stated correctly. No overreach found.
F30 — Undervaluation: DISCARD the validation, WATCHLIST the board
The metric is uvf_combine = pct(our_PPTF) − pct(combine_rank). The validation target
val_resid = realized PPTF − pick-slot regression (F8's steal score). Both terms contain
+realized_PPTF with a positive sign. The correlation between them is therefore
mechanically positive irrespective of whether combine carries any independent information.
I quantified the circularity (n=94):
HEADLINE rho(uvf_combine, val_resid) = 0.599 (the doc's flagship)
RANDOMIZE the combine signal -> rho(uvf_FAKE,..) = 0.540 <- 90% of it survives pure noise
corr(our_pct_pptf ALONE, val_resid) = 0.709 <- PPTF alone beats the full metric
corr(field_pct_combine, val_resid) = 0.007 p=0.95 <- combine's OWN link: NULL
Reading: shuffle the combine ranks into random noise and the headline correlation barely
moves (0.60 → 0.54). PPTF by itself — the term shared by both sides — correlates with
val_resid more strongly (0.71) than the full metric does. And the genuinely independent
question, "does combine rank predict the ADP-slot residual?", is a flat zero (0.007,
p=0.95). So the ρ=0.60 is not cross-validation of the combine→value mechanism; it is a
re-expression of "players with high realized PPTF have positive PPTF-residuals," which is
close to a tautology.
The doc's defense — "uvf_combine uses combine rank, val_resid uses ADP; different field
proxies, so this is real cross-validation not an identity" — is incorrect. The two
field proxies differ, but the shared +PPTF outcome term, not the proxy difference, drives
the correlation, and the proxies are themselves 0.67-redundant (combine pct vs ADP pct). The
claim "validating the F8 mechanism per-player" is stated more confidently than the data
supports. F30 does not independently validate F8; the per-player decomposition is real
but its headline statistic is an artifact.
What survives: F30 honestly flags same-season leakage and that "this is a re-expression of F8, not a new effect," and FDR finds 0/94 individually significant names. The forward NSBA4 board is a separate, internally-consistent exploit of "the field anchors on raw combine" (no PPTF involved there, it's combine-only) — keep it as a watchlist re-ordering device within tiers, but it inherits soft archetype labels (F28) and the F3 selection-bias caveat, and no name is significant. Drop the ρ=0.60 "validation" as evidence; do not cite it as confirming F8.
F31 — Field depth under difficulty: WATCHLIST
Recomputed from playoff_robustness_31.csv (the script is missing — reproducibility gap).
The two flagged rows reproduce: ESS retention 0.44, CI [0.19,0.62]; elite 1.47, CI
[1.16,1.74] — both exclude 1.0.
Leave-one-out (within nsba3) is reassuring: ESS stays 0.40–0.50 dropping any one of its 8 players; elite stays 1.24–1.55 dropping any one of its 5. So neither flagged effect is a single-player artifact within the season.
But the binding limitation is n=1 season, which no within-season robustness can fix. The bootstrap CIs resample players, not seasons, so they understate true uncertainty — there is zero out-of-sample replication and the difficulty-vs-surviving-field-strength confound is unidentified. The elite "+47%" rests on 5 players and a ratio-of-pooled-means that a couple of high-TUH playoff games can swing. Results (1) per-subject ESS drop and (3) ESS-specialist collapse are also the same underlying fact (the doc says so), not two independent confirmations. The doc grades this Low-Med and explicitly demotes it to a "tiebreaker, weakly held" — that is the correct altitude. Don't let it harden into a drafting lever.
Cross-cutting flags for the spine
- Soft archetype labels propagate. F28's labels feed F30's "bio underdrafted" / per-archetype board and F31's retention split. Both inherit the silhouette-0.23 / feature-sensitive softness. Per-archetype sub-tables (n=5–13 cells) are descriptive only.
- F30's validation should be removed from the evidence base for F8. F8's ~13 toss-pt/slot edge stands on its own design; F30 does not add an independent confirmation and its headline ρ should not be cited as one.
- Reproducibility gap: restore/commit the F31 script (
A31_*.py). Artifact-only findings on n=1 season are hard to trust without the pipeline. - What I'd trust to act on: F29's bounded null (no duo lever) and the direction of F28 (a broad-and-deep elite tier wins; specialists incl. CS are interchangeable coverage). Everything else is watchlist or weaker.