NSBA Draft Analyticsembargoed · 2026-06-06

92_archetype_skeptic.md

92 — Skeptic Audit: Archetypes / Duo Synergy / Undervaluation / Field Depth (F28–F31)

92 — Skeptic Audit: Archetypes / Duo Synergy / Undervaluation / Field Depth (F28–F31)

Date: 2026-05-30. Role: methodology adversary. Scope: re-read F28–F31 and their artifacts, re-ran spot checks. Bottom line up front: three of the four findings are honestly caveated and survive adversarial probing as directional claims; F30's headline validation is largely circular and is overstated — ~90% of its flagship ρ=0.60 is a mechanical shared-PPTF artifact, not validation of the combine→value mechanism.

All n's are tiny (344 player-seasons, 38 team-seasons, 1 clean playoff season). Reproduced F28, F29 (tables), F31 (artifact recompute) exactly. F31's script A31*.py is missing from scripts/ — only the artifact playoff_robustness_31.csv exists, so F31 is partially non-reproducible (I recomputed from the artifact and it matches the doc).


VERDICTS

Finding Verdict One-line reason
F28 archetypes watchlist Reproduces; honest about silhouette 0.23. But stability is overstated by the wrong metric — feature-perturbation ARI drops to ~0.52, and the load-bearing elite cluster has only median 0.64 membership recovery.
F29 duo synergy trust Best-built of the batch. Pseudo-replication correctly handled (team-level n=38), FDR honest, the one nominal hit explained as additive inheritance with a clean interaction test. Bounded null stands.
F30 undervaluation discard the validation; watchlist the board The flagship ρ=0.60 is ~90% a shared-+PPTF artifact; it survives randomizing the combine signal (ρ=0.54) and combine's own independent link to slot-residual is 0.007, p=0.95. "Validates the F8 mechanism per-player" is not supported.
F31 field depth watchlist nsba3-only (n=1 season). The two flagged rows are robust to leave-one-out within that season, but the elite "rise to 147%" rests on n=5 and zero cross-season replication. Correctly caveated; don't elevate above tiebreaker.

F28 — Archetypes: WATCHLIST

Reproduced exactly. Silhouette 0.234 @ k=5, seed ARI 0.986, bootstrap ARI 0.844, cluster n's and profiles match the doc to the decimal.

Where the doc oversells stability. The doc leans on seed-ARI (0.99) and bootstrap-ARI (0.84). Seed-ARI is near-tautological (identical data + features, different init). I ran the test the doc didn't: leave-one-subject-out re-clustering:

drop bio: ARI=0.53   drop ess: ARI=0.52   drop math: 0.66
drop chem: 0.71      drop phys: 0.85      drop cs: 0.94

Dropping bio or ESS reshuffles ~half the cluster assignments. So the partition is materially dependent on which features are in the vector — exactly the analyst-degrees-of-freedom surface the task flagged. The "five archetypes" are a defensible description of this 6-dim choice, not a subject-invariant structure.

The load-bearing cluster is the shakiest in membership. A2 "elite multi-science" is the only archetype claimed to reliably win and carries essentially all the VORP. Tracking A2 membership across 200 bootstraps (identifying the elite pod by max-chem centroid each time): median elite-recovery rate 0.64, min 0.14, only 39% of A2 members land elite in >70% of resamples. The direction (a broad-and-deep top tier exists and wins) is safe; the roster of who is "elite" churns more than bootstrap-ARI 0.84 implies. The doc's "trust the directions, not the labels of boundary players" is correct but undersells how many A2 members are effectively boundary.

Confidence flag: the doc's Low-Medium is right. The CS-has-no-standalone-archetype claim is fine (it's an absence claim, robust to the soft boundaries). Treat archetype labels attached to individual players (which propagate into F30's board and F31's retention split) as soft inputs, not facts.

F29 — Duo synergy: TRUST

Best-engineered finding in the batch. The central trap (474 pairs inside 38 team-seasons, pairs sharing one outcome) is identified and defused: all clean inference is at the team level (n=38) or via team-season cluster bootstrap. Verified duo_pairs.csv = 474 pairs / 38 team-seasons; the archetype-pair table reproduces.

The one nominal hit, elite+repl q=0.037, is correctly diagnosed as additive inheritance (resid_sum is additive by construction, so any pair containing an under-predicted elite inherits a positive residual) rather than interaction. The decisive interaction test — does a non-elite's own residual rise with an elite teammate? — is +0.135 ws, p=0.155, ns. That is the right test and it kills the synergy reading. FDR across the 8-test synergy family: 0/8 survive. The wrong-signed disjoint-pair "effect" is honestly flagged as a talent-density confound, not a coverage signal.

Caveat that holds: this is a bounded null — a true pair synergy below ~±0.3 win-shares is invisible at n=38. Stated correctly. No overreach found.

F30 — Undervaluation: DISCARD the validation, WATCHLIST the board

The metric is uvf_combine = pct(our_PPTF) − pct(combine_rank). The validation target val_resid = realized PPTF − pick-slot regression (F8's steal score). Both terms contain +realized_PPTF with a positive sign. The correlation between them is therefore mechanically positive irrespective of whether combine carries any independent information.

I quantified the circularity (n=94):

HEADLINE  rho(uvf_combine, val_resid)            = 0.599   (the doc's flagship)
RANDOMIZE the combine signal -> rho(uvf_FAKE,..)  = 0.540   <- 90% of it survives pure noise
corr(our_pct_pptf ALONE, val_resid)               = 0.709   <- PPTF alone beats the full metric
corr(field_pct_combine, val_resid)                = 0.007  p=0.95  <- combine's OWN link: NULL

Reading: shuffle the combine ranks into random noise and the headline correlation barely moves (0.60 → 0.54). PPTF by itself — the term shared by both sides — correlates with val_resid more strongly (0.71) than the full metric does. And the genuinely independent question, "does combine rank predict the ADP-slot residual?", is a flat zero (0.007, p=0.95). So the ρ=0.60 is not cross-validation of the combine→value mechanism; it is a re-expression of "players with high realized PPTF have positive PPTF-residuals," which is close to a tautology.

The doc's defense — "uvf_combine uses combine rank, val_resid uses ADP; different field proxies, so this is real cross-validation not an identity" — is incorrect. The two field proxies differ, but the shared +PPTF outcome term, not the proxy difference, drives the correlation, and the proxies are themselves 0.67-redundant (combine pct vs ADP pct). The claim "validating the F8 mechanism per-player" is stated more confidently than the data supports. F30 does not independently validate F8; the per-player decomposition is real but its headline statistic is an artifact.

What survives: F30 honestly flags same-season leakage and that "this is a re-expression of F8, not a new effect," and FDR finds 0/94 individually significant names. The forward NSBA4 board is a separate, internally-consistent exploit of "the field anchors on raw combine" (no PPTF involved there, it's combine-only) — keep it as a watchlist re-ordering device within tiers, but it inherits soft archetype labels (F28) and the F3 selection-bias caveat, and no name is significant. Drop the ρ=0.60 "validation" as evidence; do not cite it as confirming F8.

F31 — Field depth under difficulty: WATCHLIST

Recomputed from playoff_robustness_31.csv (the script is missing — reproducibility gap). The two flagged rows reproduce: ESS retention 0.44, CI [0.19,0.62]; elite 1.47, CI [1.16,1.74] — both exclude 1.0.

Leave-one-out (within nsba3) is reassuring: ESS stays 0.40–0.50 dropping any one of its 8 players; elite stays 1.24–1.55 dropping any one of its 5. So neither flagged effect is a single-player artifact within the season.

But the binding limitation is n=1 season, which no within-season robustness can fix. The bootstrap CIs resample players, not seasons, so they understate true uncertainty — there is zero out-of-sample replication and the difficulty-vs-surviving-field-strength confound is unidentified. The elite "+47%" rests on 5 players and a ratio-of-pooled-means that a couple of high-TUH playoff games can swing. Results (1) per-subject ESS drop and (3) ESS-specialist collapse are also the same underlying fact (the doc says so), not two independent confirmations. The doc grades this Low-Med and explicitly demotes it to a "tiebreaker, weakly held" — that is the correct altitude. Don't let it harden into a drafting lever.


Cross-cutting flags for the spine

  1. Soft archetype labels propagate. F28's labels feed F30's "bio underdrafted" / per-archetype board and F31's retention split. Both inherit the silhouette-0.23 / feature-sensitive softness. Per-archetype sub-tables (n=5–13 cells) are descriptive only.
  2. F30's validation should be removed from the evidence base for F8. F8's ~13 toss-pt/slot edge stands on its own design; F30 does not add an independent confirmation and its headline ρ should not be cited as one.
  3. Reproducibility gap: restore/commit the F31 script (A31_*.py). Artifact-only findings on n=1 season are hard to trust without the pipeline.
  4. What I'd trust to act on: F29's bounded null (no duo lever) and the direction of F28 (a broad-and-deep elite tier wins; specialists incl. CS are interchangeable coverage). Everything else is watchlist or weaker.

NSBA Draft Analytics · embargoed until after the SSB draft · ← hub