30 — Undervaluation-vs-Field Metric (value over the market)
30 — Undervaluation-vs-Field Metric (value over the market)
Date: 2026-05-30. Status: Task C. Script: scripts/A30_undervaluation.py
(reproducible). Outputs: outputs/undervaluation_board.csv (241 rows: 189
historical picks + 52 nsba4 combine pool). Env: /home/david/code/nsba/.venv/bin/python.
LEAD CAVEATS (read first). 1. Tiny, leaky sample. Validation rests on 94 drafted players with both a combine and a realized same-season game value, spread over 3 draft seasons (nsba1 contributes ~6 usable, nsba2 46, nsba3 42). Realized value is same-season PPTF — the players were drafted before the games we score them on, but our value model partly reuses those games (the survivorship / same-season leakage that bounds F8). Read magnitudes as directional. 2. This is a re-expression of F8, not a new effect. The metric formalizes the established result that the field anchors on raw combine (|ρ|~0.7) while combine predicts value only ~0.4. The undervaluation score is a systematic, per-player version of that gap. It does not discover a new edge; it operationalizes the known one. 3. No individual name is statistically significant. BH-FDR(0.10) on the 94 per-pick residuals flags 0 individually significant steals/busts. The named boards are shortlisting flags, not significance claims (extends F8 caveat 7, F12). 4. The nsba4 forward board is 100% combine — as gameable and noisy as the combine itself (F3 selection-bias caveat). It predicts who the field will misrank, which is well-grounded, but not that our projection is correct in absolute terms.
The metric
For every player-season we put our value and the field's implied value on a common within-season percentile scale (0–1, higher = better), then subtract:
value_over_field (uvf) = our_value_percentile − field_value_percentile
Positive = we rank the player above where the market did → undervalued.
Two field proxies (kept separate on purpose):
| field proxy | meaning | coverage |
|---|---|---|
uvf_adp — field = overall_pick |
the realized market price (draft slot) | drafted players |
uvf_combine — field = raw_overall rank |
the combine anchor the field mechanically follows (F8) | anyone with a combine, incl. all 52 nsba4 |
Our value = realized game PPTF percentile (F2 north-star) historically;
de-biked combine theta percentile for the combine-only nsba4 pool.
Percentile (rank) scale is deliberate: it is unit-free, robust to the nsba3 estimated-TUH proxy and to combine scale drift (F7), and it matches how a draft actually works (relative ordering, not absolute points).
Validation — do combine-undervalued players actually beat their draft slot? YES
The honest test: uvf_combine uses the combine as the field signal; val_resid
(F8's steal score) measures beating the ADP/pick-slot regression. These are two
different field proxies, so a correlation between them is genuine cross-validation,
not an identity.
- Pooled Spearman ρ(
uvf_combine,val_resid) = +0.599 (n=94, p<1e-9), bootstrap 95% CI [+0.42, +0.73] (season-clustered resample) — excludes 0. Holds every usable season: nsba2 ρ=+0.53 (p=2e-4), nsba3 ρ=+0.67 (p<1e-5). - Split test: combine-undervalued picks (
uvf_combine>0) beat their slot by +0.088 PPTF-residual vs −0.070 for overvalued picks (Mann-Whitney p<1e-4). - Robustness (win-shares as our value instead of PPTF): ρ = +0.45 (p<1e-4) — weaker (win-shares is schedule-confounded, redteam T6) but same sign and significant.
Takeaway: players the de-biked/value view ranks above their raw-combine slot systematically out-produce that slot. This is the F8 mechanism made per-player.
Biggest historical mispricings (flags, not significance)
Steals — combine-underrated, then produced (full list in board, realized_src='same'):
| season | pick | player | combine rank | realized PPTF | uvf_combine | archetype |
|---|---|---|---|---|---|---|
| nsba2 | 51 | Ray | 48 | 0.51 | +0.81 | replacement→produced |
| nsba1 | 33 | Vish | 56 | 0.74 | +0.80 | elite multi-science |
| nsba2 | 40 | Aneesh Swaminathan | 39 | 0.47 | +0.63 | bio |
| nsba3 | 54 | Ronuk Gadamsetty | 75 | 0.32 | +0.63 | bio |
| nsba2 | 60 | Mihir Kulkarni | 36 | 0.30 | +0.36 | (F12 buy-low confirm) |
Busts — field over-valued (early pick / high combine), under-produced:
| season | pick | player | combine rank | realized PPTF | uvf_combine |
|---|---|---|---|---|---|
| nsba2 | 36 | Rohan Dhillon | 11 | −0.14 | −0.77 |
| nsba3 | 42 | Mahith Gottipati | 23 | 0.00 | −0.74 |
| nsba2 | 12 | owen fei | 6 | 0.24 | −0.43 |
| nsba3 | 9 | Advai Srinivasan | 42 | −0.05 | −0.52 |
Signature is exactly F8's: steals are late picks the combine under-rated; busts are early/high-combine names that under-delivered (Rohan Dhillon #11→neg, owen fei #6). Mihir Kulkarni reappears as a steal, independently confirming F12's buy-low call.
Archetype / positional scarcity — is any type systematically under-drafted?
Mean uvf_combine (we rank above field) and mean val_resid (beat slot) by archetype
(F28 labels):
| archetype | n | mean uvf_combine | mean val_resid |
|---|---|---|---|
| bio specialist | 27 | +0.159 | +0.015 |
| ess specialist | 26 | +0.043 | −0.042 |
| math+phys specialist | 34 | +0.022 | +0.098 |
| elite multi-science (chem+phys) | 13 | +0.056 | +0.237 |
| low-output / replacement | 82 | −0.117 | −0.084 |
Two honest, partly opposing reads:
- Bio specialists are the most combine-under-drafted type (+0.159) — the field's
raw combine systematically ranks them below their game value. BUT their mean
val_residis only +0.015: they beat combine rank but barely beat their actual draft slot, because the field's live scouting already partly corrects for it. So "bio is underdrafted" is a weak, combine-only effect, not a reliable slot-beating edge. Do not over-weight. - The elite multi-science tier delivers the real residual value (+0.237 val_resid, the F28 "this is where wins come from" cluster) even though it is only mildly underpriced on combine (+0.056) — i.e. the market roughly prices the elites correctly on combine, and they still over-produce because combine understates how much the top tier matters. Conclusion: scarcity-adjust toward the broad-elite tier, not by reaching for a specialist archetype. Replacement-tier is correctly faded (−0.117 / −0.084).
This is consistent with the spine (F2/F28): draft for production rate / the broad-elite tier; single-subject specialists (incl. CS, which F28 shows has no standalone archetype) are interchangeable coverage, not a scarcity play.
FORWARD — NSBA4 combine-only board (who the field will misrank)
The field will draft nsba4 by raw combine rank (F8, |ρ|~0.7). We rank by de-biked
theta (the bias-corrected ability, F3/F10). uvf_combine = theta percentile −
raw-combine percentile. Positive = de-biking lifts them above where the raw board will
slot them. 6 meaningful buys (uvf>+0.10), 7 meaningful fades (uvf<−0.10).
Buys (de-biked ability > raw-combine slot):
| player | raw combine | theta | uvf_combine | archetype |
|---|---|---|---|---|
| Sid S | 15 | +0.20 | +0.30 | math+phys |
| Sanjay O | 16 | +0.22 | +0.21 | elite multi-science |
| Simon Z | 18 | +0.35 | +0.15 | math+phys |
| Chris W | 14 | +0.03 | +0.15 | math+phys |
| Nihar Bhave | 15 | +0.09 | +0.14 | bio |
| Kevin F | 11 | −0.06 | +0.14 | ess |
Fades (raw combine inflates them above de-biked ability):
| player | raw combine | theta | uvf_combine | archetype |
|---|---|---|---|---|
| Edward C | 18 | −0.04 | −0.23 | replacement |
| Roshan A | 21 | +0.09 | −0.22 | math+phys |
| Daniel Lu | 16 | −0.17 | −0.19 | replacement |
| Santhosh V | 15 | −0.36 | −0.16 | replacement |
| Rohan G | 22 | +0.17 | −0.14 | ess |
- Santhosh V lands on the fade list (combine inflated, de-biked low) — consistent
with F12's
santhosh_b.confirmed-biker exclusion. The metric independently flags him without being told. Treat as confirmation of the fade direction; he should be excluded from any draft entirely per F12, not merely "faded." - Sanjay O is the most actionable buy: the only elite-multi-science archetype among the top buys, and de-biking lifts him into the upper third.
- Several fades (Rohan G, Roshan A) are high raw-combine names round-1 talk will inflate — exactly where F8 says the field leaks value.
How to use it (and how NOT to): the nsba4 board is 100% combine-derived, so it predicts the field's misranking on solid ground (F8) but cannot confirm our absolute projection (F3 selection bias — combine-high no-shows/tankers break the link, and they are unobserved here). Use it to re-order within a tier, concentrated in the middle rounds (F8: round 1 is efficient), not to invent a tier. Cross every name against eligibility, the captain/keeper list, and the F10 availability/reliability prior before acting — a "buy" who never fields is worth zero.
Limitations / threats
- 94-row validation, 3 seasons, same-season leakage. The ρ=0.60 is inflated by survivorship (only drafted and ≥2-game players appear) — the true ex-ante edge is the F8 ~13 toss-points/slot figure, of which this is the per-player decomposition, not an independent confirmation of a larger number.
- Percentile metric compresses extremes. A player at the 0.95 vs 0.99 percentile looks similar; the board ranks order, not gaps. Pair with the PPTF/VORP point estimates (F13) for magnitude.
- Archetype labels are soft (F28 silhouette 0.23, bootstrap ARI 0.84) — the "bio underdrafted" and per-archetype numbers can reshuffle on a different sample. Reported as tendencies.
- No multiplicity-corrected individual significance (BH-FDR: 0 of 94 picks). Named players are flags.
- nsba4 fades/buys assume the field anchors on raw combine as in nsba1–3. If the field de-bikes this year, the edge shrinks. The metric is an exploit of a behavior, not a law.
Reproduce
/home/david/code/nsba/.venv/bin/python scripts/A30_undervaluation.py
outputs/undervaluation_board.csv schema
pool ∈ {historical, nsba4_forward}; season, overall_pick, round, player_raw,
canonical_id, discord_tag; raw_overall, raw_rank (combine); realized_pptf,
realized_ws, realized_src; field proxies field_pct_adp, field_pct_combine; our value
our_pct_pptf (hist) / our_pct_theta (nsba4); uvf_adp (vs draft slot),
uvf_combine (vs combine anchor — the headline undervaluation score); val_resid,
steal_score (F8 slot residual); theta_overall, debiked_overall, theta_cs;
archetype_label. Sort by uvf_combine desc for buys, asc for fades.