F33 — Contestation-Adjusted Individual Ability (teammate-invariant)
F33 — Contestation-Adjusted Individual Ability (teammate-invariant)
Date: 2026-05-30. Task 2 of the player-value audit. David's concern (the
"usage-rate / empty-stats" confound): PPTF = toss_points / tossups_faced, and
tossups_faced (heard) is identical for every teammate. So a lone competent player on
a WEAK team vacuums all the buzzes the team wins → inflated individual PPTF, while the same
player on a STACKED team shares buzzes → deflated PPTF. If real and large, this would
poison PLAYER rankings and set the classic draft trap ("draft the empty-stats guy, he
regresses on a better team").
This finding builds a teammate- and opponent-adjusted ability that does NOT depend on who your teammates are, then asks the only question that matters for drafting: does correcting for the confound actually reorder the board?
Bottom line (lead with this): A conditional-logit contestation model that is provably teammate-invariant (synthetic recovery below) ranks players almost identically to raw pooled PPTF (Spearman 0.93), and the rank movement it produces is uncorrelated with team strength (corr(rank_shift, team win%) = +0.03, p=0.72). The empty-stats confound David feared is directionally present but far too small to reorder the draft board at the top (top-15 overlap 13/15, top-25 overlap 20/25). Use the adjusted board only as a within-tier tiebreaker for a handful of flagged names, not as a re-ranking lever. This is a bounded null on the confound, consistent with the F14 (additive teams) / F2 (PPTF is the currency) spine — not a refutation of it.
The model
Per won toss-up in a clean game (181 games, 3,165 won-toss-up events), one player scored +4. Model the buzz as a discrete choice over the contesting set S (players present across BOTH teams):
P(player i wins the toss-up | set S) = exp(theta_i) / sum_{j in S} exp(theta_j)
a Plackett-Luce / Bradley-Terry conditional logit. We fit one latent ability theta_i per
player by penalized maximum likelihood (L2 ridge λ=1 for identifiability — the scale is
only pinned up to a constant; then center). Why this is teammate-invariant: a
non-contesting teammate stays in the denominator, so a strong player on a weak team must
still be the max over their (weak) teammates and the opponents to win the buzz. Their
ability is not inflated by teammates declining to contest, and not deflated by strong
teammates — each toss-up is a head-to-head over the whole room.
Contesting set construction (the load-bearing approximation):
- nsba1 — full lineup (tossups_present > 0): true present roster, both teams.
- nsba2 — partial lineup where it exists; else buzzed-roster fallback.
- nsba3 — no lineup data; the contesting set is approximated by the roster that
buzzed in that game (flagged per event; 1,739 of 3,165 events use this approximation).
adjusted_ability.csv carries approx_frac per player = the share of their contesting
appearances that used the buzzed-roster approximation.
Result 1 — the adjustment barely moves the ranking
- Spearman(pooled raw PPTF, adjusted ability) = 0.929 (n=181 players with ≥1 clean game).
- Top-15 by raw PPTF vs top-15 by adjusted: 13/15 shared. Top-25: 20/25 shared.
- Median |rank shift| ≈ 11 spots out of 152, but concentrated in the muddy middle (raw PPTF 0.1–0.4) where players are nearly indistinguishable anyway; the top of the board is stable (where actual draft value lives).
The adjusted top-25 (draft board) is in outputs/contestation_draft_board.csv. The top
names — sanj, arolakiv, thedoge, VulcanForge, yufei — are the same elites raw
PPTF already identifies.
Result 2 — THE CONFOUND TEST (the decisive number)
If David's usage confound were real and large, the adjustment would systematically promote
players on strong teams (they were sharing buzzes, suppressed) and demote players on
weak teams (they were vacuuming, inflated). So rank_shift should rise with team win%.
corr(rank_shift, team_win_pct) = +0.029, p=0.72 — a clean null. corr(raw_pptf, team_win%) = +0.09 and corr(adjusted, team_win%) = +0.11 — the adjustment does not even reduce the (already weak) tie between rate and team strength.
Stratified (the trap's exact prediction): mean rank_shift on weak teams (win<0.4) = −0.4 spots; on strong teams (win>0.6) = +1.5 spots. The trap's direction is faintly visible (weak-team players drift down ~0 spots, strong-team players drift up ~1.5), but the magnitude is ~1–2 board positions — negligible against the sampling noise at n≈50 per stratum.
Interpretation: within this 3-season sample, being the "lone scorer on a weak team" does NOT meaningfully inflate a player's standing relative to a teammate-invariant ability. The empty-stats trap is real in theory; in this data it is too small to act on.
Result 3 — the model is provably teammate-invariant (synthetic recovery)
To prove the estimator isn't just laundering team strength, we simulated the exact contestation structure with known abilities (10 seeds, 60 players, 12 teams, Luce-generated winners): - Recovery Spearman(true, estimated) = 0.93 (range 0.88–0.96). - corr(estimation error, teammate strength) = +0.03, range [−0.16, +0.19] across seeds — centered on zero with both signs. The estimator does not systematically borrow a player's teammates' strength into their own estimate. This is the formal guarantee that the +0.03 confound null above is a property of the method, not luck.
Result 4 — robustness to the nsba3 approximation
Refitting on lineup-only events (drop all 1,739 buzzed-roster-approximated events, leaving 1,426 nsba1/partial-nsba2 events) yields abilities that rank-correlate with the pooled fit at Spearman 0.953 (n=114 shared players). The nsba3 buzzed-roster approximation does not distort the ranking — the lineup-rich seasons carry the same ordering.
Result 5 — adjusted ability still validates at the team level (honest cost)
Aggregating player ability to a TUH-weighted team-season mean: - corr(team mean adjusted ability, win%) = +0.28 - corr(team mean raw PPTF, win%) = +0.39
The adjusted ability is a somewhat weaker team-level predictor than raw PPTF. This is expected and not a flaw: raw PPTF mechanically folds in opportunity/volume that correlates with winning, while the contestation ability is a pure per-buzz win-propensity. The point of the adjustment is per-player fairness, not team prediction — for team prediction, F2's PPTF (+0.62 with full bonus/volume) remains canonical. We report the cost transparently rather than claim the adjustment improves everything.
The flagged players (use as a tiebreaker, NOT a re-rank)
"Empty-stats" candidates (high raw PPTF, drop most under adjustment → regression risk if you pay for the raw rate). Reliable pool, contest_n ≥ 50:
| player | raw PPTF | adj ability | raw→adj rank | team win% |
|---|---|---|---|---|
| Euna Kim | 0.37 | −0.09 | 41 → 89 | 0.20 |
| Aldric Benalan | 0.35 | −0.02 | 49 → 83 | 0.20 |
| Chris Wang | 0.32 | −0.08 | 54 → 87 | 0.50 |
| michael | 0.48 | +0.31 | 27 → 58 | 0.43 |
| dan.k.memes | 0.84 | +1.20 | 8 → 15 | 0.36 |
| JoshuaW | 0.93 | +1.38 | 5 → 12 | 0.45 |
Note even the biggest "empty-stats" drops (Euna Kim, Aldric Benalan) are mid-board
players on weak teams moving from ~rank 45 to ~rank 85 — i.e. from "fringe" to "deeper
fringe." None of them were top-25 raw names, so no first-3-round pick is implicated.
dan.k.memes/JoshuaW look like "drops" by z-delta but actually keep a top-15 adjusted
rank — they are elites, not empty-stats.
"Suppressed" candidates (low raw PPTF, rise most under adjustment → possibly hidden on a strong/contested team). Reliable pool:
| player | raw PPTF | adj ability | raw→adj rank | team win% |
|---|---|---|---|---|
| abcisosm5 | 0.29 | +0.79 | 64 → 27 | 0.73 |
| akul | 0.34 | +0.82 | 52 → 25 | 0.38 |
| jucijuce | 0.27 | +0.52 | 71 → 44 | 0.27 |
| shreyase39 | 0.18 | +0.25 | 97 → 62 | 0.50 |
| Advai S | −0.05 | −0.31 | 150 → 105 | 0.00 |
abcisosm5 (on a 0.73 team) is the cleanest "suppressed on a stacked team" case — exactly
the player the confound would hide. akul rises 27 spots but sits on a weak (0.38) team,
so its rise is NOT a team-sharing story — it's that akul won buzzes against tough
contesting sets. These two are legitimate within-tier upgrades; the rest are mid-board
shuffles.
Full movers: outputs/contestation_movers.csv. Per-season deliverable schema:
data/processed/adjusted_ability.csv.
Caveats (heavy — read before acting)
-
The confound is real in theory; the null is empirical and bounded. We did not prove the empty-stats trap is zero — we showed that in this 3-season, 181-player sample it is too small (corr +0.03) to reorder the top of the board. A confound smaller than ~2 board spots is invisible here. Do not generalize "PPTF has no usage problem" to other formats.
-
nsba3 contesting sets are approximated (buzzed roster, not true lineup) for 55% of events. R4 shows this doesn't distort the ranking (Spearman 0.95 vs lineup-only), but the approximation inflates each nsba3 contesting set with players who happened to buzz at least once and omits silent present players — it makes the contest look smaller and stronger than reality, biasing nsba3 abilities mildly upward in level (not in rank).
-
Pooled across seasons.
adjusted_abilityis one number per player over all clean games; a player's growth/decline between seasons is not modeled. The per-season rows inadjusted_ability.csvrepeat the pooled value. Small-sample seasons (e.g. a 23-TUH cameo) are why we compare against pooled raw PPTF, not noisy per-season PPTF. -
Identifiability / scale. Ability is identified only up to an additive constant (ridge pins it; we center). Ridge sensitivity: Spearman ≥ 0.98 for λ∈[0.25,2], 0.976 at λ=4 — ranking is robust to the penalty. Players with few contests (contest_n < 50) are excluded from the movers/board; their estimates are heavily shrunk toward 0 and unreliable.
-
Won-toss-ups only. We model who wins a buzz, conditional on someone on the floor winning it. Dead toss-ups (no team converts) and negs are not directly modeled (a neg is just "did not win this toss-up" — the player still appears in the contesting set and simply doesn't get the win, which correctly costs them). This is a win-propensity model, not a full neg/risk model.
-
Small n everywhere. 41 team-seasons, 181 players, 3 game seasons. Trust the direction (adjustment ≈ raw PPTF; confound is small) and the two clean flagged names (abcisosm5, akul); distrust any point estimate.
Action
- Keep raw pooled PPTF / VORP as the draft currency (F2 spine intact). The teammate-adjustment does not buy a new ranking — it vindicates PPTF against the empty-stats objection for this field.
- Apply the adjusted board only as a within-tier tiebreaker for the handful of flagged names: nudge abcisosm5 and akul up a tier (they win contested buzzes their raw rate undersells); apply mild discount to Euna Kim / Aldric Benalan-type fringe players whose raw rate came on weak teams — but note none of these are early-round picks.
- Do NOT build a team-strength correction into the player projection. The data says it would correct ~0–2 board spots and add noise. (Consistent with F14: teams are additive; F2: rate is the currency.)
Reproducibility: scripts/A33_contestation.py (committed). Artifacts:
data/processed/adjusted_ability.csv, outputs/contestation_movers.csv,
outputs/contestation_draft_board.csv.