Methodology
Status: research preview under GS-18 repair. Conventional game, team, and player facts may publish only through current source, validation, rights, and lineage gates. No player rating has earned a public leaderboard. A definition below is a promise of construction, not a claim that the metric is valid or deployed.
▸ CORE — Contextual On-Field Rating Estimate
A retrospective, context-adjusted estimate of a player's contribution to scoring value per standardized opportunity, expressed in expected points above average per 100 qualifying snaps. CORE is model-conditional: it estimates contribution after accounting for teammates, opponents, situation, coaching, package, and role — it is not a claim about metaphysical football ability, career value, or the future.
CORE combines three evidence channels: context-adjusted outcome evidence (collinear but broad), observable action credit (narrow but direct), and role-specific statistical priors (regularizing, never trained on proprietary all-in-one metrics). Team, unit, opponent, coach, and package effects are modeled explicitly so a player coefficient cannot silently absorb system quality.
Current posture: scaffold. The QB ridge candidate remains a research workbench only. Its point estimates modestly exceeded raw EPA/dropback, but clustered candidate-minus-baseline intervals span zero on every audited task. The raw conventional baseline is retained; no superiority, rank, or public player value is claimed.
▸ Expected points and win probability
Game pages retain nflverse source EP/WP as attributed context. Historical challenger
packets are preserved, but the independent audit found insufficient same-bin calibration,
convergence, and portable fitted-state evidence. Every challenger is REWORK
until a newly frozen tournament passes those gates on genuinely unspent future data.
▸ FORGE — Football Overall Replacement-adjusted Game Equity
Cumulative value above a role-specific replacement baseline, in expected points first and wins only after an empirically calibrated points-to-wins conversion with published error bars. Replacement cohorts are defined result-blind from labor-market evidence (roster-fringe, practice-squad promotions, backup stratification), and every team-season publishes a reconciliation ledger — assigned player value, coaching/scheme, interactions, and a visible unassigned residual. A visible residual is more honest than false precision.
Current posture: scaffold, no player values. V1 mixed QB rushes into the RB lane and RB targets into the WR/TE lane, so those cohort conclusions are superseded. V2 requires the event to join a season-and-team roster position, quarantines unmatched or wrong-position rows, and recomputes denominators only after eligibility filtering.
▸ PULSE — Performance Under Latest Sample Evidence
A rolling recent-form descriptor with its exact evidence window always visible, shown beside full-season CORE, never in place of it. Window and half-life parameters are chosen on historical out-of-time tasks, not on making a preferred player look hot. PULSE is not automatically a forecast, and it never generates causal narratives from a trend.
Current posture: scaffold. Player pages may show explicit trailing 4/8-game factual windows; those are not PULSE values. Window selection, opponent adjustment, uncertainty, and future validation remain undone.
▸ SIGNAL — the evidence profile
Six axes: Sample (enough relevant opportunity?), Identifiability (can this player be separated from stable teammates?), Granularity (how directly does data observe the claimed role?), Noise (how unstable is the estimate?), Availability (is the period complete?), and Lineage (is every input reproducible?).
Each axis is either measured on 0–100 with its evidence attached or explicitly unmeasured with a reason. Missing is not zero, and omission is never favorable: if any required axis is unmeasured, aggregate completeness is 0 while the axis itself remains null. Population reliability is context, not a player's noise score. Lineage earns 100 only from complete verified transitive captures. SIGNAL is a scaffold evidence profile, never a player rating or a percentage chance that a rank is “correct.”
▸ Evidence tiers by position
| Tier | Typical roles | Visibility |
|---|---|---|
| A — directly observed | QB, kicker, punter, returner | Strong event and opportunity evidence on nearly every snap |
| B — partially observed | RB, WR, TE, edge, targeted DB | Participation plus direct events; non-target/non-pressure snaps less visible |
| C — presence-dominant | OL, interior DL, off-ball LB, non-target safety | Presence and unit outcomes; assignment unobserved — wider intervals required |
| D — assignment-blocked | Roles without reliable participation evidence | Counting facts and team context only; estimates render unavailable, never guessed |
▸ The trench program
Evidence tiers gate publication, not research. The current offensive-line posture is WITHHOLD INDIVIDUAL; the defensive-front posture is UNIT / IDENTIFIABILITY ONLY. A derivative family cannot count as independent agreement, team-confounded estimates do not establish player attribution, and an offensive- line simulation cannot substitute for a front-specific recovery test. Missing assignment evidence remains missing. Unit aggregates, visibility diagnostics, and honest failure evidence may continue without becoming player ranks.
▸ Special teams
LEG kicking points above expected is the first specialist scaffold. It models pre-kick make probability from lawful play-by-play, reports listed-kicker operation value with uncertainty and sensitivity checks, and never claims isolated kicker talent. Historical testing cannot promote it: a genuinely unopened future regular-season gate is deliberately required and currently false. Punting, return, and coverage constructs remain concept-stage work.
▸ Holdout integrity
The executable registry marks 2025 SPENT because it was captured, canonicalized, and published before the original seal could protect it. It may be used as disclosed historical evidence where policy permits, but never again as fresh selection, tuning, gate, validation, or promotion evidence. Research entrypoints check this policy before loading model data.
▸ The claim ladder
Every metric occupies exactly one status:
concept (named question, no values) →
scaffold (contracts, fixtures, and gates exist) →
research (real data, research lab only) →
provisional (public with required disclosures) →
validated (passed construct and out-of-time gates) —
or retired / blocked, both published with reasons.
No metric goes live because the site needs another card.