Methodology

Value Score (0–100)

Every player receives a Value Score from 0 to 100, representing their percentile rank within their position group (forwards, defensemen, or goalies) for the season. A score of 85 means the player's production was better than 85% of players at their position.

The score is built from a weighted composite of advanced metrics. For forwards, the inputs are a regularized on-ice expected-goals impact (RAPM where reliable, on-off otherwise; see the Expected Goals Model section), the player's individual expected goals per 60 minutes from our own model, points per 60 minutes, Corsi percentage, and on-ice expected goals percentage. For defensemen, the composite is production-anchored — total points and even-strength points per 60, overall points per 60, possession share (Corsi and on-ice expected-goals percentage), and individual expected goals per 60 — with the on-ice impact term (where a defenseman's career minutes make it reliable) entering only as a one-sided sanity check that can dock a defensive liability but never inflate a rating, reflecting that the blue-line market pays first for production. For goalies, the primary input is their career goals-saved-above-expected rate per shot, regressed toward league average in proportion to how many shots we have actually observed — a goalie who has faced tens of thousands of shots is trusted close to face value, while a small sample is pulled hard toward the middle. That quality signal is already shot-quality-adjusted (it accounts for how dangerous the shots faced were), and it is paired with a separate workload-and-durability term (how many games the goalie plays). We measure career performance this way — rather than a single-season rating — because in our season-over-season testing single-season goalie rate stats barely predict the next season: goaltending is close to a coin flip year to year, so instead of forecasting a number the data cannot support, we report how much of a goalie's career we can actually believe given the sample. Goalie value scores are therefore best read as lower-confidence than skaters'.

To reduce single-season noise, skaters blend the most recent three seasons with a 55/30/15 weighting (most recent heaviest). Goalies are handled differently: the quality signal pools a goalie's whole career (regressed by sample size, as described above) rather than a fixed three-season window — because goalie rate performance is so unstable year to year that a recency-heavy blend over-trusts a single noisy season — while the durability term is recency-weighted. Skaters with fewer than three seasons use available data, re-normalized.

A player's most recent season can itself be a tiny sample. A one- or two-game skater's per-60 rates are too noisy to support a confident talent estimate, so for skaters with fewer than 30 games we shrink the talent estimate toward league average in proportion to games played — a full-season player (30+ games) is unchanged, while a one-game player regresses most of the way to the middle. The shrink is one-sided: it only pulls down an inflated tiny-sample read, never lifts a genuinely weak one. This keeps a player who appeared in a handful of games from reading as a star on a few hot shifts, and pairs with the low-sample dollar treatment described below.

Sample coverage — and what it doesn't capture

Every verdict carries a sample-coverage badge — full-season, partial, or low — based purely on playing time (full season ≈ 50+ games for skaters, 40+ starts — or two seasons of 15+ — for goalies; low is under 30 games or 15 starts). We deliberately call it sample, not confidence: it tells you how much ice time the verdict is built on, and nothing else.

In particular, the badge does not measure how accurate the dollar figure is. Fair value is a single point read off a fitted curve, and that curve has its own error — about ±$2 million for a typical skater and somewhat more for goalies — that has nothing to do with how many games a player played. A player can have a full season of data (high sample coverage) and still sit on a stretch of the curve where the dollar estimate is shaky. So a persistent caveat appears on every verdict, at every coverage level, reminding you the dollar is a modeled estimate rather than a measurement.

Because of this, the dollar verdict is gated rather than always shown as a hard number:

  • Low sample — the dollar verdict is withheld entirely; there simply isn't enough ice time to put a price on the contract.
  • Within the model's margin — when the pay gap is smaller than the curve's own resolution (under about $750K), we say so instead of asserting a tiny over/underpaid figure the model can't reliably distinguish from zero.
  • Goalies — we never headline a hard dollar for goalie contracts, only the direction. Goalie value is the least repeatable year to year, so the goalie pricing curve is the noisiest one we fit.
  • Partial sample — the verdict is shown as a direction (leaning over/underpaid) with the dollar as supporting detail rather than the headline.
  • Full season, clear gap — the full signed dollar verdict is shown.

Fair Value ($)

Fair Value projects what a player's contract should cost on the open market, based on their production. It's expressed as a percentage of the salary cap (so comparisons across eras remain valid) and then converted to current dollars.

Crucially, Fair Value prices on-ice production at market — and nothing else. It puts no number on a no-movement or no-trade clause, on franchise, leadership, or marketing value, or on the premium teams pay for term itself (cost certainty, locking in a player before a rising cap). It does price term's on-ice consequence — an aging player's decline is projected to the midpoint of the contract's remaining term (see Age Curves) — but not term as a contractual asset. Those are real things teams pay for; we simply don't claim to measure them. So a contract that reads slightly overpaid against on-ice production can still be a sound allocation once that excluded value is counted — the gap is a deliberate scope limit, not by itself a bad signing.

For skaters the open-market rate is built from the value composite — the same blended, three-season measure behind the Value Score, but taken before it is flattened into a 0–100 percentile rank. That distinction matters at the very top: dozens of skaters share a Value Score of 99–100, so a curve drawn against the rank cannot tell a generational talent apart from the 40th-best player. Pricing on the composite keeps the gaps between elite players intact.

Because a forward's composite is measured against other forwards and a defenseman's against other defensemen, the two are first placed on a common scale — a replacement-level player and a 99th-percentile player at each position are pinned to shared anchors. The market rate then rises exponentially with that normalized composite (each step up in value multiplies pay), fit on roughly 1,960 skater UFA contracts from 2020–21 through 2026–27. It is floored at the league minimum and capped at the CBA maximum contract — 20% of the cap, which no deal can exceed.

The curve is centered on the middle of the league rather than forced to make league-wide surplus net to exactly zero. That is deliberate: the NHL pays its stars well below their on-ice value (and its depth somewhat above), so a model honest about the middle will — correctly — show the league's best players as collectively underpaid. Forcing the books to balance exactly would hide that. Validated on a temporal holdout (trained through 2024–25, tested on 2025–26 and 2026–27), the curve predicts signing percentage to within about 2.4 percentage points.

Goalies are priced on a separate curve fit on goalie signings alone, still keyed to Value Score (pct_of_cap = 0.0828 − 0.2239·v + 0.2271·v², with v = Value Score ÷ 100). Because goalie performance is far less repeatable season to season, the market pays even elite goalies a smaller premium than skaters, and the goalie-specific curve fits actual goalie contracts better than the shared one (2.8 vs 3.1 percentage-point error).

Even-strength impact for skaters is measured with RAPM (regularized adjusted plus-minus): a ridge regression that solves for every player's per-60 expected-goals impact simultaneously, holding teammates, competition, and venue constant. The model also controls for where each shift began: an offensive- or defensive-zone faceoff start is entered as its own term, so a shutdown defenseman who opens shifts in his own end is no longer charged for the extra expected goals that deployment invites, and an offense-sheltered player no longer banks credit for starting in the attacking zone. This isolates a defenseman's own contribution from his linemates' — a sheltered, offense-first defenseman no longer borrows credit from strong teammates, and a shutdown defenseman who absorbs hard minutes is no longer penalized for them. Limitations: the model covers 5-on-5 play and home/road venue only; special-teams value and shot-blocking beyond its effect on 5-on-5 expected goals are not captured; two players who share nearly all their ice time (a fixed defense pair) cannot be fully separated; and to stay reliable the impact term is pooled across a player's career rather than estimated season by season, so it reflects a multi-year average rather than a single season's form. Coverage begins in 2010-11; earlier seasons retain the prior on-ice differential. The zone-start control applies from 2016-17 onward (the seasons with faceoff-location data); shifts in earlier seasons carry an "unknown" start and receive no zone adjustment.

Because that impact term is the anchor of a defenseman's talent estimate, we apply two symmetric, externally-corroborated adjustments — each triggered by real on-ice usage (single-season points and ice time), never by the model's own rating, so they cannot become circular. A defenseman whose estimate leans on the model's inflated credit but lacks the production or the minutes of a top-pairing player has that inflated credit discounted — both his isolated even-strength impact and the on-ice shot-share and expected-goals share he banks from riding a strong team's play — so a sheltered specialist does not read as elite. (Only credit above his position's average is pared back; a genuinely poor share or impact is left negative, never erased.) The production side of that trigger carries a minutes floor: raw points corroborate a defenseman only when he also logs top-four ice time, so a sheltered points-getter can no longer clear the gate on box score alone. Its mirror: a defenseman corroborated as a true #1 — heavy production and heavy minutes, both required — whom the impact model under-rates is nudged, by a bounded amount, up toward the #1-defenseman tier. This corrects a known-conservative read for the productive, heavy-minutes #1 defenseman (the impact model is measurably the best available at ranking the elite tier, so we keep it and correct only this named cohort rather than re-weighting the whole model, which testing showed ranks the elite worse). The nudge is bounded and never reaches the elite apex, and it prices a corroborated #1's current value — it does not chase the market's youth-premium ceiling, so a young, corroborated #1 on a long, rich deal can still read as a mild overpay against his present production. That is the honest read.

A third adjustment — one-sided rather than symmetric, but on the same non-circular footing — addresses a related gap at the young end. Because the impact estimate is pooled over a player's whole career, a defenseman with few career minutes has that estimate regressed hard toward the league average — so a genuine young top-pairing defenseman who already plays heavy minutes and produces can be read as near-replacement, purely because he has not yet banked the ice time. When a defenseman is young, plays top-pairing minutes, and produces at a top-pairing rate — again all observable, never the model's own rating — his value is floored at a defensible young-top-pairing minimum, set deliberately below the proven-#1 floor above: a defenseman who has already proven it still prices higher. This is a floor, not a forecast — it does not credit expected development (a young player's out-year upside is handled separately, and conservatively, by the projection range), and it does not lift young defensemen as a class: one who lacks the minutes or the production is left exactly where the model puts him.

The projected signing prices on the Free-Agent Class page are built on the same value foundation as the rest of the site. A player's talent estimate is priced through the value-to-dollars curve and then declined by the empirical aging curve over a representative contract term — so forwards, defensemen, and goalies share one path, and a player's free-agent projection rests on the same model as their leaderboard fair value. Defensemen carry one added correction: past their mid-thirties the price also absorbs the market's own veteran-defenseman age-discount, grafted from real veteran-D signings, so an extreme-old defenseman prices against actual older-D comparables rather than flooring too shallow. The decline for young and peak defensemen is unchanged; forwards and goalies are unaffected. Elite forwards reaching the open market, where comparable signings are scarce, are shown with a deliberately wide range labeled as such rather than a single falsely precise figure.

Retained salary. For players with retained salary — where a trading team continues to carry a portion of the contract — verdicts compare fair value against the post-retention cap hit: the amount the holding team actually pays. The face contract value is preserved as context but is not the verdict denominator.

Two scoping tests — class-depth adjustment and NMC/NTC premiums — are documented in What We Tested and Did Not Adopt.

Availability Discount

A player who reliably misses a chunk of the season delivers less on-ice value than their per-game rate implies, so their fair contract dollars are haircut. The availability discount measures how much of the schedule a skater has actually been available for recently — their games-played rate (games played ÷ games their team scheduled), blended over the current season and the two prior seasons with the same 55 / 30 / 15 weighting the Value Score uses. Short seasons are handled honestly: the 48-game 2012–13, 70-game 2019–20, and 56-game 2020–21 schedules are measured against their own length, not against 82.

The blended rate is shrunk toward the league mean — about 90% of games — with a 40-game pseudo-count, so a thin sample is pulled toward average and only a sustained pattern of missed games moves the number. The result is a multiplier applied to both the open-market fair value and the status-adjusted fair value, after the age curve. It adjusts the dollar figure only: it does not touch the Value Score, the talent estimate, or a player's ranking among his peers.

The discount is downward-only. A skater at or above league-mean availability (roughly 74 of 82 games a season) receives no adjustment at all — there is no iron-man bonus. Below that mean, the haircut grows linearly at a half pass-through: a skater who plays about 75% of the schedule — roughly 15 percentage points below the 90% league mean, or about 62 of 82 games — is discounted about 7.5% in fair value. The maximum haircut is capped at a floor of −20%, no matter how many games are missed. Goalies are out of scope entirely (their starter/backup workload split makes games played a poor availability proxy) and always price at no discount.

Two caveats matter for reading this line honestly:

  • It measures availability, not injury. Games played does not distinguish a genuine injury from a healthy scratch. This first version deliberately measures only whether a player was available, and makes no claim about why he was not.
  • Breaking-in seasons are excluded. A prior season before a player was an NHL regular — a call-up or cup-of-coffee year (fewer than 41 games) — is dropped from the blend rather than counted as an "unavailable" season. A rookie who played a handful of NHL games two years ago wasn't injured; he wasn't yet an NHL regular, so those seasons carry no penalty.

The raw discounts are pool-preserving: after computing each skater's individual multiplier, the discounted cohort's multipliers are scaled back up toward 1.0 by a single global factor, so the league-wide sum of fair values is held essentially unchanged (within 1%). This makes the discount relative — a below-average-availability skater is priced lower than his healthier peers, while the total pool stays fixed — and it caps the realized maximum below the 20% formula limit (the deeper the discounts the more the rescaling compresses them). The published availability_mult is the post-rescaling value, so applying the formula directly to a player's GP rate will give the pre-rescaling number, which will differ slightly.

Parameters: pass-through k = 0.5, floor 0.80 (−20% cap), recenter 0.90 (league-mean availability), shrinkage pseudo-count C = 40 games.

Sensitivity Range

Every fair value is a single point estimate, but the real fair price is uncertain. Puckonomics shows a sensitivity range around it — a deliberately labeled band, not a statistical confidence or prediction interval. It answers a practical question: how far could the fair price reasonably sit, given the things that actually move the number?

The range is driven by the two real sources of error. The dominant one is the gap between the model's price and what comparable players actually signed for: across roughly 1,960 UFA skater signings, the model lands within about 2.4 percentage points of the cap two-thirds of the time (about ±$2.3M at a $6.7M player on the 2025–26 cap). That measured pricing error sets the band's base width. A second, smaller term re-prices the player one extra year into their age decline; it moves the number only slightly and is included for transparency.

The band is then widened for low-confidence players — those with thin samples (under 30 games for skaters, 15 starts for goalies). This is the opposite of what a naive "vary the player's own three-year history" approach would do: that method collapses the range to nearly zero for young players, exactly where uncertainty is highest. Widening for thin samples keeps the band honest. The widening factors are a stated assumption rather than a measured quantity, because signed contracts are almost all full-season veterans and cannot pin the thin-sample case down empirically. Goalies get a wider base band (about 2.9 points) because goalie value repeats far less reliably year to year.

The band is built in open-market percentage-of-cap terms, then discounted and floored exactly like the point estimate, so it always brackets the status-adjusted fair value shown. A player at the league-minimum floor gets a one-sided (truncated) range — you cannot be paid below the minimum. One caveat: the band is calibrated on the UFA open market; for RFA and ELC players it inherits that error scaled by the status discount, which is a reasonable default but not independently validated. The market also signs, on average, a hair above the model, but that is treated as a survivorship artifact of which deals get done, so the band is centered on the model rather than shifted.

That band is not just asserted — you can see it. Below, every current-season UFA skater’s open-market pricing residual, split by talent level. An unbiased model sits on zero at every talent level, inside the band, with no tilt:

Fair value − cap hit (% of cap) →
lower talent
-0.4pp
mid talent
-1.1pp
higher talent
-0.7pp
valued above cap hit below cap hit ±2.5pp typical pricing scatter low sample
For every current-season UFA skater (571 of them), this shows how far their actual cap hit lands from our open-market fair value, as a share of the cap. An unbiased model centers on zero. It does — mean -0.7pp, inside the ±2.5pp typical pricing scatter (the sensitivity band, not a confidence interval), with no drift across talent levels. Restricted (entry-level and RFA) players are excluded: their below-market pay is CBA cost control, not a model residual. One note: the model caps top-end value near the CBA maximum, so a handful of genuine superstars can read as “below market” — that is the cap compressing star salaries, not a pricing error.

Projection Range

Player multi-year projections carry an empirically-calibrated uncertainty range. We measured how far past projections landed from reality on a leak-free, season-by-season backtest, and calibrated the range so that about 80% and 90% of actual outcomes fall within the 80% and 90% ranges. It is asymmetric and wider for high-value players (a convex pay scale makes a star's dollar value far less certain than a depth player's), widens with contract length (fit from 1–4-year-ahead backtests, extrapolated beyond), and skews downward in out-years to acknowledge players who leave the league (whom survivor-only data can't see). It is a calibrated range, not a probabilistic guarantee, and is re-checked as new seasons land.

Separately, we tested whether a regression-to-the-mean forecast would price contracts better than this projection. It did not: about 5% worse on dollar accuracy from the smoothing itself (pricing held identical), and about 8% worse than the projection we actually ship, losing all four backtest seasons against both — it improved the underlying rate forecasts but lost on the convex pay scale. We deliberately did not tune the model to the dollar metric itself, which on four seasons would overfit; the failure held across every adjustment (exposure basis, smoothing target, an explicit curvature correction), so it is structural. We retained the current projection — and that negative result is what calibrated the ranges above.

Comparable-Deal Base Rate

Alongside the model's dollar verdict, a qualifying contract can carry a comparable-deal base rate: among past deals structured like this one — the same position, age band, value tier, cap share, and signing status — how often did the player go on to deliver surplus over the following three seasons? It is a population frequency, published separately from — and answering a different question than — the model's judgment of the specific contract. We show it only where it clears three pre-registered reliability gates, validated on 3,473 comparable signings from 2015–2025. Calibration is the whole story; we make no discrimination or “beats-a-baseline” claim here.

Gate 1 — Calibration: does the rate mean what it says?

When the base rate says 30%, do about 30% of those deals actually pay off? We sort every historical comparable into ten equal groups by predicted rate and check the realized rate against the prediction — walk-forward, so each season is scored only from earlier ones. All ten deciles land within 10 percentage points of their prediction.

Predicted vs. realized surplus rate, by decile
DecilePredictedRealizedErrorDeals
19%10%0.8 pp268
218%21%2.6 pp268
323%21%1.8 pp268
430%37%7.2 pp269
537%42%4.3 pp268
644%42%2.1 pp268
754%55%0.4 pp269
863%59%3.9 pp268
972%71%1.1 pp268
10 †82%78%3.9 pp269

† The top decile is the one to watch. The richest, highest-predicted deals realized surplus 78.4% of the time against a predicted 82.3% — about 3.9 points optimistic. That gap is inside our tolerance, so the gate passes, but we surface it rather than hide it: the base rate slightly over-promises at the very top of the market.

Gate 2 — Resolution: does it separate good structures from bad?

A base rate is only useful if it spreads deals apart instead of quoting the same league average to everyone. Across the comparable pool, cohort realized rates span 43.9 percentage points from the lowest group to the highest — against an unconditional league rate of 40.9% — with a Brier skill score of 0.211 (a resolution measure; higher means the cohort rates track outcomes more sharply). The cohorts carry real, differentiated signal.

Gate 3 — External anchors: does it agree with something we don't control?

The first two gates check the rate against our own later valuations — a self-consistent yardstick, not an independent one. Gate 3 asks whether the base rate lines up with something the model does not compute: the market's own re-signing behavior. It does, modestly. Cohorts with a higher base rate re-sign at a higher rate — a rank correlation of 0.284 across 1,999 pairs, with 82.3% of high-base-rate players earning a raise versus 58.0% of low. That is the real external anchor, and it is a moderate one.

We are deliberately clear about what is weaker. Raw future production correlates only faintly with the base rate for forwards (+0.078). For defensemen the aggregate signal is status-confounded: it is carried almost entirely by unrestricted free agents (UFA-D rank correlation 0.35), while the cost-controlled entry-level and restricted buckets sit at essentially zero (-0.06 and -0.03). Read the base rate as a structure-and-market signal, not a clean forecast of on-ice output.

What we did not find: our “underpaid” call adds no lift

We also tested whether our own underpaid verdict adds predictive lift beyond the comparable-deal base rate — whether flagging a deal underpaid meaningfully raised the odds it delivered, within a cohort. We found it indistinguishable from zero. That is why we publish a base rate, not an “our call doubled the odds” claim: the honest quantity is the frequency of the comparable group, full stop.

Limits that travel with every base rate

  • Self-consistent, not an oracle. “Delivered surplus” is measured on our own valuation model — the same one behind the rest of the site. It is internally consistent but is not an independent judge of value.
  • It mostly reflects contract structure. Cheap and entry-level deals return surplus far more often than expensive veteran deals, so a high base rate is often a statement about the deal's price, not a forecast of stardom.
  • The cheapest tier discriminates least. Sub-2%-of-cap deals dominate the pool and almost all clear, so within that band the rate spreads outcomes by only about ten points. There, read a base rate as a coarse “these usually work” rather than a fine-grained edge.
  • The current-season level is extrapolated. The validation runs on completed signings (2015–2025), and surplus rates have trended upward across that span. The rate attached to a current-season contract is a projected continuation of that era trend — the newest class has not had three seasons to play out — not a rate observed on a finished cohort. The ordering of cohorts is stable; only the absolute level for the newest season carries this extra extrapolation.

Trade Calculator

The Trade Calculator weighs the two sides of a proposed trade on cap-efficiency surplus — the value a player delivers over the remaining term of his contract, minus what he is paid (see Fair Value). It is deliberately a surplus tool, not a talent tool: a star on a fair-market contract carries little surplus, so a star-for-picks deal can read “too close to call” even when one side clearly receives the better player. Each side's surplus is summed, and one side is named the winner only when the gap exceeds the combined uncertainty band — otherwise the verdict is “too close to call.”

Draft picks are valued from a curve of historical draftee surplus by draft slot, built from every pick's realized cost-controlled-window value with busts included, so the curve is an honest expected value rather than a survivorship-inflated one. Because no public source reliably projects a future draft order, first-round picks are priced at a coarse, user-stated tier — lottery, mid, or contender — each a smooth factor off the round-one average rather than a single noisy slot; later rounds use the round average. Every projected pick carries the full round-one uncertainty band and is flagged low confidence, and future-year picks are discounted for the time until their value is realized.

Each asset's surplus carries a band derived from the model's measured pricing error, scaled over the contract term (see Sensitivity Range). Per-year valuation errors are partly persistent within a player — a misestimate tends to survive a re-pricing — so the term scaling sits between fully independent (square root of term) and fully persistent (linear), using an empirically estimated lag-1 residual autocorrelation of about 0.37, measured only across genuine re-signings. Asset bands combine in quadrature, and the “too close to call” threshold is one combined band: the tool will not crown a winner inside the noise.

The Trade Calculator checks only the value of the assets. It does not check roster fit, cap space, or cap compliance — a trade it calls a clear win may be impossible to execute under a team's cap. And because it scores surplus rather than talent, it under-credits a great player on a fair contract; a talent-in-context line is planned but deferred until it can be shown in stable units (see Limitations).

Contract Status Adjustment

Not all contracts are negotiated on equal footing. Entry-level contracts (ELCs) are capped by the CBA, and restricted free agents (RFAs) negotiate with less leverage than unrestricted free agents (UFAs).

Puckonomics reports two fair values: an open-market value (what the player would earn as a UFA) and a status-adjusted value that accounts for the leverage dynamics of their actual contract situation. The discounts are measured, not assumed: at an equal Value Score, entry-level players sign for roughly 60% less than the UFA curve predicts, and restricted free agents receive a smaller, value-dependent discount — largest (around 14%) for lower-value RFAs and tapering toward zero for those whose open-market value already approaches unrestricted dollars. The Pay Delta on the leaderboard uses the status-adjusted value, so an ELC star correctly shows massive surplus without implying their team made an error.

A player's status is his class this season, not when his deal ends. We derive it from the CBA rules using data we already have: age at the June 30 cutoff, accrued NHL seasons, and career games. Entry-level deals are identified by contract type; everyone else is classified as UFA (27+, or seven accrued seasons, or a Group-6 career fringe player) or otherwise RFA. This matters because a multi-year contract records the status a player will reach at expiry. A 22-year-old on the first year of a long second contract becomes a UFA when it ends, but is a restricted free agent right now. Pricing that in-season, rather than off the at-expiry label, is what keeps a young player on a big second deal from being valued as though he had full open-market leverage he does not yet have.

The contract track record on a player page measures realized surplus differently from the forward Pay Delta above. Because it asks what a deal has actually delivered, each elapsed season is valued at the player's open-market value minus the cap hit they were really paid — with no entry-level cap or restricted-status discount applied. That way an elite player's ELC and RFA years show the full surplus their team banked, rather than the smaller, signability-discounted figure used to price a contract going forward.

Age Curves

Player production follows predictable aging patterns by position. Forwards typically peak between ages 22 and 25, defensemen between 23 and 28, and goalies between 25 and 31. These curves are used to contextualize a player's current production within their expected career arc. When a contract's remaining term is known, the decline is projected to the midpoint of that term rather than the current season alone, so a long deal for an aging player is discounted more.

Expected Goals (xG) Model

From the 2010–11 season onward, expected goals are computed by our own model rather than imported. It is a gradient-boosted decision-tree model (LightGBM) trained on unblocked shot attempts (shots on goal, missed shots, and goals; the "Fenwick" basis). Blocked shots are excluded because the league records where the block happened, not where the shot came from. Shootout attempts and penalty shots are excluded entirely.

Each shot's chance of becoming a goal is estimated from 99 features. The core describe the shot itself — distance to the net and its square, angle (including shots from behind the goal line), a distance–angle interaction, shot type, man strength, whether it was a rebound or came off a rush, and the rebound geometry (how far and how fast the puck moved across before the shot). Layered on top are contextual enrichments: shooter and goalie finishing/vulnerability priors, on-ice possession embeddings, sequence-pressure scores, and — for 2021–22 onward — NHL EDGE tracking data (shot speed, skating speed, zone time). Empty-net attempts flow through the model like any other shot rather than receiving a fixed value.

The model is retrained on all seasons since 2010–11 with the most recent season (2025-26) held out to report honest calibration. Calibration is audited per season, and each season's league-total expected goals is scaled to match actual goals scored — a single model spanning a rising-scoring era lands on the average environment and drifts year to year, so each season is normalized to ground truth. This is a uniform per-season adjustment, so it does not change any player's relative ranking or value.

We report calibration, not discrimination: whether a predicted 8% chance really becomes a goal about 8% of the time. That is the property that makes an expected-goals number trustworthy as a rate, and it is the honest story here — modeling sophistication is not the positioning.

Held-out calibration (2025-26)
Calibration error (ECE)
0.7%
Brier score
0.0600
Observed goal rate →
Predicted goal rate →
≥200 shots <200 shots (wide sampling error) perfect calibration
Each point is a bucket of shots; its size is how many shots landed there. The vast majority sit at low predicted values, right on the line, so calibration error is just 0.7% across the ten buckets. The few buckets above ~55% predicted are the rarest, most dangerous chances (only a few dozen to a few hundred shots each across the held-out season), so their observed rates carry wide sampling error (shown as whiskers) and wander around the line rather than proving a real bias.

The stat labeled xG IMPACT/60 on player pages is an on-ice expected-goals impact rate, in expected goals per 60 minutes — not wins. For skaters whose career 5v5 ice time clears our reliability floor, it is the player's career xG-RAPM: a regularized estimate of on-ice expected-goals impact that adjusts for teammates and competition. For players below that floor (and for seasons before 2010-11, where we don't compute it), it falls back to a raw on-off expected-goals differential per 60 — the team's xG rate with the player on the ice minus the rate with him off it — which makes no teammate or competition adjustment. This descriptive on-ice rate is shown for transparency; the value score uses the same RAPM-based term, blended and standardized as described above.

Data Sources

Traditional statistics, rosters, play-by-play, and shift data come from the NHL's public API. Expected goals, the on-off differential, and goals saved above expected are computed by our own model from 2010–11 onward (earlier seasons retain values imported from MoneyPuck's publicly available datasets, which also serve as an ongoing validation reference). Contract information is maintained from public records.

Historical seasons sit on differing published bases. Recent seasons are priced on the current model using 5-on-5 possession computed from our own play-by-play — every season from 2011-12 onward. The oldest seasons (2008-09 through 2010-11) remain frozen on their prior published values from an earlier version of the model, pending deeper coverage. We recompute and republish older seasons on the current basis one at a time, only once each clears our validation checks — so a value from a much earlier season may reflect a slightly older methodology than the latest one.

Data is refreshed by running the pipeline manually; there is no fixed weekly schedule. The "Data as of" timestamp on every page shows when the most recent pipeline run completed.

What We Tested and Did Not Adopt

These are adjustments we built and validated against historical data before deciding not to ship. In each case, we share the result rather than silently omitting the test.

Role Scarcity

We tested whether positional scarcity in a UFA class (few quality free agents at F, D-L, D-R, or G) predicts signings clearing above model price, 2010–11 through 2025–26. After removing the shared time trend and the 2020–21 flat-cap shock, no within-role scarcity effect survives, so fair value includes no class-depth adjustment.

Contract Clauses (NMC/NTC)

We tested whether a no-move/no-trade clause carries a signing premium after controlling for player quality, age, and status. The apparent premium was disqualified by sparse, non-random clause coverage in the data, so no clause adjustment is applied.

Deployment / Zone Starts

We tested whether offensive-zone deployment inflates a player's on-ice share metrics enough to warrant a discount. After separating coaching selection (good offensive players get O-zone starts), team-system effects, and special-teams contamination from the true zone-start effect, the defensible causal component sits below our run-to-run noise floor, so no deployment discount is applied.

Defenseman Handedness

We tested whether the market pays right-handed defensemen above model price at the same quality — a static handedness premium — across 1,744 defense signings, 2008–09 through 2025–26. After controlling for signing year, player quality, age, term, and contract status, the right-shot effect is −0.07% of cap (95% CI −0.19 to +0.05), below our pre-registered +0.20% floor; the small raw gap in recent seasons reflects player quality and term mix, so fair value includes no handedness adjustment.

Limitations

Puckonomics is a free analytical tool, not a definitive verdict. The model intentionally uses a simple, transparent approach rather than a black-box prediction. Known limitations include:

  • Even-strength impact (RAPM) covers 5-on-5 and venue only and is pooled across a player's career; special teams and single-season form are not captured by it.
  • Goalie valuation is inherently volatile due to high season-to-season variance in performance.
  • The model does not account for intangibles like leadership, playoff performance, marketability, no-movement / no-trade protection, or franchise value — Fair Value prices on-ice production only (see Fair Value).
  • Contract comparables are influenced by market conditions (cap trajectory, free agent class depth) that the model doesn't capture.
  • Historical coverage begins in 2008–09; our own xG model covers 2010–11 onward (the NHL play-by-play and shift data it needs don't exist before that), with earlier seasons retaining imported values.
  • The xG model has no pre-shot-movement features — cross-ice passes and one-timers are systematically underrated relative to models that track them.
  • The xG model has no shooter-talent term and no adjustment for arena-specific shot-location recording bias.
  • 3-on-3 overtime is treated as even strength, and the rebound window is a fixed 3 seconds.

The generational ceiling

The dollar model is fit on players who actually reached the open market. The two or three best players alive essentially never do — they re-sign — so we have only a handful of comparable signings for that tier, all against lower historical caps. Rather than extrapolate an “elite premium” from a sample that small (which our out-of-sample tests showed adds error, not accuracy), the model prices these players on measurable on-ice value and we show the real market comparable alongside it on the free-agent view. The difference is the generational, term, and marketing premium we don't claim to measure — a data limitation, not a flaw in the player's rating. We investigated closing it two ways (re-weighting the value composite; adding convexity to the top of the dollar curve); both failed leakage-free out-of-sample testing, so we deliberately leave it unmodeled and labeled.