Cap & Crease
Est. 2026 — Vol. I — Sep 11 Data Feed Active
Methodology
Why these systems exist and how they think. For definitions and keys, read the Glossary.
Fair Market Value — what the market actually pays
A player's trade value depends on the gap between their on-ice production and what it costs to replace them on the open market. Fair Market Value is what the market actually pays a skater or goalie with a given profile — not what we think they deserve, but what GMs have historically signed for.
The skater FMV model is trained on 1,996 one-way standard contracts signed between 2017 and 2026, fitted separately for forwards and defensemen. Production and deployment enter as monotone piecewise-linear splines with knots at the median and 85th percentile — a straight line was badly misspecified, under-predicting both the cheapest and the most expensive contracts while over-predicting the middle. The spline slopes are constrained non-negative so the price curve can only rise with production: an elite scorer is never penalized for scoring more.
Walk-forward validation (trained on pre-July 2024, scored on post): forwards R² = 0.70 with a mean error of $1.21M at a $104M cap ceiling; defensemen R² = 0.60 and $1.33M. Roughly two-thirds of what a skater signs for sits in these features; the rest is leverage, cap room, and how many clubs were bidding. The published range reflects that uncertainty rather than hiding it behind a single figure.
A separate backtest of the offensive power curve — the convex mapping from points pace to offensive value — tested the model's 1.6 exponent against 1,439 contracts matched to prior-season production. The optimal exponent across all skaters is 1.5 (R² = 0.5558) versus the model's 1.6 (R² = 0.5548), a gap of one-tenth of a percent. The market prices forward production slightly less convexly (optimal 1.4) and defenseman production nearly linearly (optimal 1.0–1.1), consistent with defensemen being valued more for deployment than scoring. The engine now uses the backtest-calibrated exponents: 1.6 for forwards (within 0.01% of optimal) and 1.1 for defensemen, eliminating the largest position-specific calibration gap in the original model.
Player Gravity — a modelled territorial field
Player Gravity v3 is a position-relative territorial influence index. It combines on-ice chance impact, transition proxies, and defensive suppression into an offensive-zone well, neutral-zone well, and defensive-zone dome. The rink field is a model visualization of those components, not a literal tracking map.
The current field force is a bounded display composite, not expected goals and not a direct measurement of defender attention. The public scope is MIXED SITUATIONS because v3 combines all-situations, 5v5, 5-on-4, 4-on-5, and regular-season EDGE aggregate inputs. Signal Stability is based mainly on agreement between current and baseline on-off values, with a legacy defenseman pair-driver adjustment; it is not a fitted portability model. Reliability is a 0–100 coverage/stability index, not a probability, and coverage is a hard ceiling on it. Profiles below 20 games or two-thirds weighted coverage are marked INSUFFICIENT and receive no tier or percentile.
Evidence, not just calibration. A year-over-year backtest (2015-2026, 5,722 consecutive-season pairs) shows gravity force persists at r=0.68 — a genuinely stable trait. It does not forecast a player's own future results (it predicts his next-season on-ice xGF% worse than simply carrying his prior forward), which is why it is not a scoring number and stays out of valuation. But asked the question gravity is actually about — a player's effect on his teammates' 5v5 chances, with his own shots removed — force predicts that effect out of sample, on next season's partly-different linemates, at r=0.42 for forwards and r=0.25 for defensemen. Stripping out the mechanical on-off term does not weaken it, so the signal is the player's own quality lifting his line, not a math artifact. Gravity is therefore a measure of playmaking and territorial impact on others, not a scoring projection. The neutral-zone transition well — the model's one novel claim — still has no historical tracking data and remains unvalidated; the full accounting lives in the Gravity model card.
V3 tiers and percentiles are calibrated separately within forwards and defensemen from the verified 2025-26 qualified population. They describe rarity within position, not equivalent impact across positions; the model does not publish a combined v3 league percentile.
Gravity v3 has three independent release channels: public display, an X-NAV transition contribution, and a simulation contribution. All three fail closed and are off in the public-launch baseline. Enabling display cannot change a valuation or simulated season; either value channel requires its own held-out validation and release decision.
When the X-NAV channel is enabled, it receives only the transition portion of v3. Direct offensive production and defensive suppression are valued elsewhere. When the simulation channel is enabled, its separately bounded term can affect team strength. Neither channel is active merely because a field is displayed.
Territorial Gravity v4 is a separate 5v5 expected-goal contract and diagnostic path. In v4 output, positive wells and domes represent estimated expected-goal impact in each phase of play. It is hard-locked off until an authorized event/shift dataset supports fitting, uncertainty intervals, held-out validation, season-specific calibration, and shadow testing. It does not currently affect X-NAV or simulation.
STRAND DNA — identity, not grades
Two 70-point wingers can be entirely different players. STRAND (Stylistic Trait & Rating Analysis for NHL Development) profiles how a player creates value — scoring pace, chance creation, suppression, usage trust, deployment difficulty — as an identity strand rather than a single grade, so roster fit and role redundancy become visible at a glance.
The double helix encodes eight skater dimensions: four offensive (scoring pace, expected goals, net on-ice value, ice time) and four defensive (defensive point shares, chance suppression, quality of competition, zone deployment). When a dimension has no data, it is honestly greyed out rather than silently filled with a neutral value — the helix shape only shows what is actually measured.
A stability backtest across 9,506 consecutive-season pairs (2008–2025, minimum 20 GP) measured which identity traits persist year to year. Every STRAND dimension is a solid or strong signal — none are noise. Expected goal pace (r = 0.89) and ice time (r = 0.88) are the most persistent skater traits, more stable than any goalie metric. Scoring pace (r = 0.83) and net on-ice value (r = 0.77) are strong skills. Usage difficulty (r = 0.70), zone deployment (r = 0.59), and chance suppression (r = 0.55) are solid signals. For comparison, the best goalie metric (freeze rate) persists at r = 0.72 — most skater identity traits are stickier than any goalie stat.
The identity profile is real: forwards are slightly more stable than defensemen across most traits, and star players are more predictable than depth players. Some trait pairs share signal — ice time and quality of competition correlate at r = 0.77, scoring pace and expected goals at r = 0.79 — but they measure different facets of the same underlying skill rather than being redundant. STRAND traits are normalized by position and rendered with league context. It answers the question a grade cannot: what does this player actually do?
Goalie Evaluation — a separate model, not a forced fit
Goalies are not scored on points, so they cannot share a skater's valuation pipeline. G-NAV prices goalies through goals saved above expected, save percentage, workload, and a fitted fair-market-value model trained on 260 goalie contracts. The result feeds the same trade-value currency as F-NAV and D-NAV, but the inputs and aging assumptions are entirely different.
Not all goalie stats are equal. A stability backtest across 769 consecutive-season pairs (2008–2025, minimum 1,000 minutes) measured which metrics actually persist year to year: freeze rate (r = 0.72) and rebound control (r = 0.69) are genuine repeatable skills. High-danger save percentage (r = 0.40), overall save percentage (r = 0.30), and GAA (r = 0.34) carry moderate signal. GSAx per 60 (r = 0.13) and medium-danger save percentage (r = 0.06) are nearly random from one season to the next — a single year tells you almost nothing about the next. The app weights accordingly: the percentile profile emphasizes the stable metrics, and all eight are regressed toward population means using their measured stability coefficients before any evaluation.
Goalies peak later than skaters. The backtest aging curve (131 goalies, signings-derived birth years) shows near-plateau through age 30, a 3% annual decline at 31–33, 4.5% at 34–36, and steeper decline after 37 — with survivorship bias flattening the oldest group, since only elite goalies still play at that age. The valuation engine uses a goalie peak age of 30 (versus 27–28 for forwards), so a 29-year-old goalie in his prime is not penalized the way a 29-year-old skater past his peak would be.
Hot goalie regression is real and strong: 78% of goalies with a save percentage at or above .915 declined the next season, with an average drop of 0.93 percentage points. The model captures this — a .925 season regresses toward .913 in the next projection, not .925.
The Team Model — what a club page actually means
A club is never one number. "Contention" collapses at least five different questions — where a team sits in today's standings, what its current roster can do, how deep its prospect and pick capital runs, how much cap room it has, and whether it can even ice a legal lineup — and answering them with one label is how a page ends up looking self-contradictory. The Ledger keeps them as named, separately sourced fields rather than one shared field two systems silently overwrite.
Phase and competitive window are the clearest case. Phase is the standings tier — conference rank, division rank, and points percentage, read straight from the current NHL table and static with respect to anything a user does. Competitive window is a read of the CURRENT roster's valuations, and only Armchair GM produces one, because only Armchair GM lets a roster diverge from what actually took the ice. A club can be a Contender by phase and Rebuilding by window at once — both are true, they are just answering different questions, and every page names which one it is showing rather than printing an unlabeled "Phase" that means the standings on one page and the roster on another.
Cap space is one calculation curated once, not recomputed per route. The published figures are measured against a $95.5M reference ceiling and already encode real accounting the app does not model on its own — LTIR relief, buried contracts, bonus overages — so a live cap-ceiling change shifts every club's room by exactly the ceiling delta rather than being re-derived by summing contract rows, which would silently drop that accounting and produce a different number on every route that tried it independently. (This is a real, previously-shipped fix: the league and team routes once disagreed by exactly $8.5M — $104.0M minus $95.5M — for all 32 clubs, because one route applied the rebasing and the other did not.)
Lineup legality is a named, machine-readable count, not a pass/fail badge. Every club is checked against the NHL's real dressed-lineup minimum — 12 forwards, 6 defensemen, 2 goaltenders — with the exact shortfall ("2 forwards and 1 goaltender," never a bare "illegal") computed by one pure counter shared by the Armchair GM simulation gate and the league API payload, so a team that has let its blue line walk in free agency shows the same deficit everywhere the count appears rather than a simulator that quietly fails and a roster page that does not say why.
F-NAV, D-NAV, and G-NAV on the Teams page are a client-side positional SUM of each rostered player's already-computed NAV — forwards, defensemen, and goaltenders bucketed and added, with any below-replacement player clamped at zero so the three splits always reconcile back to the combined total. As of NAV-02/NAV-03, the per-player values being summed are themselves genuinely position-specific, not one shared formula relabeled: calcForwardNAV, calcDefenseNAV (a fitted defensive model, holdout-validated against real team-level defensive outcomes, live for any defenseman with 20 or more games played this season), and calcGoalieNAV (diagnosed against real team-level goaltending outcomes and confirmed sound) are separate, real engine paths. The Teams-page chart's SUM step itself remains a display aggregation on top of those numbers, not an additional model of its own. The population is always the signed active roster only — never draft picks, and never the unsigned reserve list — so a team's F/D/G split is traceable to the exact same player snapshots as its combined roster total.
Expiring contracts are read by the calendar year a player's rights actually reach the market, not collapsed into a single "expiring" bucket that only lights up the season it happens. A player already signed through next season with a known following-year free-agent class is real cap-planning information today, not an unexplained zero until the year it becomes current — the ledger names the year for every club's pending UFA and RFA decisions, however far out they are known.
The GM Audit — plausibility, not just math
A trade can be mathematically balanced and still be one no general manager would make. The GM Audit is the rule layer that thinks like a front office: clause and cap legality, retention mechanics, roster slots, timeline alignment, surplus gaps, and whether the deal fits each team's contention window.
The audit publishes its objections as named flags rather than a silent score, so a rejected deal always says why — and a verdict can be argued against on the record.
The Simulation Engine — consequences on the record
Armchair GM's three-year Cup Run exists because trades are hypotheses and seasons are experiments. The engine rolls the whole league forward — aging, retirements, breakouts, drafts, cap growth, AI cap compliance — so a deadline rental or a rebuild pivot is tested against consequences, not vibes.
Goalies evolve across offseasons rather than carrying a frozen stat line. Each rollover regresses GSAx and save percentage toward population means using the backtest stability coefficients, applies the goalie age curve, and adds stochastic noise — so a team riding a hot backup into Year 2 will honestly face the regression that history says is coming.
Seeded randomness keeps runs replayable: the same decisions in the same season produce the same league, so outcomes are attributable to choices.
Data pipeline & acknowledgements
Nightly snapshots combine NHL API rosters and NHL EDGE tracking with MoneyPuck's public analytics, layered over multi-season baselines so single-season noise never masquerades as signal. In fixed-weight composites such as Gravity v3, missing evidence contributes no term, which shrinks the estimate toward neutral and lowers the reported reliability.
The Ledger stands on the shoulders of the public hockey-data community. Sincere thanks to the NHL, MoneyPuck, CapWages, and Hockey-Reference — their work makes independent analysis like this possible. Contract data is now entered by hand here rather than queried from CapWages, but the baseline this was built on came from them and the debt is worth naming. Full source credits with links are in the Glossary's Data & Sources section.
Cap & Crease is free, independent, and built nights-and-weekends. If the analysis has earned a spot in your bookmarks, a coffee keeps the data flowing.
☕ Buy Me a Coffee