MLB model handbook

FUSE decision and execution guide

Official MLB FUSE is a market-free baseball model with two internal probability heads. A decision is eligible only when both heads independently choose the same side at 60% or greater. APEX and LINX never vote on that decision. Unified and SOLO are append-only prospective challengers; they are displayed for comparison and cannot create an official order, settlement, or P&L entry.

Last updated
July 31, 2026
Guide and legal documents
FUSE handbook

Scope and strategy roles

This handbook applies only to the MLB implementation of FUSE. Football FUSE uses a different ClubElo/Davidson method and does not inherit the MLB two-head rule.

Champion is the only execution-eligible MLB FUSE strategy. Unified and SOLO are challenger tracks. They write immutable decisions for forward comparison, but their picks never enter Today’s official signals, Yesterday’s settlements, the official FUSE curve, paper orders, or official P&L.

  • Champion: calibrated Logistic and calibrated XGBoost heads must both clear the same 60% side threshold.
  • Unified: the fixed 50/50 mean of the same two heads must clear 60%; diagnostic only.
  • SOLO: the frozen calibrated XGBoost head must clear 60%; single-head candidate only.
  • A true standalone successor would require a separately registered artifact, preregistered rule, and its own prospective proof record. SOLO is not that successor.

The six market-free inputs

FUSE creates a home-win probability from a frozen six-feature baseball snapshot. Odds, implied probability, bookmaker consensus, Polymarket price, APEX, and LINX are prohibited as model inputs. Each feature is oriented so that a larger value indicates greater home-side advantage.

  • Away starter FIP minus home starter FIP.
  • Home starter K/BB minus away starter K/BB.
  • Home team win percentage minus away team win percentage.
  • Home run differential minus away run differential.
  • Home expected wOBA minus away expected wOBA.
  • Away bullpen fatigue minus home bullpen fatigue.

Training and calibration

One frozen artifact produces a Logistic probability and an XGBoost probability. Each raw head is calibrated separately with its registered Platt calibrator; missing, non-finite, or out-of-contract values fail closed.

The current artifact is fundamentals_v1.20260724T104554494274Z.candidate.pkl, checksum ffd2b4111a75c20434f183be1c39ce3a55a377587e4f66aa6645f886736b3f97. Changing the artifact, features, calibration, threshold, or decision rule requires a new version and a new prospective evaluation boundary.

Official Champion decision rule

For the home side, both calibrated home probabilities must be at least 0.60. For the away side, both home probabilities must be at most 0.40. Every other combination is an abstention. Official confidence is the weaker of the two agreeing heads after it is converted to the selected side.

This is an internal two-head agreement rule, not agreement between APEX and LINX. Those models may appear beside FUSE as external comparison only and cannot approve, veto, or reverse a FUSE side.

A 60% threshold is a selection rule, not a guarantee that the selected games will win 60% of the time.

Unified and SOLO challenger rules

Unified selects a side when the fixed arithmetic mean of the calibrated Logistic and XGBoost home probabilities reaches 0.60 or 0.40. SOLO applies the same threshold to the frozen calibrated XGBoost head alone.

Both tracks are forward-looking diagnostics. Their records are append-only and remain operationally isolated from Champion. A challenger can be promoted only through an explicit versioned decision after sufficient prospective evidence; a favorable retrospective slice cannot promote it automatically.

When a decision may be frozen

The decision job accepts regular-season games only and runs in a one-to-three-hour window before scheduled first pitch. Its feature snapshot must be no more than 90 minutes old, both probable starters must exist and match the snapshot, and all six features, artifact metadata, and runtime checks must be complete.

The job writes one immutable result per game and strategy: pick, pass_confidence, or pass_missing. Re-running the job cannot rewrite an earlier decision. Missing evidence causes abstention rather than an inferred pick.

Order and execution logic

Champion chooses the side before any market quote is consulted. Only after the side is frozen may a venue quote be read for execution availability and later comparison. Market movement can neither select nor veto the side.

For new orders, the registered fuse_confidence_band_v1 policy freezes the paper target from Champion’s recorded weaker-head confidence: US$25 at 60.00%–60.99%, then US$5 more for each full percentage point, capped at US$75 from 70% upward. The reference bankroll is fixed at US$10,000 and total executed FUSE paper exposure is capped at US$150 per ET game date. Neither market price, APEX/LINX, recent P&L, nor a winning or losing streak can alter that target.

Only after the target is frozen does the execution layer check quote availability and whether quoted capacity supports the full target. A missing quote, insufficient capacity, or the daily cap leaves the executed stake and P&L unavailable; no partial order is created. The order record must be written before first pitch; a missed service window is recorded as timing_missed and can never be backfilled as a post-start order. Earlier public orders remain sealed under their historical US$50 policy and are never recalculated. These are auditable paper-simulation controls, not personalized sizing. Unified and SOLO never create an order.

How to read the thinking wall

The thinking wall is a readable rendering of frozen pregame evidence: selected side or abstention, both internal head probabilities, starter identities, feature support and risk, decision time, and decision identifier. It must not add facts after the game.

Support, risk, and neutral labels describe the signed feature snapshot; they are not SHAP values, causal claims, or global feature importance. Weather, late injuries, line movement, APEX, and LINX are not current FUSE decision inputs. A pass or missing-data state must never leak an implied pick.

What the historical diagnostic does—and does not—show

At the frozen historical cutoff, Champion recorded 33 wins in 46 selections (71.7%); Unified recorded 42 in 63 (66.7%); and the 17 games added by relaxing two-head agreement recorded 9 wins (52.9%). The difference is not statistically decisive: the two-sided Fisher comparison is p=0.229, and the samples are small.

The XGBoost head alone recorded 52/76 (68.4%) in the same retrospective reconstruction and 19/30 (63.3%) outside Champion’s 46 selections. These are research diagnostics, not a forward SOLO record. All Champion and Unified selections in this slice were home-side picks, and coverage was only 46/1,448 (3.2%) for Champion and 63/1,448 (4.4%) for Unified.

  • Champion exact 95% interval: 56.5%–84.0%.
  • Unified exact 95% interval: 53.7%–78.1%.
  • Incremental 17-game exact 95% interval: 27.8%–77.0%.
  • Calibration used 2025 out-of-fold predictions rather than a fully independent calibrated test set.
  • Historical starter identifiers may reflect final actual starters, creating hindsight risk.

These limitations are why Champion remains unchanged while SOLO is tested prospectively instead of being promoted from the retrospective result.

Prospective proof gates

FUSE targets a win rate above 55%; this is a research objective, not a promised outcome. The proof protocol excludes a 150-pick burn-in and requires a frozen terminal window of 700 official picks before any durable performance claim.

Coverage must reach at least 20% of eligible games, decision capture at least 95%, and uncertainty must be reported. The registered stability check uses 20,000 seven-day block-bootstrap samples so clusters of games are not treated as fully independent.

  • Compare Champion, Unified, and SOLO on the same prospective opportunity set.
  • Report selections, abstentions, wins, losses, accuracy, coverage, confidence intervals, and drawdown—not win rate alone.
  • Keep rule, artifact, cutoff, and checksum visible; start a new cohort after any material change.
  • Do not promote on a short hot streak, a retrospective filter, or a challenger’s higher headline rate.

Website display contract

Today’s official signals, Yesterday’s settlements, and the official FUSE curve represent Champion only. The thinking wall’s official decision is also always Champion; when the backend explicitly publishes a challenger snapshot, the wall may show it only as a separately labelled, non-execution diagnostic. The comparison archive may show Champion, Unified, and SOLO together under the same challenger labels.

Game cards may reveal the frozen Champion state before first pitch only to entitled viewers under the normal access rules. Schedule modules may show whether FUSE is locked, abstaining, awaiting data, or published without manufacturing a side. APEX/LINX agreement is always labelled external comparison.

Every displayed FUSE target, executed stake, confidence, policy version, risk limit, execution status, and reason must come from the immutable backend order record. The website may format those fields but must never infer or recalculate a stake. A blocked or unavailable order displays an em dash with its recorded reason.

Claims the record does not support

Do not describe 71.7%, 66.7%, or 68.4% as an expected future win rate. Do not call SOLO a fully independent model, claim that the 55% goal has been achieved, treat missing P&L as zero, or merge challenger results into Champion.

FUSE is allowed to be selective and to abstain. The purpose of this contract is not to create more picks; it is to preserve a falsifiable decision process whose failures can be measured without rewriting history.

Continue readingHow to read WiseLine

Need clarification?