MechaPip ELO Edge ยท Invite only. Blackbox. No public source.
How Elo ratings work
Desk notes by Elo Star, Head of the Sports Desk โ meet the house.
This page explains the rating math behind MechaPip forecasts. The live board publishes calibrated Elo win probabilities. Every sport uses the same pure Elo engine, enriched with matchup data: player form, team advanced stats, injuries, and probable starters. The engine itself stays private. This is not a clone kit, not source code, and not financial or gambling advice.
What this number is
A MechaPip forecast is a win probability for one side of a scheduled game. If the board shows the home team at 62%, that is a model estimate that the home side wins about 62 times out of 100 similar matchups. It is not a ticket, not a price target, and not a promise.
Probability is a long-run rate. A 62% favorite still loses often. A 55% favorite is a coin with a slight lean. The interesting claim is calibration: among games labeled about 60%, close to 60% should actually go that way over a large sample.
Forecasts are for entertainment and analysis. They are not financial advice and not gambling advice. Prediction-market rules vary by place. This site does not sell event contracts.
Elo ratings
Arpad Elo's system (used in chess, then borrowed everywhere) stores one strength number per team. After a final, the winner takes points from the loser. The expected score is:
P(A beats B) = 1 / (1 + 10^(-(R_A - R_B + HOME_ADV) / 400))
That curve is the reason a 200-point gap feels like a solid favorite and a 50-point gap does not. Sports teams fit the same picture: a season is a long paired-comparison tournament.
The update rule
After every final, each rating moves by a fraction of the error between the outcome and the expectation:
R_new = R_old + K * MOV_multiplier * (outcome - expected)
- K factor. How strongly one result moves a rating. Tuned per sport: baseball plays 162 games and learns gently; football plays 17 and must react hard.
- MOV multiplier. Margin of victory: a blowout moves ratings more than a one-run win, capped per sport so a 30-point NBA rout does not overreact.
- Home edge. Added to the home rating when the site is not neutral. Tuned per sport. A packed NFL Sunday is not a college hoop gym.
Season regression
Between seasons, ratings regress toward 1500 so last year's champion does not freeze the next opener. Rosters turn over, coaches change, and a fresh season starts from memory, not from last October's momentum.
The data layer
Elo is the skeleton. Matchup data is the muscle. Each forecast time adds:
- Player form. Season player stats are scored into a per-team strength adjustment, so a team playing far better than its rating gets the benefit.
- Team advanced stats. Per-sport composite metrics (pace, EPA, expected goals, and similar) z-scored against the league and attached to the rating.
- Injuries. Out listings can move a rating before the jump ball.
- Probable starters. Where the sport cares (pitcher, goalie, quarterback), a named starter can nudge the team rating.
- Rest. Days since last game adjust the matchup context.
- Altitude and referee lean. Sport-specific nudges where the sample supports them.
Complete data is the rule: every sport carries full game history, teams, players, injuries, live scores, starters, and advanced stats โ all refreshed from ESPN and league APIs every cycle.
Calibration
A rating curve is not automatically a probability. If the model says 70% and those games only hit 58%, the curve is overconfident. Calibration is a reliability pass: learn a small map from raw probability to observed frequency, then apply that map at publish time.
MechaPip uses Platt scaling:
P_cal = sigmoid(a * logit(P_raw) + b)
One slope, one intercept, fitted from settled results per sport. It keeps the raw curve's shape and stretches it so that games labeled 60% land near 60%, not 75% and not 51%.
How one forecast is built
One game, live board, current default:
- 1. History. Past results update each team's Elo rating via the K-factor and margin rules above.
- 2. Pregame data. Player form, team advanced stats, injuries, probable starters, rest, and sport-specific extras adjust the ratings used for this kickoff.
- 3. Raw Elo chance. The logistic curve with home edge produces a home win probability.
- 4. Calibration. Platt scaling stretches that raw number so buckets match outcomes.
- 5. Publish. Favorite and probability go to the board, the stream HUD, and the day's posts.
What you see on mechapip.com is step 4: calibrated Elo plus pregame data.
What live uses
- Live board (default)
- Pure Elo ratings. Pregame data layer. Platt calibration. Seven sports. This is the number on mechapip.com.
- Live Elo updates
- ESPN finals are ingested as they settle, ratings are updated in place, and the same ratings feed the board, the stream, and the daily posts without a restart.
The stack is deliberately one engine: pure Elo plus complete data. No Glicko-2, no tree ensembles, no secondary model kinds. Accuracy is a hit rate on the board, not a dollar figure and not a promise that any given night hits the mark.
Seven sports, same math, different noise
Same engine. Different sample sizes, different starters, different calendars.
- NBA
- 82-game grind. Back-to-backs, altitude, star load. Rest and injuries matter every night.
- NFL
- Short season, fat uncertainty. Weather, byes, and quarterback-out flags punch above their sample. Referee lean is a small extra where used.
- MLB
- 162 games, starter swings, park context. Longest memory, lowest per-game noise, and a strong pitcher-start adjustment.
- NHL
- Goalie quality can outweigh a hot skater. Travel and back-to-backs show up as rest and starter nudges.
- WNBA
- Compact schedule, travel density, rest clustering. Same Elo math, tighter window.
- NCAAF
- Roster churn every year. Heavy season regression. FCS/FBS context and Saturday volume.
- NCAAMBB
- Hundreds of teams. Transfer portal noise. Conference strength lives in the prior, not in a handwritten power ranking.
FAQ
Why can a 55% favorite lose?
Because 55% means 45% the other way. Sports are noisy. One bounce, one missed call, one bullpen night. A well-calibrated 55% should win about 11 times in 20, not 20 in 20. If 55% favorites won 90% of the time, the model would be lying about its confidence.
Why do college sports look different?
Huge team counts, yearly roster reset, and uneven schedules. Season regression pulls everyone toward 1500 when the roster turns over, so a mid-major with a thin sample does not get a fake lock against a blue blood.
Why do injuries move the number?
A rating is a team average. An Out listing says tonight's team is not that average. The engine nudges the rating before computing the probability. It is not a second model. It is a correction so yesterday's strength does not pretend a star is suiting up.
Why is this page not source code?
The engine is invite-only. This host explains the public ideas: Elo ratings, matchup data, calibration. It does not publish a kit to rebuild the bot. Daily numbers live on www.mechapip.com.
Access is invite only
The engine is a closed blackbox. No public source. No public clone. Invited operators get the board and the runbook. Everyone else watches forecasts on mechapip.com.
Business inquiries: mechapip@mechapip.com
Live forecasts stay on www.mechapip.com. This host does not replace the board.