How the model works
Per90 builds its own probability for every player prop it covers, from match-level data, and compares that against the bookmaker’s price. This page explains how — and just as deliberately, what the model refuses to claim. If any of it reads like hand-waving, that is a bug.
Training panel
120,409
player-match rows, 5 leagues
Markets priced
7
fouls · tackles · shots · SOT · fouls won · cards · saves
Latest refit
v20260812-1340
per-market versions in the fit table below
Projections, latest run
35,163
regenerated through the day
This page in four lines
- Each market gets a counting distribution fitted on 120,409 player-matches, shrunk toward role-and-league norms, priced at each player’s projected minutes.
- Every model must beat a naive baseline on a held-out season and calibrate within tolerance before anything it says is published.
- Beating a baseline is not beating the market — the test that counts is closing line value, scored in public, losses included.
- The edge is finite, taking it closes it, and books limit winners — the end of this page says so plainly.
Why build a model at all
Most “value” tools work by stripping the margin out of a sharp bookmaker’s odds and calling what remains the true probability. That works where sharp books make real markets — but sharp books barely price fouls, tackles or cards. On player props there is no reliable market to borrow a probability from, so this product builds its own: a statistical projection for each player in each fixture, fitted on 120,409 completed player-match performances.
The same fact is why these markets are worth pricing. Books lean on season averages and recent form for props, adjust thinly for referee and opponent, and the result is the softest prices on the board. A model that measures what they skip is the entire bet this product is making.
Counting stats need a counting distribution
Fouls, tackles and shots are counts, and player counts are streakier than a simple average suggests — the technical term is overdispersed. The model fits a negative binomial distribution per market, which carries a dispersion parameter for exactly that. The obvious simpler choice, Poisson, systematically underprices the extremes — and the extremes are where over bets live.
The negative binomial also earns its keep on minutes: it is a Gamma-mixed Poisson, so scaling the mean to a player’s expected minutes leaves the dispersion structure intact — a 60-minute projection is the same distribution at reduced exposure, not a different model. That property is what makes minutes-scaling legitimate rather than convenient.
The output is not one number but the full distribution: the probability of exactly 0, 1, 2, 3+ fouls. That is what the distribution ladders on the match pages show, and the probability of clearing a line is read straight off it. Where a line falls inside the ladder’s open tail and the split cannot be read honestly, the site shows a dash rather than an invented number.
Recent form is mostly noise, so it gets little weight
A player averaging 1.4 fouls over his last five matches has told you almost nothing — at that sample size the measurement error swamps the signal. The market over-weights exactly this kind of streak, and that is a large part of why prop prices are soft. A model that chased the same streaks would have no edge by construction.
Instead, each player’s rate is estimated by empirical Bayes shrinkage: his observed rate is blended with a prior built from players in his position and league, weighted by how much evidence he has personally supplied. How strong that pull is — the prior weight — is fitted per market by cross-validation, not guessed, and the fitted values tell their own story: a goalkeeper’s save count shrinks hardest because it is mostly the opponent’s shot volume wearing his name.
Last-5 and last-10 figures are still shown across the site, because you expect to see them — always next to the season rate, so you can see when a streak is just a streak.
What actually moves each market
Per market, the covariate that does the most work — which is often not the player’s own form.
- Fouls
- Who referees the match. A strict official lifts every player on the pitch, which is the part most books price off a season average.
- Tackles
- Possession share. A tackle count depends far more on how much of the match is played at your own end than on a player’s own form.
- Shots
- The player’s own shot rate, shrunk toward others in his position, then adjusted for how many shots the opponent concedes.
- Shots on target
- Shot volume first, conversion second — literally: the projection is the shots model thinned by the player’s own on-target rate, so a high-volume low-accuracy shooter and a sniper price differently at the same shot count.
- Fouls won
- How much a player runs at defenders. Dribble volume drives this far more than the referee does.
- Cards
- Fouls first, then how readily the referee reaches for a card. Two officials averaging four cards a match can do it for opposite reasons.
- Saves
- Opponent shot volume. A keeper’s own record carries so little that this market shrinks harder than any other on the board.
The referee deserves the special mention: foul and card projections carry the official’s own foul and card rates, shrunk toward the league mean when his sample is thin, plus his cards-per-foul — which separates a strict referee from one who merely officiates scrappy fixtures. Referee pages under trends show the same numbers the model uses.
Minutes are the biggest single input
No rate survives a player who is substituted at half time. Every projection is built on expected minutes, and the site is explicit about how sure it is: confirmed once the lineup is out an hour before kickoff, expected when a probable lineup exists, provisional when neither does and minutes come from historical start rates.
Minutes are not one number. A rotation player’s match is a 78-minute start or a 20-minute cameo — almost never the 45-minute average of the two — so the projection is priced as a mixture of those two scenarios, each with its own distribution, rather than one distribution at the blended minutes. The blend has the same average and the wrong shape: too much weight on middling counts, too little on the tails, and the tails are where over bets live.
The mixture is also conditional on the player taking part. UK books void player props when a player takes no part in the match — stake returned, not lost — so the honest price for the bet is the probability given that he plays. Averaging in no-shows would price a refund as a loss and understate every rotation player.
Provisional is not a small print — it changes the rules. A provisional edge must clear a 10% bar where a confirmed-lineup edge clears 4%, because publishing uncertainty at the same threshold as knowledge would be pretending the difference away.
From probability to price
The model’s probability becomes a fair price (1 divided by the probability), and the edge is how far the book’s price beats it. Stakes are quarter Kelly capped at 2.5% of bankroll — never full Kelly, which is ruinous the moment the model is even slightly wrong, and every model is slightly wrong.
Per-bet sizing is not the whole story, because edges arrive in correlated clusters. A strict referee lifts every fouls projection in his match at once — if that spills six edges onto the board, they are one opinion wearing six names, and staking each as if it stood alone would put six independent-looking bets on a single judgement. So advised stakes are capped across correlated groups: all edges in one match share a 5%-of-bankroll budget, and all edges under one referee share 7.5% — measured across every fixture in the multi-day publication window, not within one match, since an official can take two matches in it. Inside a single match the 5% cap is the one that binds; the referee cap exists for the window, where two of one official’s matches could otherwise carry 10% of a bankroll on a single judgement about him. Stake figures live in the bankroll calculator and the downloadable record — a tool you point at your own bankroll, not advice printed beside a pick.
Whether an edge is published is a separate question with its own gates — minimum edge by confidence, sample floors, price freshness, sanity checks against other books, and a verified settlement rule per bookmaker. The full list is on the value board, and every rejection is logged with its reason.
The match predictions are a separate goals model
Everything above prices player events. The match predictions come from a different model built for a different question: each league gets a Dixon-Coles goals model — an attack and a defence strength per team, a home advantage, and a correction for the low-scoring results that plain independence gets wrong — fitted with recent matches weighted more heavily, the weighting itself chosen by cross-validation rather than picked.
Multiplying the two sides’ expected goals through that model gives a probability for every scoreline, and those sum to the win, draw, both-teams-to-score and over/under figures on the cards. The headline number is each side’s expected goals because that is the model’s actual output; the “most likely score” is shown with its probability attached because even the likeliest football scoreline is a roughly one-in-eight event — a 1–1 or 1–0 favourite is a property of how goals distribute, not a hedge.
It faces the same discipline as the player models: it only renders after beating a naive baseline on held-out matches and calibrating within tolerance, and a side that changed division this summer is marked provisional — promoted teams enter at a deliberately below-average prior, relegated teams carry their old league’s strength translated down by a backtested gap, and real matches replace both within weeks.
Where the fit stands — and its honest limits
The improvement metric, precisely: percentage reduction in out-of-sample log loss (mean negative log-likelihood per player-match) against a naive predictor — the player’s unshrunk season per-90 rate scaled to minutes played — on a held-out season the model never trained on. Log loss scores both calibration and discrimination at once: a model that just quotes the base rate cannot score well on it. Each market serves its own separately gated version, and every figure below is read live from that version’s stored validation artefact — nothing in this table is typed by hand, so it cannot disagree with itself.
| Market | Version | vs naive baseline | vs shrunk rate | Worst in-band miss | Gate |
|---|---|---|---|---|---|
| Fouls | v20260811-1858 | +34.2% | — | 0.047 | pass |
| Fouls won | v20260811-1858 | +34.1% | — | 0.116 | pass |
| Tackles | v20260811-1858 | +30.9% | — | 0.049 | pass |
| Shots | v20260811-1858 | +32.2% | — | 0.070 | pass |
| Shots on target | v20260811-1905 | +48.2% | — | 0.034 | pass |
| Cards | v20260811-1905 | +60.9% | — | 0.022 | pass |
| Saves | v20260812-1340 | +9.5% | — | 0.021 | pass |
A dash under “vs shrunk rate” means that model’s stored validation artefact predates the second baseline; the number appears when the market next refits, and this page does not backfill measurements it did not make.
Calibration error is the worst single miss between predicted and observed frequency across probability deciles holding at least 30 observations — not the average miss, the largest, and the sample floor is stated because an unstated gate rule is how this page’s one past scoring bug got in. The pass bar is 0.100. When the model says 60%, it should happen about 60% of the time, and the number above is how far the worst-behaved decile strayed from that.
What these numbers do not prove
- One validation fold. The data subscription exposes two seasons, so “train on the past, test on the future” could only be done once. Every figure above rests on that single split until more seasons accumulate.
- Cards topping the table is partly the baseline’s failure, not our triumph. Cards are rare events, and a naive season-average predictor is at its worst on rare events — so the market with the biggest improvement number is also the one facing the weakest opponent. Read the ranking as “how much a model helps over averages, per market”, not as an ordering of which market we price best. The comparison that would settle that — our probability against the de-vigged bookmaker price on the same bets — becomes computable as captured prop prices accumulate, and will be published here when it does.
- Two markets failed this gate, were benched, and earned their way back — with one honest caveat. An audit questioned a suspicious 0.000; the investigation found the gate as it then stood only counted misses that were individually statistically significant, which let shots on target’s worst deciles — the model quoting 45–55% where reality lands 33–35% — score as perfect. The hardened gate (10 August) measured honestly, failed shots on target at 0.187 and cards at 0.177, and both markets were blocked: no builder legs, no published edges, caveats on every projection. The rebuild — model shots first, then each player’s on-target conversion; model fouls first, then cards per foul — was not designed in response to the failure: the project specification prescribing exactly that decomposition was committed to version control on 9 August, the day before the gate failed, and the git history shows it. The caveat owed to a sharp reader: the re-pass (0.034 and 0.022, 11 August) was scored on the same held-out season that produced the failure, because two seasons of data offer no second one. The design predates the failure; the truly out-of-sample test is the closing-line record below.
- Saves clears the bar by little. +9.5% is real but modest — consistent with a market where the keeper’s own contribution is small. Expect fewer published edges there, not more confident ones.
- Everything is provisional until lineups flow. The expected-lineups feed has not delivered yet, so every current projection uses the widest minutes estimate and carries the highest publication bar.
- Beating a baseline is not beating the market. The only test that counts is closing line value on published calls, in public, on the track record. And “closing line” is named now, before the record starts, so it can never look retrofitted: the last captured price for the same selection at the same book before kickoff — raw, vig included, because these markets are quoted over-only and there is no under to de-vig against. That slightly understates true closing value and is only comparable within a book; both biases run against us, which is the direction we accept. Until that sample exists, treat this page as a description, not a promise.
Capacity, and what happens to winners
Player prop markets are soft because they are small. The same thinness that leaves mispriced fouls lines on the board also means a line can move off a few hundred pounds of action. An edge published here is finite: take it, and you are part of what closes it. If this record ever has many readers, the closing line will start moving toward our fair price faster — which the CLV record would score as success, and which also shrinks what is left for the next reader. Both things are true at once.
And bookmakers limit winners. Soft books stay soft by restricting accounts that beat them, not by getting sharper — expect stake limits on these markets if you follow published edges and win, measured in months, not years. That is their business model working as designed, not a malfunction. Spreading across books extends the runway; nothing extends it indefinitely.
So this product scales as public measurement — a model, its calls, and its closing-line record, verifiable by anyone — not as a money tap that pours faster the more people hold a cup under it. Anyone selling the second thing is describing a product that cannot exist, and the honest version of this page says so before the record starts rather than after someone asks.
The web layer never computes a probability — every number shown was produced by the model pipeline and stored, and the site reads it. Where a value cannot be derived honestly, the site shows a dash and says why. Terms used above are defined in the glossary.