I Built a Prediction Bot Just to Prove Your Betslip Will Lose

I spent weeks building a prediction engine for the top five European football leagues. The goal was simple, bordering on a developer cliché: write enough Python to find an edge in the market.
The crazy part is that the underlying predictions actually worked. Nearly 70% of my strong predictions were right. But when I ran a strict historical betting simulation, the system still lost money.
Figuring out exactly why my highly accurate model was bleeding out taught me more about the gambling industry than any dataset ever could. It boils down to a masterclass in margins, concentrated errors, and the mathematical reality of why the house always wins.
The Illusion of Being "Right"
To build the engine, I wired together three concepts: Elo ratings updated for goal difference, a time-weighted Dixon–Coles goals model to estimate attack and defence, and an XGBoost classifier trained on expected-goals (xG) form and rest days. Everything was evaluated in a season-by-season forward approach, meaning the models only ever saw matches that happened strictly before the one they were trying to predict.
Initially, the aggregate numbers looked beautiful. I defined a "strong prediction" as any match where the model's highest-probability outcome was at least 60%, with a massive 25-point lead over the second choice. For this group, the headline hit rate was 69.3%.
But matching an average confidence with an average hit rate across a group is not proof of a truly calibrated forecaster, because massive errors can simply cancel each other out across different probability ranges. Furthermore, in sports betting, you don't bet on every match; you only pull the trigger when the price justifies the risk.
So, I filtered those strong predictions to create a ledger of "selected simulated bets," betting only when the recorded Bet365 historical odds exceeded my model's break-even odds. *
The moment I applied that price filter, the population shifted drastically:
Selected simulated bets: 925
Winning bets: 550
Selected-bet hit rate: 59.5%
Let's translate that into South African Rands(R). If you put down R1 on each of those 925 selected bets, you would have wagered a total of R925.00. When the dust settled, your tickets would have returned R877.37, leaving you with a net loss of R47.63. That is a -5.1% return on your money.
How does a model correctly predict the future 69.3% of the time in its strong group, yet burn cash in a simulation? Because winning picks and profitable picks are two entirely different things.
The Overround (Or Why the Math Never Equals 100%)
The secret to the house edge is a concept called the overround.
In a perfectly fair market, the implied probabilities of all possible outcomes would equal exactly 100%. But bookmakers are not running a fair market. Take a standard football match with decimal odds of 1.60 for the Home win, 4.20 for the Draw, and 5.50 for the Away win. If you convert those into implied probabilities, they sum to approximately 104.49%.
That excess above 100% is the market’s overround. Now, the overround doesn’t guarantee you will lose every wager; if an outcome’s true probability is high enough, an individual selection can still have a positive expected value. However, the overround acts as a toll that artificially shrinks the quoted odds, giving you very little margin for error.
The Break-Even Trap
This brings us to the math every bettor needs to internalize before risking a cent.
| Implied Probability | Break-Even Odds |
|---|---|
| 75% | 1.33 |
| 60% | 1.67 |
| 50% | 2.00 |
| 25% | 4.00 |
The formula is brutal and simple: Break-even odds = 1 / probability.
If your model calculates that a team has a 62% chance of winning, pulling the trigger only makes mathematical sense if the bookmaker offers odds above 1.61. If they offer 1.50, you will eventually lose money over time even if your 62% prediction is flawlessly accurate.
Being right is not an edge. You have to be right at the right price.
Concentrating the Mistakes
When a bettor finds a discrepancy between their prediction and the bookmaker's odds, they usually think they have found "value."
When my code flagged a high-probability outcome and the market offered a generous price, my initial thought was that the bookies had slipped up. But the massive drop from a 69.3% hit rate on strong predictions to a 59.5% hit rate on selected bets reveals a stark reality: adding a favourable-price filter radically changes the population being evaluated.
Filtering for apparent value likely just concentrated the model’s mistakes. The lower hit rate doesn't strictly prove that the bookmakers possessed secret information the model missed, but it strongly suggests that when you heavily disagree with a highly efficient closing market, you are often just isolating your own blind spots.
A Frontend for Transparency
I turned the experiment into a website where anyone can inspect the results for themselves.
The Football Experiment Website
The Overview shows the headline findings: the hit rate for strong predictions, the number of selected bets, and the simulated return. Below that, an interactive historical ledger lets you filter selections by league or season, search for a match, and expand individual entries to inspect the final score and the model’s home-win, draw, and away-win probabilities. Each selection shows the recorded odds, whether it won, and its simulated profit or loss.
The Behind the Model page goes deeper. It breaks down performance by league, season, and confidence range, while a cumulative profit-and-loss chart shows how the simulation unfolded. It also explains the modelling approach, the tools I used, and why I chose them.
Readers can follow the original data-source links, inspect the provenance records, download the historical selections and match predictions, or open the public GitHub repository to examine the code.
Everything comes from a fixed historical snapshot. The site gives readers the evidence to explore what the model predicted, which selections it made, and how those selections performed.
The Takeaway
Betting has never been more frictionless, heavily subsidized by algorithmic marketing and living in our pockets 24/7. But the core product design, the overround, the odds structure, the immediate access, remains perfectly optimized to slowly bleed players dry.
If you are building a predictive model for the sheer fun of the engineering challenge, it’s a brilliant way to learn data science. But if you actually want to deploy it to beat the sportsbook, you have to accept the reality of the game you are playing: price affects profitability independently of how frequently a selection wins.
My code was a highly accurate forecaster. The global betting market was simply a better one. That isn’t a failure; it’s a measurement.
Explore the data: The Football Experiment Website
Inspect the code: GitHub Repository
Author: Keabetswe Mmakola
Legal Notice & Disclaimer
1. Informational Purposes Only
All content, articles, models, and code snippets published here are provided for educational, informational, and entertainment purposes only.
2. Not Financial or Betting Advice
Discussions regarding sports analytics, betting markets, odds, and predictive modeling do not constitute financial, investment, or gambling advice. The author is not a licensed financial advisor. Any bets placed or financial decisions made based on this content are done strictly at your own risk. Past performance of any model does not guarantee future results. If you choose to gamble, only risk capital you can afford to lose.
3. "As-Is" Software and Code
All code, scripts, and technical architectures are shared "as-is" without express or implied warranties of any kind. The author makes no guarantees regarding the accuracy, reliability, performance, or security of any provided code.
- The ledger spans August 2021 to September 2026, including a partial 2026/27 season.