Skip to the copy
WIRE

Connecting to the wire

Desk story / History

AI and Machine Learning in Sports Betting Models

The sports bettor who uses machine learning is purchasing a probabilistic edge in a market priced by aggregated human belief. The mechanism is straightforward. The implications are not.

Copy sheetSM365
Filed
Byline
Jerry Boyd
Length
773 words, about 4 minutes
Copy ID
SM365-BCD4E24D
ai machine learning sports betting model with probabilistic edge and aggregated belief pricing
ai machine learning sports betting model with probabilistic edge and aggregated belief pricing

Sports betting markets, like all betting markets, price in the aggregate belief of all market participants. The closing line reflects what the collective wisdom of professional bettors, syndicates, and public bettors believes will occur. No individual bettor possesses perfect information. All bettors possess the same information, or close to it.

The expected utility maximizer in sports betting wins if and only if their model of the future differs from the market's model of the future in a favorable direction. Machine learning models offer this advantage in specific domains where data quantity exceeds human cognitive capacity.

Consider an Austrian economist's framing: the bettor is not trying to predict the future. The bettor is trying to identify a discrepancy between his subjective valuation of an outcome and the market's valuation. Machine learning excels at this in domains where the relationship between variables is non-linear and involves high dimensionality.

A team's expected score in an NBA game depends on: recent form, home-court effect, rest days, specific player injuries, player-on-player matchups, weather (irrelevant for indoor sports), schedule strength, coaching changes, salary cap implications of that season's spending. A human model might incorporate eight to ten of these variables. A machine learning model can incorporate 500 variables if training data permits.

The Mechanism

The model is trained on historical data. Every NBA game from the past decade, with all relevant variables tagged. The model learns the relationship: when these inputs occur, this outcome follows, with this probability. The model is backtested. Does it predict better than the closing line. If yes, the model has found an edge. The bettor deploys capital accordingly.

The practical obstacle: the edge, if it exists, is small. Perhaps two percent. The closing line is accurate to within one percent on average. The machine learning model is accurate to within 0.98 percent on average. The difference is marginal. The bettor needs to deploy large enough capital to profit meaningfully from a marginal edge.

This is where capital constraints bind. A model with a 0.02 probability edge on a one-hundred-dollar bet makes two dollars. The cost of tracking transactions, paying for the model, paying for the computing infrastructure, all of that might exceed two dollars. The bet is unprofitable at small scale.

But a model deployed across 10,000 bets can generate meaningful profit. The variance across individual bets is high. The expected value across the portfolio is positive. The player who has sufficient capital can exploit this edge repeatedly until the market adjusts or the edge disappears through overfitting.

The fundamental question in machine learning for sports betting is not whether the model can find an edge. The question is whether the edge persists long enough to recoup the cost of generating the model.

Market efficiency is the constraint. If the model identifies an edge, other models will identify the same edge. The closing line will adjust. The edge will compress. This process accelerates as more capital pursues the same opportunities.

Historically, the first wave of machine learning models in sports betting found substantial edges. Early 2000s models identified patterns in line movement that persisted for years. By 2015, those same patterns had been arbitraged away. The market had learned.

Current machine learning models must update continuously. Static models become obsolete. A model trained on 2015-2019 data is useless in 2024 because team composition, player skill distribution, and coaching philosophy have all shifted. The bettor must retrain constantly, burning computational resources and capital.

The Risk Structure

A machine learning model that works for the backtested period may not work forward. The model may have discovered spurious relationships. It may have overfit. It may have discovered relationships that were real in the past but are no longer true.

The deployed model carries hidden risks. If a model was trained on data from teams with a certain average salary cap, and one season all teams increase spending by 20 percent, the model's predictions shift. The model has not learned the relationship between salary cap and performance. It learned the relationship between salary cap and performance in a specific era. The era has changed.

Capital requirements are non-trivial. A deployed model, if it has a 0.02 edge per bet and the average bet is 100 dollars, requires approximately 100,000 dollars of bankroll to sustain a 95 percent confidence that the model will remain profitable through normal variance swings.

The Austrian economics lens clarifies the core issue: the bettor is not predicting. The bettor is arbitraging. The machine learning model is the tool of arbitrage. It identifies where the market's price differs from the bettor's estimate of true probability. It is not a crystal ball. It is a lens.

Share the copy