FootIQ prediction methodology

This page describes how FootIQ produces its predictions, from raw match data to a published pick. It explains the methods without publishing internal code or parameters that change over time.

Last updated:

Summary

FootIQ fits three statistical goal models per competition on real match data, combines them with learned weights, calibrates the resulting probabilities, blends in bookmaker prices, applies quality gates and has an AI model review every remaining candidate. Only picks that clear the minimum probability of their market are published. All probabilities are estimates; none is a guaranteed outcome.

1. Data

Match data comes from API-Football: finished matches of the current and previous season for every covered competition (goals, half-time scores and, where available, expected goals), plus fixtures, standings, injuries and suspensions, lineups and bookmaker odds. History is refreshed regularly; xG is collected gradually because not every match has it.

2. Time-weighted analysis

Recent matches say more about a team's current level than old ones. Every match is weighted by w = exp(−ξ · age in days) with ξ = 0.0019, so a match from one year ago counts about half as much as one from today. This is the optimum reported by Dixon and Coles (1997) for football data.

3. The statistical models

Dixon–Coles
A Poisson model of home and away goals with team attack and defence parameters, a home advantage and the Dixon–Coles correction for low scores (0-0, 1-0, 0-1, 1-1), fitted by weighted maximum likelihood.
Bivariate Poisson
Models both teams' goals jointly with a shared component, which captures that goals of the two teams are not fully independent. Uses venue-specific strengths.
Empirical Bayes (hierarchical)
Shrinks each team's attack and defence towards the league average in proportion to how much data exists (Gamma–Poisson). Uses xG instead of goals where available, which is less noisy over short periods.
Time-split goals
Estimates each team's share of goals in the first and second half, applied to all models for half-based markets (goal in the 1st or 2nd half, half-time/full-time).

Each model produces expected goals for both teams and a full score-probability table, from which every market probability is computed.

4. Combining the models

The models are combined per market with weights learned from a walk-forward backtest: each week in the recent past is predicted by models fitted only on earlier data, and the weights that gave the best log loss are chosen (shrunk toward equal weights for robustness). The most recent 30% of weeks are held out and used only to measure accuracy.

5. Calibration

An isotonic calibration curve per market maps model probabilities to observed frequencies, so a stated 70% corresponds to roughly 70% in the past. Once enough live outcomes exist (300 per market), curves learned from FootIQ's own labelled predictions replace the backtest curves.

6. Bookmaker prices

Where odds exist, the bookmaker margin is removed and the implied probability is blended with the calibrated model probability. The weight is learned per market and is higher where the models showed less out-of-sample skill. A candidate is rejected when the two views clearly conflict: when the bookmaker probability is more than 12 percentage points above the calibrated model probability, or, for small probabilities such as exact scores, when one is more than 1.5 times the other. When the model is more than 12 points above the bookmaker, the pick is only flagged for the AI review, because a large gap is more often a model error than a real edge. Missing odds do not disqualify a pick on their own; Correct Score has stricter rules when no price exists.

7. Quality gates

8. AI decision layer

Every candidate that passes the gates is reviewed by one AI model (currently from OpenAI) with the complete match dataset. It may accept, reject or lower a probability, but never raise it. If no model is configured or it fails, deterministic rules decide. See AI football predictions.

9. Ranking and publishing

Within each category the qualified picks are ranked by final probability (ties broken by lower model disagreement, then higher data quality), and the category's daily limit is filled from the top. Categories are independent. Picks are published before kick-off and are never published within 15 minutes of kick-off.

10. Confidence

The "confidence" shown with a prediction is its final probability in percent, after calibration, bookmaker blending and any AI reduction. See prediction confidence.

11. Settlement

Predictions are settled automatically from the official regular-time score (90 minutes plus stoppage time; extra time and penalties do not count). Half-based markets use the half-time score. If a match is cancelled, abandoned or postponed, its pending predictions are void and do not count in win rates.

12. Learning from results

Every evaluated candidate (published or not) is labelled after full time. These outcomes are used to re-learn the model weights, the bookmaker weight and the calibration, so the system is measured against reality continuously.

13. Limitations