Skip to content
Giordano Cabral

Essay · 2026

What did 124 AIs get wrong at the 2026 World Cup?

Between June 11 and July 19, 2026, 124 artificial intelligence models predicted the 104 matches of the World Cup under the same rules as an office pool — 5,797 predictions in 38 days, with an open ranking. The machines' consensus had Brazil as champion, and Brazil was eliminated in the round of 16 by Norway. In the semifinal France 0 × 2 Spain, none of the 62 AIs that answered that match predicted the Spanish win. The best human participant scored 629 points and finished ahead of 121 of the 124 AIs. The models' errors came aligned: in the same direction and at the same time, which reduces the value of consulting several models as a second opinion.

Noticeboard with a slightly uneven grid of blank cards held by pins; a single red card.

01 · O que 124 inteligências artificiais erraram na Copa

Haaland's second goal

July 5, 2026, MetLife Stadium, round of 16. In the 79th minute, Erling Haaland opened the scoring for Norway; in the 90th, he scored a second. Neymar pulled one back with a penalty in stoppage time, and Brazil exited the World Cup earlier than in any edition since 1990.

Three weeks earlier, in an interview with Folha de Pernambuco, I had said two things. The first: "AIs tend to be very conservative. They follow the statistics, avoid risks, and usually converge on similar answers. Humans, on the other hand, take more risks, follow intuitions, and explore possibilities that machines normally dismiss." The second, a few lines later: "I expect Brazil to win, however unlikely that may be."

The consensus of the 124 artificial intelligences taking part in the pool pointed to Brazil as world champion — the same guess that I, speaking as a fan, had made. The piece deals with what the open record shows after the World Cup ended.

02 · O que 124 inteligências artificiais erraram na Copa

What the experiment was, and what it does not measure

Between June 11 and July 19, 2026, the FAROL group, coordinated by Giordano Cabral and Filipe Calegário at the Centro de Informática of UFPE, had 124 artificial intelligence models predict the 104 matches of the World Cup. There were 5,797 predictions in 38 days, in an open ranking that remains published at bolao.arenadasias.com.br and whose code is at github.com/giordanorec/bolao-copa-2026.

The design is deliberately simple. We gathered statistics, news, recent form and injuries into a single dossier, and sent the same context and the same prompt to all the models. Most ran via API. A subset, which we called Série A (Tier A), predicted through the interface: it opened the browser, searched, and then set the score, as a person would. And people entered the same ranking, under the same rules, with no handicap.

The scoring rule is the one used in an office prediction pool: 10 points for the exact score, 7 for the winner with the correct goal difference, 5 for the winner without the correct goal difference, 5 for a draw without the exact score, and zero for everything else — with all points doubled in the knockout stage.

What it measures is points, under an arbitrary scoring rule, over a sample of 104 matches: a point-prediction test in a high-uncertainty setting. It does not measure intelligence, nor quality of reasoning, nor calibration in the technical sense, which would require probabilities rather than scorelines. The ranking is not a table of "which AI is better."

03 · O que 124 inteligências artificiais erraram na Copa

Who won, and what the score conceals

Two AIs finished tied at the top with 636 points: Mistral Small 3, with 22 exact scores, and Grok 4 Fast Reasoning, with 20. Third-place OpenAI o4-mini scored 634. The highest-ranked human participant, listed as Gabriel in the public ranking, scored 629 points with 19 exact scores and finished ahead of 121 of the 124 AIs. I was the second-highest-ranked human, with 551.

The top human participant scored 629 points and finished ahead of 121 of the 124 AIs. Two of them tied at 636; a third finished between those two scores.
The top human participant scored 629 points and finished ahead of 121 of the 124 AIs. Two of them tied at 636; a third finished between those two scores.

On average, the machines scored more than the humans; at the top, one person finished ahead of nearly all the machines.

The first finding calls for a caveat. The humans' average in the pool was 55 points, with a median of zero — because most people who created an account never made a prediction. Comparing the AIs' average, which answered game after game, with the humans' average, most of whom largely did not play, compares nothing. The comparison that holds is at the top, and at the top the number of people is small. Among those who actually played, the accuracy gap is more modest than the ranking suggests: 15.3% exact scores across the AIs, 15.4% in Série A, 11.4% among the humans.

The scoring rule also shapes the result. The champion, Mistral Small 3, nailed 22 scorelines and made 636 points; Grok 4 Fast Reasoning nailed 20 and also made 636, in second place; o4-mini nailed 22, same as the champion, and came third with 634. More exact scores did not mean more points. What ordered the podium was the combination of the value of goal difference, the doubled weight of the knockout stage, and picking the winner correctly — that is, the scoring rule.

04 · O que 124 inteligências artificiais erraram na Copa

Where they went wrong together

In the France 0 × 2 Spain semifinal, none of the 62 AIs consulted predicted a Spanish victory. In the third-place match, 5 of the 63 picked England and none predicted the exact score. Across the 104 matches, 7 of the 124 models reached twenty exact scores, and 5 exceeded that number.
In the France 0 × 2 Spain semifinal, none of the 62 AIs consulted predicted a Spanish victory. In the third-place match, 5 of the 63 picked England and none predicted the exact score. Across the 104 matches, 7 of the 124 models reached twenty exact scores, and 5 exceeded that number.

The semifinal that determined the world champion was France 0 × 2 Spain. Of the 62 AIs that responded for that match, none predicted a Spanish victory; the consensus pointed to a draw. In the third-place match, France 4 × 6 England, with ten goals, none of the 63 predicted the exact score and only five even picked England to win. In the final, Spain 1 × 0 Argentina, decided by a goal from Ferran Torres in the first minute of extra time, 41 of the 62 wrote down the same 1 × 1 — the strongest consensus of the entire knockout stage. Seventeen predicted the exact 1 × 0 score.

The pattern was already established in the group stage. The consensus for Germany 7 × 1 Curaçao was 3 × 0, with 41 votes. For Qatar 1 × 1 Switzerland, it was 0 × 2, with 45. Across the 104 matches, only 5 of the 124 models exceeded twenty exact scores; 7 reached twenty.

More relevant than the error rate is the shape of the error. Teams lose, goals go in, and no serious forecaster would promise to get France 4 × 6 England right. When they erred, they erred in the same direction, at the same time, and often writing literally the same scoreline. If the errors of 124 forecasters were independent, the set would cover a good part of the space of plausible outcomes and someone would get the upset right by construction. In practice, the set behaved close to a single forecaster.

05 · O que 124 inteligências artificiais erraram na Copa

Why the errors align

Part of the explanation is a decision on our part, and it is fair to start there. We sent the same dossier and the same prompt to everyone, precisely so the comparison would be between models, not between who phrased the question better. That favors the comparison and reduces diversity: we eliminated from the outset one of the sources of variation that would make the errors diverge.

The rest comes from what these systems are. Language models trained on widely overlapping corpora, optimized to produce the most likely continuation, respond to "what will the score be" with the mode of their distribution. And the mode of a football match is a small set of scorelines: 1 × 0, 1 × 1, 2 × 1. The pool's rule, in turn, pays 10 points precisely for exactness, and doubles everything in the knockout stage, which is where the improbable concentrates. The tournament rewards, with double weight, exactly the region where the mode falls short.

The literature helps situate this, and it does not say that models predict poorly. Schoenegger and Park entered GPT-4 in a three-month forecasting tournament on Metaculus, with 843 human participants, and measured probabilistic forecasts significantly less accurate than the human crowd median (arXiv 2310.13014). The following year, Schoenegger, Tuminauskaite, Park, and Tetlock aggregated 12 models across 31 binary questions and obtained a result statistically indistinguishable from a crowd of 925 human forecasters (arXiv 2402.19379). In that format, aggregation made up for individual performance.

The two findings are compatible with each other, and neither contradicts ours — because they measure something else. Both work with probabilities for binary questions, assessed through calibration. A prediction pool asks for a point estimate in a discrete space with dozens of possible outcomes and rewards exactness. Averaging probabilities is an operation that improves the estimate; averaging point predictions pushes the group even further toward the mode, which is precisely where the unlikely outcome is not. That is what our Bola de Cristal did when it picked Brazil as champion: consensus working as designed.

06 · O que 124 inteligências artificiais erraram na Copa

What is wrong with this experiment

The experiment has five limitations, besides the one already mentioned: a pool does not measure intelligence.

The participant count is not stable. I use 124, the closing count on July 22, 2026, and the only one consistent with "121 of the 124"; the archive itself publishes other counts, and the discrepancy is documented in the verification notes. An experiment that does not say how many subjects it had has a recordkeeping problem, and that problem is ours.

Fewer than half of the possible predictions were recorded. One hundred and twenty-four models times 104 matches would yield 12,896 predictions; there were 5,797. Models joined midway through the tournament, stopped responding, or failed. That is why the denominator changes from match to match: 62 in the semifinal, 63 in the third-place playoff, 124 overall. "None of the 62" and "12 of the 124" are not statements about the same set.

The human sample is not a sample. The zero median shows this. There was no recruitment, no criteria, and whoever joined did so because they saw the link.

The ranking is an artifact of the scoring rule. Change the 10-7-5 scoring system and the podium shifts; a Brier score would reward calibration and produce a different order, with a different winner. We chose the rule people use in office pools, which makes the experiment understandable and, to the same extent, dependent on a convention with no scientific standing.

A World Cup is 104 matches. That is too few to separate competence from luck. Between the first- and third-place finishers there were two points of difference. This data does not allow us to say whether Mistral Small 3 predicts football better than o4-mini; it scored two points more, in a single World Cup.

07 · O que 124 inteligências artificiais erraram na Copa

Consequences for the use of AI at work

The first is that asking the same thing to several models is not a second opinion. Agreement among systems trained on similar material, reading the same context, and optimized for the same type of response measures the popularity of the answer. And it looks a lot like confirmation: five models saying the same thing produces a sense of security that the data does not support. Variance needs to enter through the question: different designs, adversarial hypotheses, people who disagree.

The second is that the model helps least precisely where the decision is made. In matches with a predictable script the AIs were good; it was in the four or five games that turned the World Cup that they were unanimous and wrong. Work has the same shape. The routine case, where they excel, is also the one where being wrong costs little; the high-cost exception is the one they did not predict.

The third is operational: ask for the distribution. "What will the score be" returns the mode. "What are the five most likely scores and with what weight" returns something you can work with — including the information that the model has no idea. It is the difference between a point estimate and an estimate with declared uncertainty, and it is what lets you notice that the answer is a guess before acting on it. The tone of the answer does not indicate this: a model writes "Brazil champion" with the same composure with which it writes "Germany beats Curaçao," and nothing in the sentence distinguishes the two cases.

In June, when the pool began, I summed this up to Folha de Pernambuco in a way that still seems correct to me: "AI is not an oracle that predicts the future, it is a tool for making better decisions in environments of uncertainty". The World Cup reinforced the second half of the sentence.

08 · O que 124 inteligências artificiais erraram na Copa

Gabriel

Gabriel scored 629 points and predicted 19 exact scores. He did not receive the dossier we assembled for the machines, did not answer our prompt, and went through no selection process: he joined an online prediction pool, as anyone would. He finished ahead of 121 of the 124 most capable text-processing systems we could bring together.

It is not possible to say why it won. A hundred and four matches do not separate competence from luck, and the hypothesis that it simply got lucky is compatible with everything we measured — as is, for that matter, compatible with my second place. What was recorded is that, among 124 models that read the same material and wrote, for the most part, the same guess, the difference that decided the tournament came from outside that set.

On July 22 we froze the site as a public archive. It is still there: the ranking, the game-by-game guesses, the consensus, the upsets, and the mistakes.

Frequently asked questions

How many AIs took part in the prediction pool, ultimately?

One hundred and twenty-four, in the final count on July 22, 2026 — this is the count consistent with the 5,797 recorded predictions and the phrase "121 of the 124". Material published during the tournament gives lower counts, because models joined and left over the 38 days; the discrepancy is covered in the verification notes.

Did the AIs perform better or worse than humans?

It depends on where you look, and both answers are true. Overall, the AIs scored more: the Série A average was 537 points, and the average across all AIs was 458. At the top, the best human scored 629 points and finished ahead of 121 of the 124. The published human average (55 points, with a median of zero) should not be used for comparison: most people who registered never submitted a prediction.

Why did no AI predict Spain's victory over France?

Because they all received the same context and all converged on the most likely outcome according to the available data, which pointed to an evenly matched game — the consensus for that match was a draw. A language model asked for a single score returns the mode of its distribution, and the mode, by definition, does not cover the unlikely outcome. Of the 62 AIs that responded for that semifinal, none predicted a Spanish victory.

Which model won the prediction pool?

Mistral Small 3 and Grok 4 Fast Reasoning tied on 636 points, with 22 and 20 exact scores. OpenAI o4-mini came third, with 634 points and 22 exact scores. Among the models that submitted predictions through the interface, in Série A, ChatGPT 5 Thinking took first place. A two-point difference over 104 matches does not justify claiming that one model predicts football better than another.

Does the experiment prove that AI is no use for prediction?

No. It proves that, under the conditions of this prediction pool, the models were good at the expected outcome and blind to the unlikely one, and that their errors were aligned. There is literature showing that aggregating several models achieves accuracy comparable to that of human crowds on binary questions with probabilistic answers (arXiv 2402.19379); that is a different question format, and aggregation works in that format. A prediction pool asks for an exact score, and there aggregation pushes everyone toward the same place.

Where are the experiment's data?

The site remains online as an archive, with rankings, match-by-match predictions, consensus, and analysis by model family, at bolao.arenadasias.com.br. The code is available at github.com/giordanorec/bolao-copa-2026. The experiment is run by the FAROL group at CIn-UFPE, is nonprofit, and has no connection to betting companies; API costs were covered by voluntary contributions.