We gave an AI swarm four live Polymarket markets, 24 runs and five different language models. It cost $5.96. The swarm gave roughly the same answer to almost every question, and what it answered depended more on how we wrote its input files than on the question itself.
Across 24 runs on four Polymarket markets, MiroFish's swarm medians ranged from 20 to 55 against market prices of 31.5% to 85.5%, and mixing five LLMs changed the median on only one market, by 5 points.
Here is what we did, what we found and every forecast, locked in before the markets resolve.
TL;DR
- Five models, one opinion. A swarm mixing five LLMs gave the same median forecast as a swarm using one model: identical on three markets, 5 points apart on the British Columbia one.
- A default, not a forecast. On three of four markets the swarm landed between 20 and 35, reasoning that "there is no official confirmation yet." In the final runs, at most 2 of 44 agents voted YES on the Saudi market and none at all on the other two.
- Inputs decide the output. In the pilot, a date the persona generator dropped set off a cascade in one run: 61% of agents voted YES, versus 2% in an otherwise identical run.
- The tool anchors itself. Every time MiroFish's generator put a forecast into its own opening posts, it was 30%.
- Scorecard coming. Markets resolve October 16 to 24, 2026. We will publish Brier scores for every source, including where smart money was wrong.
What MiroFish is
MiroFish is an open-source "swarm intelligence" engine. You feed it documents. It extracts the people and organizations in them, turns each into an AI agent with a persona and memory, and lets them post, comment and argue on a simulated Reddit-style feed built on CAMEL-AI's OASIS framework. Afterwards a report agent writes up what happened, and you can interview any agent.
The pitch is that a crowd of simulated stakeholders can rehearse how events unfold. Prediction markets are the natural test: Polymarket already aggregates real traders with real money on the line. If a swarm adds signal, it should show up next to the market price.
How we set it up
Markets. Four open Polymarket markets, picked from OrcaLayer's database: non-sports, resolving one to three weeks out, priced between 15ยข and 85ยข, with at least five smart-money wallets in them.
| Market | YES means |
|---|---|
| Saudi pipeline | Saudi Arabia's government officially announces the East-West pipeline is operational, Sep 23 to Oct 15 |
| Gemini 4 | Google makes Gemini 4 available to the general public by Oct 15 |
| BC Premier | Lorne Doerkson becomes the next Premier of British Columbia |
| Bank of Russia | The Bank of Russia leaves its 14% key rate unchanged on Oct 23 |
Inputs. For each market, a few short "seed" documents built only from public, sourced reporting up to October 4, plus the exact resolution wording from Polymarket.
Two swarm configurations.
- One model: every agent runs on GPT-6 Luna.
- Mixed: each agent is assigned one of five cheap, fast models for the whole run: Gemini 3.8 Flash, GPT-6 Luna, Qwen 3.8 Flash, DeepSeek V4.1 Flash or GLM-5.3 Flash. Reasoning was off or at the lowest level.
Protocol. 25 rounds, Reddit-style feed only, one action per agent per round. Swarm size was set by MiroFish itself from the entities in the seeds: 44, 21, 23 and 15 agents. Three runs per configuration per market, 24 runs in total. After each run every agent was polled with the exact market question and asked for a 0 to 100 probability, then YES or NO, then one sentence of reasoning.
Baselines at the same moment.
- Polymarket: the YES price.
- OrcaLayer Smart Money: the share of smart wallets, and of their invested money, holding YES. This is positioning, not a probability.
- MiroFish's own report agent: the number in MiroFish's written report. It reads the facts graph built from the seeds but not the simulation itself.
Isolation and budget. MiroFish ran in an isolated Docker container with localhost-only ports and no access to OrcaLayer's database. Only public information went in. Budget cap $20, with an automatic stop at $15, later lowered to $12. Actual spend: $5.96, of which $3.84 was the 24-run series and $2.12 the two pilots.
The forecasts, locked in on October 4-5, 2026
| Market | Polymarket | Smart Money, wallets on YES | Smart Money, money on YES | Swarm, one model (0 to 100) | Swarm, mixed (0 to 100) | MiroFish report agent (0 to 100) |
|---|---|---|---|---|---|---|
| Saudi pipeline* | 31.5% | 40.4% | 40.5% | 25 | 25 | 17.5 |
| Gemini 4 public by Oct 15 | 75.5% | 47.2% | 72.5% | 20 | 20 | 62.5 |
| Doerkson next BC Premier | 61.5% | 22.0% | 53.8% | 30 | 35 | 19 |
| Bank of Russia holds | 85.5% | 35.7% | 68.3% | 55 | 55 | 30 |
Swarm figures are the median agent probability, median of three runs. The report agent figure is the median of six reports. Snapshots were taken between 22:41 on October 4 and 00:23 on October 5 (Kyiv time).
*On the Saudi market we corrected three agent personas by hand. See finding 3.
Finding 1: five models, one opinion
The case for a mixed swarm is diversity. Different models have different training data and different blind spots, so mixing them should produce a richer debate.
It didn't. Medians for one model versus five: 25 and 25, 20 and 20, 30 and 35, 55 and 55. Per-model differences inside the mix were small. Qwen was the most cautious on the Saudi and Gemini 4 markets, GPT-6 Luna the most confident on the Bank of Russia market, and all five stayed in the same corridor.
The reason is structural. Every agent's persona is generated from the same seed documents. They share the same facts, the same framing and the same gaps. Swapping the model behind a persona changes the wording, not the conclusion.
Finding 2: "no confirmation yet" is not a forecast
Polymarket priced the first three markets at 31.5%, 75.5% and 61.5%. Very different questions, very different odds. The swarm said 25, 20 and 30 to 35. Across 18 final runs on those three markets, 5 agent votes in total were YES, all on the Saudi market.
The reasoning repeats across topics. A typical answer: "Google has not announced a date, so access by October 15 is possible but unconfirmed." Or: "Polls do not establish that the Conservatives will win." Absence of confirmation becomes 20 to 35, almost regardless of what is being asked.
Inside a single run the agents cluster tightly. In each run, between half and nine-tenths of the answers fell within five points of the most common number. No agent on the first three markets went above 55; on Gemini 4 the ceiling was 35. With that little spread, the swarm's median is essentially the opinion of one model that has read the seed files, repeated 15 to 44 times.
The only market where the swarm leaned YES was the Bank of Russia, and that is the one market where YES means "nothing changes".
Two honest caveats. The swarm only knows its seed files and has no internet access, so if the Gemini 4 market sits at 75.5% because of news or rumors we didn't include, the swarm can't see it. And "nothing has been confirmed" is sometimes the right answer. The scorecard will show which.
Finding 3: the swarm amplifies its inputs
This is the finding we didn't plan for, and the one most worth knowing if you use tools like this.
Pilot 1. Our first version of the Saudi question forgot to specify a date window. A Saudi statement from April 2026, about an earlier outage, technically satisfied it. Same mixed swarm, same run: MiroFish's report agent said 100% in five out of five reports, while only 3 of 45 agents (6.7%) voted YES. The tool's own summary and its own agents told opposite stories.
Pilot 2. We fixed the question. Now 1 of 45 agents (2%) in the one-model swarm voted YES, and 27 of 44 (61%) in the mixed swarm. Same question, same seed files.
The cause. Our seed file dated the April statement correctly. But when MiroFish wrote the agent personas, it dropped the date: the "Saudi Ministry of Energy" persona now described the April recovery as its "subsequent official statement", as if it had followed the September shutdown. In the mixed run, the Ministry agent (on DeepSeek) repeated it in a comment, and the "Saudi Press Agency" agent (on GLM) published it as a fresh official announcement. Twenty-two agents picked it up. At the poll, every one of the 27 YES votes cited the "official statement already issued". In the one-model run, nobody happened to voice that line, so the cascade never started.
The fix. We rewrote that seed with an explicit date in every sentence and a line stating that, as of October 4, no announcement had been made since the September shutdown. The cascade disappeared: both configurations settled at 25. A check still found three personas with undated March and April events next to September ones. We inserted the dates from the seed by hand, kept the originals, and flagged that market in the table above.
The lesson: a swarm does not average out errors in its inputs. It socializes them. Whoever writes the seed files is doing most of the forecasting.
Finding 4: the generator anchors itself
To start a simulation, MiroFish writes a handful of opening posts. On both pilots and on two of the four final markets, those posts contained an explicit forecast nobody asked for. It was the same number every time: 30%. One read "My estimate is 30%", another "a 30% chance that the Bank leaves the rate unchanged".
We removed those lines before the final runs. Left in, every agent starts its debate from a number that came from the tool, not from the question.
What this experiment doesn't show
- Four markets are an illustration, not a statistic. We'll publish the outcomes, but four resolutions can't rank forecasters.
- Smart Money is positioning, not probability. On Doerkson, 22% of smart wallets hold YES, but those wallets account for 53.8% of the smart money invested in the market. Head counts and dollar weights tell different stories, and neither is a probability.
- We tested MiroFish as a forecaster. Exploring how stakeholders might react to a scenario is a different use, and we didn't test it.
- Small swarms. MiroFish found 15 to 44 entities in our seeds, so these are dozens of agents, not thousands.
Who got it right
Coming after resolution: around October 16 for the Saudi pipeline and Gemini 4, October 23 for the Bank of Russia, and after the October 24 election and swearing-in for British Columbia. For each market and source we'll report the outcome, the Brier score and the absolute error, without averaging four points into a fake winner.
The takeaway
A swarm of agents built from the same documents is closer to one opinion said many times than to a crowd. Its answer tracks its inputs: the seed files, the persona generator and even the numbers it writes for itself.
A prediction market is noisy too, but its participants are independent people who read different sources and put money behind what they think. If you want to know what the crowd really believes, watch where the money goes. That's what we track at OrcaLayer: the wallets and the smart-money positioning behind every open Polymarket market.