Forecasting study · Method
Method, blinding and errors
The study ran in seven phases on 30 September 2026, on one MacBook, with Claude agents doing the reading, the searching and the forecasting. A person set the rules and approved each step that left the machine.
Building the claim register
A crawler copied every public page of pistomechanics.org. Agents then extracted each claim the pages make about belief, kept the exact sentence it came from, and rewrote it in plain terms. A script confirmed that every quoted sentence appears word for word on the page it cites. Each claim was graded for whether a study could test it: 371 can be tested as written, 818 only once a term is given an operational definition, and 169 cannot be tested at all. The untestable claims are listed with what would have to change to make each one testable.
Matching each claim to earlier work
Agents grouped the testable claims into 30 families, such as placebo and expectation, hypnosis and suggestion, or cultural evolution. They built a reading list from web searches for each family, then judged every claim against it and recorded the closest earlier statement. Identical and renamed mean the claim was already published, in the same words or in new ones. Extends means earlier work covers part of it and the claim adds scope, a mechanism or a prediction. The two remaining verdicts are contradicts and no precedent found.
Every DOI was checked against Crossref and every other link was fetched. Some first-pass agents had judged from memory. Every claim marked new or contradicted, 153 in all, therefore went to a second agent, which had to run and log at least two web searches per claim. That second pass kept 90 verdicts and changed 63. The old verdict stays in the register next to the new one.
Finding pending trials
ClinicalTrials.gov was searched for fifteen terms, from placebo and nocebo to psilocybin and targeted memory reactivation. The search kept only trials with no posted results and a primary completion date of July 2026 or later. The forecasters' training data runs to June 2026, so nothing they could have read about these trials includes an outcome. OSF registrations from April 2026 onward and Registered Reports with in-principle acceptance were added.
A study entered the pool when belief, expectation or suggestion was the thing manipulated or the proposed mechanism, and when it named a primary outcome and a comparison. That left 256 studies. Each was checked for published results in its registry and in Europe PMC. Nineteen Registered Reports carry no completion date, and some were accepted in 2022 or 2023, so their data may predate June 2026; they are flagged and were left out of this round.
Choosing and forecasting the ten
Each topic contributed its trials with the earliest primary completion date. Placebo and expectation, psychedelics and hypnosis gave two trials apiece, and nocebo, persuasion, mindset, and sleep and memory one each. Each was searched on the open web for results on the day of forecasting. None had any. The search did turn up an earlier, non-randomised pilot of the spinal cord injury trial's therapy; it stays, as prior knowledge any forecaster could legitimately have.
For each trial, a protocol file was written from the registry record, with one addition fixed in advance: the single outcome and contrast that will decide whether the hypothesis counts as supported. Several trials register more than one primary outcome, and the food-intake trial registers 28, so these choices are judgements. They are published with the forecasts, and scoring will follow them.
Each forecast came from a fresh agent with its own instruction file, run in shuffled order. The agent wrote its answer to a separate folder, so no forecaster could read another's. Tool access cannot be removed from these agents, so the bans on web access and on opening other files were given as instructions and then audited. A script read all 100 transcripts for web calls, sub-agents and any file outside the permitted set, which for the baseline group meant anything from pistomechanics. It found none.
Scoring
A script checks each registry and Europe PMC for results, and will run monthly once the schedule is approved. A trial with a hit goes to a queue, and a person reads the paper and records whether the main hypothesis was supported under the published rule. Forecasts are scored with the Brier score for each group and each trial. Lower is better. Every resolved trial appears in the score table whether the framework forecast it well or badly. A trial that is terminated or withdrawn is marked void and listed, never dropped.
The first run of that script queued 69 of the 256 studies. The one forecast trial among them turned out to have only reviews citing its registration number. The queue needs a filter for reviews and protocol papers before the monthly check runs.
Errors, in the order they happened
- Search cap. The first session reached its limit of 200 web searches while 145 claims were still unchecked. A second session checked them all.
- Citation checker. A bug in the checker made agents replace correct citations with ones it would accept. The bug was fixed and the original sources were restored. Later, 28 DOIs the checker had never confirmed, because Crossref had rate-limited it, were re-queried, and all 28 resolve to the titles cited.
- An email address in a request. One agent put Eron Falbo's email address in the header of a Crossref request. He was told at the time. Every script and prompt since forbids it.
- An unregistered DOI. Lavrič and Rutar's 2026 article in Sociology of Religion has a DOI that Crossref lists as pending. It is cited by its publisher's address and will be checked again.
- A results check that lost its output. The first run over the ClinicalTrials.gov pool saved only at the end and ran for about an hour and ended without saving. It had also lost 13 OSF checks to rate limits. The script now saves after every study and waits when rate-limited.
- OpenAlex skipped. Without an API key, OpenAlex shares a small daily quota among everyone on the same network, and it was used up. A key would have meant opening an account in Eron Falbo's name, so the check ran on the registries and Europe PMC alone, and every affected study is flagged.
- False exclusions. The check treated any reference a sponsor labelled RESULT as a published result. Five sponsors had used the label for background reading published years earlier. All 28 automatic flags were read by hand: 15 were the OSF registration's own record, five were background reading, and the rest were protocol papers, reviews or pilot studies from before the trial.
- Two agents, one folder. Two verification agents shared a scratch folder and one overwrote the other's helper script. Both checked their output. No wrong row reached the register.