# Forecast protocol (Phase 4) Two arms, run as separate subagents with no shared context. Each arm forecasts each study five times; every run is kept. | | BASELINE | FRAMEWORK | |---|---|---| | Study protocol | `forecasts/{study_id}/protocol.md` | same file | | Required reading | `forecasts/primer_baseline.md` | `forecasts/primer_framework.md` | | Also available | nothing else | `data/claims.csv`, `corpus/` | | Must cite | nothing | `claim_ids_used` from `data/claims.csv` | | May see pistomechanics material | **no** | yes | The two required-reading files are matched in length (within 10% by character count; see `forecasts/primer_lengths.txt`). The FRAMEWORK arm additionally has the claim registry and the corpus on disk. That extra material is the treatment being tested, not a leak. Neither arm may browse the web. A forecaster who searched for the study could find its results, and the pool only admits studies whose results were not public when checked, so a search would either be useless or contaminate the run. Each run is a fresh agent. Runs do not see one another's output. ## Output: one JSON file per run `forecasts/{study_id}/{agent}_{run}.json`, where agent is `baseline` or `framework` and run is 1-5. ```json { "study_id": "NCT00000000", "agent": "baseline", "run": 1, "p_primary_hypothesis_supported": 0.42, "effect_direction": "favours intervention | favours control | null | other: ", "effect_size_90pct_interval": {"low": 0.05, "high": 0.45, "metric": "Cohen's d on "}, "moderator_prediction": " or \"none\"", "claim_ids_used": [], "rationale": "<= 120 words" } ``` - `p_primary_hypothesis_supported` is the probability that the study's own primary hypothesis, as stated in `protocol.md`, will be reported as supported on its primary outcome at the stated alpha. - The interval is in the study's own metric if `protocol.md` names one with a scale; otherwise it is a standardised mean difference (Cohen's d) or an odds ratio, and `metric` says which. The sign convention is positive = in the hypothesised direction. - `claim_ids_used` is empty for BASELINE. For FRAMEWORK it lists every claim_id that the rationale relies on; at least one is required.