# Baseline primer: mainstream base rates for forecasting pending clinical and behavioural studies Prepared 2026-09-30. A summary of mainstream evidence on how clinical trials and psychology experiments usually turn out, with meta-analytic effect sizes for thirteen domains. Every figure comes from a cited source; figures that could not be verified were left out. Bracketed numbers point to the references. Part 1 gives the outside view (how often a design succeeds at all), Part 2 the inside view per domain, Part 3 the recurring moderators, Part 4 a forecasting procedure. --- ## Part 1. Base rates for trial and study outcomes ### 1.1 Drug-development phase transitions The largest industry dataset (BIO/Informa/QLS, 2011 to 2020) reports the share of programmes that advance to the next phase [1]. Advancing is not the same as meeting a primary endpoint (programmes also stop for commercial or safety reasons), but it is the best available proxy. | Area | Phase I to II | Phase II to III | Phase III to NDA/BLA | NDA/BLA to approval | Approval from Phase I | |---|---|---|---|---|---| | All indications | 52.0% | 28.9% | 57.8% | 90.6% | 7.9% | | Psychiatry | 52.7% | 26.8% | 56.3% | 91.2% | 7.3% | | Neurology | 47.7% | 26.8% | 53.1% | 86.7% | 5.9% | | Oncology | 48.8% | 24.6% | 47.7% | 92.0% | 5.3% | Phase II is the bottleneck: fewer than three in ten Phase II psychiatric programmes reach Phase III. For psychiatry the chance of approval is 13.8% from Phase II and 51.4% from Phase III [1]. A second analysis (406,038 trial records, 2000 to 2015) counts differently and reports higher levels: 13.8% from Phase I to approval, 48.6% Phase II to III, 59.0% Phase III to approval; 15.0% overall for central nervous system drugs and 3.4% for oncology [2]. Both agree that Phase II is where most programmes stop. ### 1.2 Trial-level success in psychiatry and pain - **Antidepressants, FDA files.** The FDA judged 38 of 74 registered antidepressant trials (51%) positive; in the published literature 94% looked positive, because negative trials went unpublished or were written up as positive. Publication inflated the apparent effect size by 32% [3]. - **Antidepressants over time.** Across 85 FDA-registered trials (1987 to 2013), placebo and drug responses rose together, so the effect size stayed flat (0.30, then 0.29). The share of positive trial arms went from 47.8% to 63.8%, a difference that was not statistically significant [4]. In other words, roughly half of adequately run trials of *approved* antidepressants miss their primary endpoint. - **Scale of the true antidepressant effect.** A network meta-analysis of 522 double-blind trials (116,477 participants) found all 21 antidepressants better than placebo on response [5]. Analyses of FDA data found that the placebo response equalled about 82% of the drug response [6], and a patient-level meta-analysis found that the drug-placebo difference grows with baseline severity and is clinically meaningful mainly in very severe depression [7]. - **Esketamine.** Of three short-term Phase III trials in treatment-resistant depression, only the flexible-dose adult trial (TRANSFORM-2) met its primary endpoint (MADRS difference -4.0 points at day 28, 95% CI -7.31 to -0.64) [8]. The fixed-dose trial and the trial in patients aged 65 and over did not [9, 10]. - **Pain.** Placebo responses in US neuropathic-pain trials grew after 1990, but not in European or Asian trials, and were larger in bigger and longer trials [11]. A systematic review of neuropathic-pain drugs estimated that publication bias overstated treatment effects by about 10%, and reported numbers needed to treat for 50% pain relief of 7.7 for pregabalin and 7.2 for gabapentin [12]. **Takeaway.** For a confirmatory, placebo-controlled trial in psychiatry or pain, even a treatment that genuinely works has something like a coin-flip chance of hitting its primary endpoint in any single trial. A treatment of unknown efficacy starts lower. ### 1.3 Positive-result rates: standard reports versus Registered Reports Scheel et al. compared standard psychology papers with Registered Reports (accepted before results are known): the positive-result rate was 96% in standard reports and 44% in Registered Reports [13]. For a study with a locked, pre-registered analysis plan, a success rate near half is a better prior than the 90%-plus of the journal literature. ### 1.4 Replication rates in psychology and the social sciences | Project | What was replicated | Result | |---|---|---| | Open Science Collaboration 2015 [14] | 100 studies from three psychology journals | 36% of replications significant; mean effect r fell from 0.403 to 0.197 (half) | | Many Labs 1, 2014 [15] | 13 classic effects across many labs | 10 of 13 replicated | | Many Labs 2, 2018 [16] | 28 effects, 60+ labs each, 36 countries | 14 of 28 replicated; median d fell from 0.60 to 0.15; 21 of 28 replication effects smaller than the original | | Many Labs 3, 2016 [17] | 10 effects | 3 of 10 replicated | | Social Sciences Replication Project, 2018 [18] | 21 experiments from *Nature* and *Science*, 2010 to 2015 | 13 of 21 (62%) replicated; replication effects about half the original | | Nosek et al. 2022 review [19] | 307 replications pooled across projects | 64% significant in the same direction; effect sizes 68% as large as the originals | | Tyner et al. 2026 (SCORE) [20] | 274 claims from 164 social and behavioural science papers | 151 of 274 (55.1%) replicated; median r fell from 0.25 to 0.10 | Between a third and two thirds of published findings replicate, and replication effects are typically half the original or smaller. Treat any published original effect size as an overestimate by default. ### 1.5 Effect-size shrinkage from small to large studies - In 85,002 Cochrane forest plots, very large effects came mostly from small trials, and 90% shrank as trials accumulated (median odds ratio of first trials 11.88, pooled 4.20) [21]. - Trials with fewer than 50 patients reported effects 48% larger than trials with 1,000 or more; the smallest quarter 32% larger than the largest [22]. - Pilot effect sizes are imprecise; powering a main trial on them yields underpowered trials [23]. - Comparator choice matters as much as size. In psychotherapy trials for depression, effects against a waiting list (g = 0.95) were about 50% larger than effects against care as usual (g = 0.63) [24]. Rule of thumb: expect a large confirmatory trial to show roughly half the effect of the pilot or first small trial, less still with a more active comparator. ### 1.6 Single-arm pre-post studies - **Regression to the mean.** People enrol when symptoms peak, and scores drift back toward their average on re-measurement [25]. - **Within-group change in control arms is large.** In the second Phase III MDMA trial, the placebo-plus-therapy arm had a within-group effect size of d = 1.25 [26]. A meta-analysis of ketamine and esketamine trials found an overall placebo-arm response of d = -1.85, about 72% of the drug-arm response [27]. Control arms in antidepressant trials improve by a standardised mean change of about 1.0, and in esketamine trials by 1.12 [28]. **Takeaway.** A single-arm study with a symptom outcome will almost always show significant pre-post improvement; this says little about a later controlled trial. The real question is how large the change is and how the authors describe it. ### 1.7 Publication lag and outcome switching - Of NIH-funded trials registered on ClinicalTrials.gov, 46% were published in a peer-reviewed journal within 30 months of completion, and about a third were still unpublished after a median of 51 months [29]. - Among 147 adequately registered trials in high-impact journals, 31% showed discrepancies between the registered and the published primary outcome, and where the effect on significance could be judged, the change usually favoured a significant result [30]. - COMPare checked 67 trials in five top medical journals: 58 (87%) had outcome discrepancies needing a correction letter, primary outcomes were correctly reported in 76% of cases on average, and only 23 letters were published [31]. **Takeaway.** When a question turns on a registered primary outcome, check whether a report could satisfy its wording by shifting to a secondary outcome, subgroup or time point. --- ## Part 2. Domain summaries Effect sizes are standardised mean differences (d, g or SMD) unless stated otherwise. Heterogeneity (I²) is given where the source reports it. ### 2.1 Placebo and open-label placebo - **Conventional placebo versus no treatment.** The Cochrane review found SMD -0.23 (95% CI -0.28 to -0.17), mainly on patient-reported continuous outcomes such as pain and nausea, with wide variation between trials [32]. - **Open-label placebo (OLP), clinical trials.** Von Wernsdorff et al. pooled 11 randomised trials (back pain, cancer-related fatigue, ADHD, allergic rhinitis, depression, IBS, hot flushes): SMD 0.72 (95% CI 0.39 to 1.05), I² = 76% [33]. - **OLP, non-clinical samples.** In 20 experimental studies with healthy or sub-clinical samples, OLP had a moderate effect on self-reported outcomes (SMD 0.43, k = 17) and no effect on objective outcomes (SMD 0.02) [34]. - **OLP, updated synthesis.** Fendel et al. (60 RCTs, 63 comparisons, n = 4,554) report a smaller pooled effect than the early reviews: SMD 0.35 (95% CI 0.26 to 0.44), I² = 53%. Effects were larger in clinical samples (0.47) than non-clinical (0.29), and larger for self-report (0.39) than objective outcomes (0.09). Neither the level of suggestion nor the type of control moderated the effect [35]. - **OLP versus double-blind placebo.** In a 262-patient IBS trial, OLP beat no-pill control (improvement 90.6 vs 52.3 points on the IBS severity scale, d = 0.43) and did not differ from double-blind placebo (d = 0.10) [36]. **Pattern.** As the OLP literature grew, the pooled effect shrank (0.72, then 0.35). A new OLP trial with a self-reported outcome in a clinical sample has a reasonable chance of success; one with an objective primary outcome has a poor chance. ### 2.2 Nocebo - **Experimental pain.** Petersen et al. found nocebo hyperalgesia of moderate to large size. Studies that combined verbal suggestion with conditioning found effects from g = 0.76 to g = 1.17, larger than studies using verbal suggestion alone (g from 0.64 to 0.87 depending on the analysis) [37]. - **Statin side effects.** In the SAMSON n-of-1 trial, 60 patients who had stopped statins for side effects rotated through statin, placebo and no-tablet months. Mean symptom scores were 16.3 on statin, 15.4 on placebo and 8.0 with no tablets; about 90% of the symptom burden on statin was also present on placebo [38]. - **Vaccine trials.** Across 12 COVID-19 vaccine trials, 28.7% to 35.0% of placebo recipients reported systemic adverse events; the nocebo response accounted for 76.0% of systemic adverse events after the first dose and 51.8% after the second [39]. **Pattern.** Nocebo effects on self-reported symptoms are reliable. Studies designed to *produce* nocebo effects in the lab usually succeed. Studies designed to *reduce* nocebo effects (for example, by reframing side-effect information) should be forecast more cautiously. ### 2.3 Verbal suggestion and expectation manipulations - Peerdeman et al. meta-analysed expectation interventions for pain in patients: overall g = 0.61, I² = 73%. Verbal suggestion alone gave g = 0.75 (k = 18), conditioning paired with suggestion g = 0.65 (k = 3), and guided imagery g = 0.27 (k = 6). Effects were medium to large for experimental and acute procedural pain, and small for chronic pain [40]. **Pattern.** Expectation effects are largest for short, self-reported outcomes (lab and procedural pain). Forecast a single-session manipulation aimed at a chronic outcome weeks later as small at best. ### 2.4 Hypnosis - **Experimental pain.** In 85 controlled experimental trials (3,632 participants, mostly crossover), hypnosis produced analgesia across pain outcomes (g = 0.54 to 0.76). Highly suggestible participants who received direct analgesic suggestions showed a 42% clinically meaningful pain reduction, and medium-suggestible participants 29%. Efficacy depended strongly on hypnotic suggestibility [41]. - **Umbrella review.** Across 49 meta-analyses (261 primary studies), effects ranged from d = -0.04 to 2.72; 25.4% were medium and 28.8% large, with the strongest evidence for procedures and pain [42]. - **Invasive medical procedures.** An updated meta-analysis of 20 RCTs (1,250 patients) found reductions in anxiety (SMD -0.43) and pain (SMD -0.35) versus standard care [43]. These effects are smaller than older reviews suggested. - **IBS.** Gut-directed hypnotherapy improved global symptoms and pain (12 studies, 1,158 patients), including in groups [44]; higher-volume interventions did better [45]. - **Anxiety.** In 17 trials, hypnosis reduced anxiety with a mean weighted effect size of 0.79 at the end of treatment, and more when combined with other psychological treatment than when used alone [46]. - **Hypnotisability as a trait.** Hypnotisability is a stable trait; one cohort showed a test-retest correlation of 0.71 over 25 years [47]. **Pattern.** Hypnosis reliably reduces short-term self-reported pain and anxiety; size depends on hypnotisability, direct suggestion and comparator. The most recent procedural meta-analysis gives 0.35 to 0.45, below the 0.54 to 0.79 of earlier syntheses. ### 2.5 Psilocybin for depression - *Psilocybin versus escitalopram* (phase 2, n = 59). The primary outcome (QIDS-SR-16 at 6 weeks) showed a difference of 2.0 points favouring psilocybin, 95% CI -5.0 to 0.9, P = 0.17: not significant. Secondary outcomes favoured psilocybin but were not corrected for multiplicity [48]. - *COMP360, phase 2b* (n = 233, treatment-resistant depression). A single 25 mg dose beat 1 mg on MADRS change at week 3; 10 mg did not. Response at week 3 was 37%, 19% and 18% for 25, 10 and 1 mg; remission was 29%, 9% and 8%. At week 12, sustained response was 20% for 25 mg against 10% for 1 mg [49, 50]. - *Usona phase 2* (n = 104, major depression, niacin active placebo). Psilocybin 25 mg reduced MADRS by 12.3 points more than niacin at day 43 (95% CI -17.5 to -7.2) [51]. - *COMP005, phase 3* (single 25 mg dose versus placebo). MADRS difference of -3.6 points at week 6 (95% CI -5.7 to -1.5) [52]. - *COMP006, phase 3* (581 dosed; 25 mg, 10 mg and 1 mg; two doses three weeks apart). Met its primary endpoint with a MADRS difference of -3.8 points versus 1 mg at week 6; 39% of the 25 mg arm had at least a 25% MADRS reduction. The sponsor reports that benefit persisted to week 26, and a rolling NDA submission is under way, with completion expected in Q4 2026 [53, 54]. The pattern is shrinkage: a 12.3-point difference in the Usona phase 2 trial, about 3.6 to 3.8 points in phase 3. **Functional unblinding and expectancy.** Participants can usually tell whether they received a psychedelic. Effect sizes in psychedelic RCTs are probably inflated by de-blinding and expectancy, which most trials neither measure nor report [55]. Three findings: - Control arms in psilocybin trials improve much less than control arms in other antidepressant trials (standardised mean change 0.50, against 1.00 for SSRI trials and 1.12 for esketamine trials; placebo response rates 19%, 33% and 42%). The between-group effect for psilocybin (0.70) was about 2.3 times that of SSRIs (0.27) or esketamine (0.30), mostly because its control arms did worse [28]. - When psilocybin-assisted therapy was compared with *open-label* conventional antidepressants, both effectively unblinded, the difference was 0.3 HAM-D points (95% CI -1.39 to 1.98), not significant [56]. - In the escitalopram comparison trial, pre-treatment expectancy predicted response to escitalopram (r = -0.49) but not to psilocybin, while baseline suggestibility was associated with response to psilocybin only [57]. **Moderators and fade.** Dose (25 mg beats 10 mg), treatment-resistance, the strength of the control condition, and whether participants have prior psychedelic experience or strong positive expectations. Single-dose effects fade over weeks to months for many patients; the phase 2b sustained-response rate at week 12 was 20% [50]. The sponsor reports better durability with re-dosing in COMP006 [54]. ### 2.6 MDMA-assisted therapy for PTSD - **Phase 3, MAPP1** (n = 90, severe PTSD). CAPS-5 change -24.4 with MDMA against -13.9 with placebo plus the same therapy; between-group d = 0.91 [58]. - **Phase 3, MAPP2** (moderate to severe PTSD). Least-squares mean change -23.7 against -14.8; d = 0.7. Within-group d was 1.95 for MDMA and 1.25 for placebo plus therapy [26]. - **Blinding.** In the FDA briefing, 95.7% of MDMA recipients and 84.1% of placebo recipients in MAPP1 correctly guessed their arm; in MAPP2 the figures were 94.2% and 75% [59]. - **Regulatory outcome.** On 4 June 2024 the FDA advisory committee voted 2 to 9 on effectiveness and 1 to 10 on benefit outweighing risk. On 9 August 2024 the FDA issued a complete response letter asking for an additional Phase III trial [60, 61]. The published CRL raised durability, abuse-related adverse events, and prior MDMA use among participants. Around the same time, *Psychopharmacology* retracted three papers from the earlier phase 2 programme over unethical conduct at one site [62]. **Pattern.** Two positive phase 3 trials did not produce approval. A positive topline and a favourable regulatory decision are separate events. ### 2.7 Ketamine and esketamine - **Speed and fade.** In a meta-analysis of eight placebo-controlled RCTs (n = 183), ketamine produced remission odds ratios of 7.06 at 24 hours, 3.86 at day 3 and 4.00 at day 7 [63]. Another meta-analysis of 21 studies found larger effects for repeated infusions than for single ones at 4 hours, 24 hours and 7 days [64]. Single-infusion effects fade within about a week. - **Placebo response.** Across 14 double-blind RCTs (n = 1,100), the placebo-arm response was d = -1.85, about 72% of the active-arm response [27]. - **Against ECT.** In a non-inferiority trial in non-psychotic treatment-resistant depression, response was 55.4% with ketamine against 41.2% with ECT (difference 14.2 points, 95% CI 3.9 to 24.2) [65]. - **Esketamine.** One of three short-term phase 3 trials met its primary endpoint (see 1.2) [8, 9, 10]. **Pattern.** Ketamine reliably shows short-term separation from saline placebo; blinding is weak because of dissociative effects, and effects fade quickly without repeat dosing. ### 2.8 Growth-mindset interventions - **National study.** A preregistered, nationally representative US trial (12,490 ninth-graders in 65 schools, two 25-minute online sessions) raised core-course grade-point average among lower-achieving students by 0.10 grade points; effects were larger where peer norms supported the message [66]. - **Meta-analyses.** Sisk et al.: d = 0.08 on achievement (29 studies, N = 57,155); where a manipulation check existed, the effect was significant only when the check failed [67]. Macnamara and Burgoyne found d = 0.05 across 63 studies (N = 97,672), not significant after correcting for publication bias; d = 0.04 in studies where the intervention demonstrably changed mindsets; and d = 0.02 (95% CI -0.06 to 0.10) in the six highest-quality studies [68]. Burnette et al. found d = 0.14 in targeted, high-fidelity samples (95% prediction interval -0.08 to 0.35) [69]. **Pattern.** Average effects are near zero; positives concentrate in lower achievers and supportive schools. Forecast a general-population trial as likely null on achievement. ### 2.9 Persuasion, misinformation and inoculation - **Political persuasion.** In 59 randomised experiments (34,000 people, 49 ads from the 2016 US presidential campaign), persuasive effects were small whatever the ad's tone, timing or audience [70]. A meta-analysis of 49 field experiments found that the best estimate of the effect of campaign contact on candidate choice in US general elections is zero [71]. Where advertising effects exist, most decay within about two weeks [72]. - **Debunking.** A meta-analysis of 52 experiments (N = 6,878) found large effects of presenting misinformation (d = 2.41 to 3.08) and of debunking (d = 1.14 to 1.33), and a substantial persistence of misinformation after debunking (d = 0.75 to 1.06). Debunking worked better when audiences generated counter-arguments themselves [73]. - **Inoculation (prebunking).** A meta-analysis of 42 studies (N = 42,530) found that inoculation lowered the rated credibility of misinformation (d = -0.36), slightly raised the credibility of true information (d = 0.20), and improved credibility discernment (d = 0.20) and sharing discernment (d = 0.18) [74]. Short inoculation videos improved recognition of manipulation techniques in six lab studies and a YouTube field study (n = 22,632) [75]. Effects decay over weeks; a memory-targeted booster one week later extended them to about a month [76]. - **Accuracy prompts.** An internal meta-analysis of 20 experiments (N = 26,863) found that prompting people to consider accuracy reduced intentions to share false headlines by about 10% relative to control [77]. - **Dialogue with an AI model.** A 2024 experiment with 2,190 participants who believed in conspiracy theories reported a roughly 20% reduction in belief after a personalised dialogue with GPT-4 Turbo, lasting two months [78]. *Science* later published an editorial expression of concern about reporting inconsistencies and the public dataset [79]. **Pattern.** Survey outcomes measured right after exposure, on items the intervention addressed, reliably move; behavioural, real-world and delayed outcomes move much less. Discernment effects are about a fifth of a standard deviation. ### 2.10 Self-affirmation - **Health behaviour.** Self-affirmation increased acceptance of health messages (d = 0.32), intentions (d = 0.14) and behaviour (d = 0.32) [80]. - **Educational achievement.** A well-powered replication in a setting where an earlier large field experiment had found lasting benefits found no effect, precise enough to rule out benefits larger than 0.10 standard deviations [81]. **Pattern.** Small effects on intentions and message acceptance; inconsistent effects on long-term achievement, with large replications often null. ### 2.11 Targeted memory reactivation (TMR) - Hu et al. meta-analysed 91 experiments (212 effect sizes, N = 2,004). Re-presenting learning cues during sleep improved later memory, g = 0.29 (95% CI 0.21 to 0.38). Cueing worked in NREM stage 2 (g = 0.32) and slow-wave sleep (g = 0.27), but not in REM sleep or during wakefulness [82]. **Pattern.** A small, fairly reliable lab effect with cueing in deep sleep. Forecast home-based cueing, REM cueing or non-memory outcomes as smaller and less certain. ### 2.12 Sleep and memory - Berres and Erdfelder meta-analysed 823 effect sizes from 271 samples. Sleep after learning improved episodic memory relative to wakefulness, g = 0.44; after adjustment for selective reporting, g = 0.28. Effects were larger for free recall (g = 0.49) than recognition (g = 0.38), and larger for natural night sleep than for daytime naps or sleep-deprivation designs [83]. - A review of replication attempts concluded that sleep benefits for memory are smaller, more task-dependent, less tied to slow-wave sleep, less robust and less long-lasting than earlier work suggested [84]. ### 2.13 Memory reconsolidation interventions - **Behavioural (retrieval-extinction).** Schiller et al. reported in 2010 that retrieving a fear memory shortly before extinction training prevented the return of fear [85]. A high-powered, pre-registered direct replication of the critical conditions found no benefit of retrieval-extinction over standard extinction; the authors also found that the original paper had misreported its exclusion criteria [86]. - **Pharmacological (propranolol).** Kindt et al. reported in 2009 that propranolol given with memory reactivation erased the fear response in humans [87]. A 2022 meta-analysis found reduced recall of aversive material in healthy adults (g = -0.51, 14 studies, n = 478) and reduced symptoms in PTSD, addiction or phobia (g = -0.42, 12 studies, n = 446) [88]. A letter re-analysing the same evidence argued that, once errors were corrected, there was no effect of propranolol versus placebo on traumatic memory reconsolidation [89]. **Pattern.** Striking early single-lab findings, followed by failed or contested replications. Forecast new reconsolidation trials with a low prior, especially where the claim is durable prevention of relapse rather than a short-term difference. --- ## Part 3. Recurring moderators and why effects fade | Moderator | What the evidence shows | Where | |---|---|---| | Baseline severity | Drug-placebo differences in depression grow with severity; small in mild to moderate cases | [6, 7] | | Comparator strength | Waiting-list controls inflate effects by about 50% relative to usual care; weak placebo-arm responses inflate psilocybin effects | [24, 28] | | Blinding integrity | Most psychedelic and MDMA participants correctly guess their arm; open-label comparisons remove most of the psilocybin advantage | [55, 56, 59] | | Expectancy | Predicts response to conventional antidepressants; baseline suggestibility predicts response to psilocybin | [57] | | Hypnotisability and suggestibility | Hypnotic analgesia is much larger in high- than low-suggestible people | [41, 47] | | Outcome type | Placebo and OLP effects appear on self-report and largely vanish on objective measures | [32, 34, 35] | | Chronic versus acute outcomes | Expectation effects large for lab and procedural pain, small for chronic pain | [40] | | Subgroup and context | Mindset effects appear in lower achievers and supportive school contexts only | [66, 69] | | Dose and repetition | Higher psilocybin dose and repeat ketamine infusions give larger effects | [49, 53, 64] | | Trial location and size | US pain-trial placebo responses grew over time and with size | [11] | **Why effects fade at follow-up.** (1) Regression to the mean inflates early change scores [25]. (2) Expectancy and novelty effects are strongest right after the intervention, and single-dose drug effects wear off (ketamine within about a week; single-dose psilocybin over weeks to months) [50, 63]. (3) Memory for a one-off message decays; persuasion and inoculation effects fade over days to weeks without reinforcement [72, 76]. (4) Dropout at follow-up is rarely random, so analyses of completers look better than intention-to-treat. (5) Follow-up analyses are often secondary and less protected against selective reporting than primary endpoints [30, 31]. --- ## Part 4. How to turn this into a forecast The adjustments below are judgement calls anchored on the cited figures, not outputs of a fitted model. **Step 1. Classify the study.** Design (single-arm, waitlist-controlled, active or usual-care comparator, double-blind placebo); phase (pilot, phase 2, phase 3, replication); sample size; outcome type (self-report, clinician-rated, objective); time point; pre-registration; and exactly what counts as success in the question. **Step 2. Pick a starting base rate.** - *Single-arm pre-post with a symptom outcome:* the chance of a significant improvement is very high (often above 90%); regression to the mean and placebo response alone usually produce it [25, 27, 28]. - *Randomised against waitlist or no treatment, self-reported outcome:* high, because the comparator is weak [24]. - *Double-blind placebo-controlled confirmatory trial in psychiatry or pain, for a treatment with good earlier evidence:* about a coin flip. Around half of trials of approved antidepressants are positive [3, 4], and one in three esketamine phase 3 trials hit its primary endpoint [8, 9, 10]. - *Phase 2 programme in psychiatry or neurology moving on to Phase 3:* about 27% [1]. - *Pre-registered psychology experiment or Registered Report:* about 44% positive [13]. - *Direct replication of a published psychology finding:* about 35% to 65% significant, depending on the field [14, 16, 19, 20]. **Step 3. Adjust the expected effect size before judging power.** Halve any effect size taken from a pilot, a first small trial or a single original study [14, 16, 21, 22, 23]. Use the domain meta-analysis in Part 2 rather than the most-cited single study. Where later syntheses have shrunk (OLP 0.72 to 0.35; psilocybin 12.3 points in phase 2 to 3.6 to 3.8 in phase 3; hypnosis for procedures around 0.35 to 0.45), use the later figure. **Step 4. Check power against the adjusted effect.** For a two-arm trial, detecting d = 0.3 with 80% power needs roughly 175 participants per arm, and detecting d = 0.2 needs roughly 390 per arm (standard two-sample calculation). If the realistic effect needs a sample the study does not have, lower the probability of a significant primary result sharply. **Step 5. Adjust for design features.** - *Blinding:* if participants can tell their arm (psychedelics, MDMA, ketamine, hypnosis versus nothing), the observed effect will be larger than the pharmacological or specific effect; this raises the chance of a positive primary result but lowers the chance that the result survives regulatory or replication scrutiny [55, 56, 59, 60]. - *Comparator:* the more the comparator resembles the intervention (active placebo, usual care, an established treatment), the smaller the expected difference [24, 28, 56]. - *Outcome type:* self-report outcomes favour positive results; objective outcomes lower them, sometimes to near zero for expectation-based interventions [34, 35]. - *Chronic conditions and long follow-up:* lower the expected effect [40, 50, 72, 76]. - *Subgroups:* if the question is about the whole sample but the evidence is positive only for a subgroup (low achievers, high suggestibles, severe patients), forecast the whole-sample result as weaker [7, 41, 66]. **Step 6. Separate the events in the question.** "Meets the primary endpoint", "authors describe the result as positive", "is published by a given date", "gains regulatory approval" and "replicates" are different events with different probabilities. Authors describe results as positive far more often than primary endpoints are met [3, 13, 31]; publication often lags completion by years [29]; approval can fail after positive phase 3 trials (MDMA) or succeed after mixed ones (esketamine) [8, 60]. **Step 7. Sanity-check.** If the forecast implies an effect above the best meta-analytic estimate, or a success probability above the base rate for its design, name the specific reason (stronger dose, targeted subgroup, weaker comparator) or move back toward the base rate. **Quick reference: expected effect sizes from recent syntheses.** | Intervention and outcome | Pooled estimate | Source | |---|---|---| | Placebo versus no treatment, all conditions | SMD -0.23 | [32] | | Open-label placebo, all RCTs | SMD 0.35 (self-report 0.39; objective 0.09) | [35] | | Expectation interventions for pain | g = 0.61 (verbal suggestion 0.75; chronic pain small) | [40] | | Nocebo hyperalgesia, suggestion plus conditioning | g = 0.76 to 1.17 | [37] | | Hypnosis for invasive procedures | anxiety SMD -0.43; pain SMD -0.35 | [43] | | Psilocybin phase 3, MADRS difference at week 6 | -3.6 to -3.8 points | [52, 53] | | MDMA-assisted therapy, phase 3 | d = 0.7 to 0.91 (functionally unblinded) | [26, 58] | | Conventional antidepressants versus placebo | about 0.3 | [4, 28] | | Growth mindset, highest-quality studies | d = 0.02 | [68] | | Inoculation, sharing discernment | d = 0.18 | [74] | | Targeted memory reactivation | g = 0.29 | [82] | | Sleep versus wake, episodic memory, bias-adjusted | g = 0.28 | [83] | | Propranolol reconsolidation, clinical | g = -0.42 (contested) | [88, 89] | --- ## References Short form: first author, year, journal, DOI or stable URL. 1. BIO / Informa Pharma Intelligence / QLS Advisors (2021). Clinical Development Success Rates 2011-2020. https://go.bio.org/rs/490-EHZ-999/images/ClinicalDevelopmentSuccessRates2011_2020.pdf 2. Wong CH et al. (2019). Biostatistics. doi:10.1093/biostatistics/kxx069 3. Turner EH et al. (2008). N Engl J Med. doi:10.1056/NEJMsa065779 4. Khan A et al. (2017). World Psychiatry. doi:10.1002/wps.20421 5. Cipriani A et al. (2018). Lancet. doi:10.1016/S0140-6736(17)32802-7 6. Kirsch I et al. (2008). PLoS Med. doi:10.1371/journal.pmed.0050045 7. Fournier JC et al. (2010). JAMA. doi:10.1001/jama.2009.1943 8. Popova V et al. (2019). Am J Psychiatry. doi:10.1176/appi.ajp.2019.19020172 9. AJMC report on TRANSFORM-2, noting TRANSFORM-1 missed its endpoint. https://www.ajmc.com/view/transform2-results-on-esketamine-published-along-with-caveats 10. Ochs-Ross R et al. (2020). Am J Geriatr Psychiatry. doi:10.1016/j.jagp.2019.10.008 11. Tuttle AH et al. (2015). Pain. doi:10.1097/j.pain.0000000000000333 12. Finnerup NB et al. (2015). Lancet Neurol. doi:10.1016/S1474-4422(14)70251-0 13. Scheel AM et al. (2021). Adv Methods Pract Psychol Sci. doi:10.1177/25152459211007467 14. Open Science Collaboration (2015). Science. doi:10.1126/science.aac4716 15. Klein RA et al. (2014). Soc Psychol. doi:10.1027/1864-9335/a000178 16. Klein RA et al. (2018). Adv Methods Pract Psychol Sci. doi:10.1177/2515245918810225 17. Ebersole CR et al. (2016). J Exp Soc Psychol. doi:10.1016/j.jesp.2015.10.012 18. Camerer CF et al. (2018). Nat Hum Behav. doi:10.1038/s41562-018-0399-z 19. Nosek BA et al. (2022). Annu Rev Psychol. doi:10.1146/annurev-psych-020821-114157 20. Tyner AH et al. (2026). Nature. doi:10.1038/s41586-025-10078-y 21. Pereira TV et al. (2012). JAMA. doi:10.1001/jama.2012.13444 22. Dechartres A et al. (2013). BMJ. doi:10.1136/bmj.f2304 23. Kraemer HC et al. (2006). Arch Gen Psychiatry. doi:10.1001/archpsyc.63.5.484 24. Cuijpers P et al. (2024). Epidemiol Psychiatr Sci. doi:10.1017/S2045796024000611 25. Barnett AG et al. (2005). Int J Epidemiol. doi:10.1093/ije/dyh299 26. Mitchell JM et al. (2023). Nat Med. doi:10.1038/s41591-023-02565-4 27. Matsingos A et al. (2024). Front Psychiatry. doi:10.3389/fpsyt.2024.1346697 28. Hieronymus F et al. (2025). JAMA Netw Open. doi:10.1001/jamanetworkopen.2025.24119 29. Ross JS et al. (2012). BMJ. doi:10.1136/bmj.d7292 30. Mathieu S et al. (2009). JAMA. doi:10.1001/jama.2009.1242 31. Goldacre B et al. (2019). Trials. doi:10.1186/s13063-019-3173-2 32. Hrobjartsson A & Gotzsche PC (2010). Cochrane Database Syst Rev. doi:10.1002/14651858.CD003974.pub3 33. von Wernsdorff M et al. (2021). Sci Rep. doi:10.1038/s41598-021-83148-6 34. Spille L et al. (2023). Sci Rep. doi:10.1038/s41598-023-30362-z 35. Fendel JC et al. (2025). Sci Rep. doi:10.1038/s41598-025-14895-z 36. Lembo A et al. (2021). Pain. doi:10.1097/j.pain.0000000000002234 37. Petersen GL et al. (2014). Pain. doi:10.1016/j.pain.2014.04.016 38. Wood FA et al. (2020). N Engl J Med. doi:10.1056/NEJMc2031173 39. Haas JW et al. (2022). JAMA Netw Open. doi:10.1001/jamanetworkopen.2021.43955 40. Peerdeman KJ et al. (2016). Pain. doi:10.1097/j.pain.0000000000000540 41. Thompson T et al. (2019). Neurosci Biobehav Rev. doi:10.1016/j.neubiorev.2019.02.013 42. Rosendahl J et al. (2024). Front Psychol. doi:10.3389/fpsyg.2023.1330238 43. Walter N et al. (2025). J Psychosom Res. doi:10.1016/j.jpsychores.2025.112117 44. Adler EC et al. (2025). Neurogastroenterol Motil. doi:10.1111/nmo.70037 45. Krouwel M et al. (2021). Complement Ther Med. doi:10.1016/j.ctim.2021.102672 46. Valentine KE et al. (2019). Int J Clin Exp Hypn. doi:10.1080/00207144.2019.1613863 47. Piccione C et al. (1989). J Pers Soc Psychol. doi:10.1037/0022-3514.56.2.289 48. Carhart-Harris R et al. (2021). N Engl J Med. doi:10.1056/NEJMoa2032994 49. Goodwin GM et al. (2022). N Engl J Med. doi:10.1056/NEJMoa2206443 50. Review, psilocybin for treatment-resistant depression: an update (COMP360 phase 2b response, remission, week 12). https://pmc.ncbi.nlm.nih.gov/articles/PMC10801413/ 51. Raison CL et al. (2023). JAMA. doi:10.1001/jama.2023.14530 52. COMP005 phase 3 topline, HCPLive. https://www.hcplive.com/view/compass-pathways-comp360-psilocybin-shows-benefit-in-phase-3-trd-trial 53. COMP006 phase 3 topline, HCPLive. https://www.hcplive.com/view/comp360-psilocybin-meets-primary-endpoint-second-phase-3-trial-trd 54. Compass Pathways press release, 7 July 2026, COMP006 six-month data. https://ir.compasspathways.com/News--Events-/news/news-details/2026/Compass-Pathways-Announces-Six-Month-Data-from-Second-Phase-3-Trial-Confirming-Rapid-and-Durable-Profile/default.aspx 55. Muthukumaraswamy SD et al. (2021). Expert Rev Clin Pharmacol. doi:10.1080/17512433.2021.1933434 56. Williams ZJ et al. (2026). JAMA Psychiatry. doi:10.1001/jamapsychiatry.2025.4809 57. Szigeti B et al. (2024). Psychol Med. doi:10.1017/S0033291723003653 58. Mitchell JM et al. (2021). Nat Med. doi:10.1038/s41591-021-01336-3 59. FDA (2024). Briefing document, NDA 215455 midomafetamine. https://www.fda.gov/media/178984/download 60. FDA advisory committee votes on MDMA for PTSD, HCPLive (2024). https://www.hcplive.com/view/live-updates-fda-psychopharmacologic-advisory-committee-meeting-mdma-ptsd 61. Lykos Therapeutics (2024). Complete response letter press release. https://www.prnewswire.com/news-releases/lykos-therapeutics-announces-complete-response-letter-for-midomafetamine-capsules-for-ptsd-302219182.html ; CRL contents: https://www.hcplive.com/view/fda-releases-crl-detailing-safety-concerns-mdma-assisted-therapy-ptsd 62. BioPharma Dive (2024). Journal retracts MDMA-assisted therapy papers. https://www.biopharmadive.com/news/lykos-journal-mdma-retractions-maps-ptsd-unethical-conduct/723945/ 63. McGirr A et al. (2015). Psychol Med. doi:10.1017/S0033291714001603 64. Coyle CM & Laws KR (2015). Hum Psychopharmacol. doi:10.1002/hup.2475 65. Anand A et al. (2023). N Engl J Med. doi:10.1056/NEJMoa2302399 66. Yeager DS et al. (2019). Nature. doi:10.1038/s41586-019-1466-y 67. Sisk VF et al. (2018). Psychol Sci. doi:10.1177/0956797617739704 68. Macnamara BN & Burgoyne AP (2023). Psychol Bull. doi:10.1037/bul0000352 69. Burnette JL et al. (2023). Psychol Bull. doi:10.1037/bul0000368 70. Coppock A et al. (2020). Sci Adv. doi:10.1126/sciadv.abc4046 71. Kalla JL & Broockman DE (2018). Am Polit Sci Rev. doi:10.1017/S0003055417000363 72. Hill SJ et al. (2013). Polit Commun. doi:10.1080/10584609.2013.828143 73. Chan MS et al. (2017). Psychol Sci. doi:10.1177/0956797617714579 74. Lu C et al. (2023). J Med Internet Res. doi:10.2196/49255 75. Roozenbeek J et al. (2022). Sci Adv. doi:10.1126/sciadv.abo6254 76. Maertens R et al. (2025). Nat Commun. doi:10.1038/s41467-025-57205-x 77. Pennycook G & Rand DG (2022). Nat Commun. doi:10.1038/s41467-022-30073-5 78. Costello TH et al. (2024). Science. doi:10.1126/science.adq1814 79. Science. Editorial expression of concern on Costello et al. 2024. doi:10.1126/science.aej2383 80. Epton T et al. (2015). Health Psychol. doi:10.1037/hea0000116 81. Hanselman P et al. (2017). J Educ Psychol. doi:10.1037/edu0000141 82. Hu X et al. (2020). Psychol Bull. doi:10.1037/bul0000223 83. Berres S & Erdfelder E (2021). Psychol Bull. doi:10.1037/bul0000350 84. Cordi MJ & Rasch B (2021). Curr Opin Neurobiol. doi:10.1016/j.conb.2020.06.002 85. Schiller D et al. (2010). Nature. doi:10.1038/nature08637 86. Chalkia A et al. (2020). Cortex. doi:10.1016/j.cortex.2020.04.017 87. Kindt M et al. (2009). Nat Neurosci. doi:10.1038/nn.2271 88. Pigeon S et al. (2022). J Psychiatry Neurosci. doi:10.1503/jpn.210057 89. Steenen SA et al. (2022). J Psychiatry Neurosci. doi:10.1503/jpn.220072-l