How to read a clinical trial
Randomised controlled trials are the gold standard of clinical evidence — but misreading them is surprisingly common. This guide walks through every section of a trial paper and the key questions to ask at each step, drawing on CONSORT reporting standards, Cochrane RoB 2, and GRADE methodology.
Clinical trial phases at a glance
Trials are conducted in sequential phases, each answering a different question. The FDA's clinical research framework and the EMA's guidance on clinical trials both outline this progression from first-in-human to post-marketing.
| Phase | N (typical) | Primary question | Success threshold |
|---|---|---|---|
| 0 / FIH | 10–15 | Pharmacokinetics, microdosing | PK meets prediction |
| I | 20–100 | Safety, MTD, dose escalation | Acceptable safety profile |
| II | 100–500 | Preliminary efficacy, dose selection | Signal of activity |
| III (pivotal) | 500–30,000+ | Confirmatory efficacy vs. standard of care | Pre-specified primary endpoint met |
| IV (post-marketing) | Thousands | Long-term safety, new populations, labelling | Surveillance, REMS if required |
For oncology, the FDA Accelerated Approval Programme allows Phase 2 data on surrogate endpoints to support conditional approval pending confirmatory Phase 3 data.
Step 1 — Identify the PICO question
Every trial should answer a well-defined clinical question structured as Population · Intervention · Comparator · Outcome. The PICO framework (NCBI) guides both trial design and critical appraisal.
Who was enrolled? Check inclusion/exclusion criteria — highly selected populations may limit generalisability. The ClinicalTrials.gov registry holds eligibility criteria for every registered trial.
Exact dose, route, duration, co-interventions. Is this the dose used in practice? What was adherence like?
Placebo, active control, or best supportive care? Placebo-controlled trials have the cleanest causal inference but may not reflect real-world choices.
Hard outcomes (death, MI, hospitalisation) vs. surrogate endpoints (LDL, HbA1c, LVEF). Not all surrogates reliably predict patient-important outcomes — see Fleming & DeMets (NCBI).
Step 2 — Assess the study design
The CONSORT 2010 checklist lists 25 items that a well-reported RCT should address. Focus on four design features that most influence internal validity.
Randomisation & allocation concealment
Randomisation (sequence generation) distributes confounders. Allocation concealment (hiding the upcoming assignment) prevents investigators from enrolling patients selectively. Without concealment, even proper randomisation can be subverted. Look for centralised randomisation, sealed envelopes, or IVRS/IWRS systems. The Cochrane RoB 2 tool scores this as a separate domain.
Blinding
Open-label trials carry performance and detection bias. Single-blind (patient unaware) reduces placebo effects. Double-blind (patient + assessor unaware) is the standard for subjective outcomes. Triple-blind (+ statistician) reduces analysis bias. Some outcomes (e.g. all-cause mortality) are objective and less susceptible to unblinding. Check the risk of deblinding in practice.
ITT vs. per-protocol analysis
Intention-to-treat (ITT) analyses all randomised participants as assigned, preserving the benefits of randomisation and providing a conservative estimate of efficacy. Per-protocol (PP) analysis excludes non-adherers and may overestimate efficacy. Modified ITT (mITT) excludes only those with no post-baseline data. For non-inferiority trials, both ITT and PP should confirm NI — if they diverge, be cautious.
Handling missing data
Missing outcome data breaks the randomisation. Methods include last observation carried forward (LOCF — often too optimistic), multiple imputation, and pattern mixture models. The ICH E9(R1) Addendum on Estimands (EMA) requires sponsors to pre-specify how intercurrent events (discontinuation, rescue medication) are handled.
Step 3 — Interpreting statistics
Statistical literacy is essential for avoiding both false positives and false negatives. The NEJM "Statistical Concepts in Clinical Trials" series provides an accessible foundation.
Effect measures
| Measure | Formula | When used | Interpretation |
|---|---|---|---|
| RR | P(event|Tx) / P(event|control) | Binary outcomes, RCTs | RR < 1 = lower risk with treatment |
| OR | odds(Tx) / odds(control) | Case-control, logistic regression | Overestimates RR when event rate > 10% |
| HR | Hazard(Tx) / Hazard(control) | Time-to-event (survival) outcomes | Assumes proportional hazards; check log-log plot |
| ARR | P(control) − P(Tx) | Absolute risk reduction | More clinically meaningful than RRR alone |
| NNT | 1 / ARR | Clinical impact summary | Smaller = greater benefit; always report 95% CI |
| MD / SMD | Mean(Tx) − Mean(control) | Continuous outcomes | SMD (Cohen's d) allows cross-scale comparisons |
P-values and confidence intervals
A statistically significant p-value (<0.05) does not equal clinical significance. Always pair it with the 95% confidence interval to gauge precision and clinical relevance. A very large trial can yield a tiny but statistically significant effect that is clinically meaningless. The 2019 Nature statement signed by 800+ statisticians called for moving beyond binary "significant / not significant" framing.
Power and sample size
Trials are sized to detect a pre-specified minimum clinically important difference (MCID) with 80–90% power at α = 0.05. An underpowered trial that finds no effect cannot conclude safety of equivalence — it may simply lack precision. The CONSORT explanation for sample size reporting (NCBI) explains required elements.
Subgroup analyses
Step 4 — Assess risk of bias
The Cochrane RoB 2 tool evaluates five domains for each outcome:
Sequence generation and allocation concealment. High risk if open allocation or predictable sequence.
Unintended co-interventions, unblinding, or non-adherence that affect the estimate.
Differential dropout between arms; reason for missingness (MCAR vs. MAR vs. MNAR).
Unblinded assessors for subjective outcomes; differential outcome ascertainment.
Outcome switching; reporting only favourable time-points or subgroups. Check against protocol registered on ClinicalTrials.gov.
For a broader evidence synthesis, apply the GRADE framework which rates certainty of evidence as High / Moderate / Low / Very Low based on risk of bias, inconsistency, indirectness, imprecision, and publication bias.
Step 5 — Applicability to your patient
Internal validity (was the trial conducted rigorously?) is distinct from external validity (do the results apply to my patient?). Ask:
- Does my patient resemble trial participants? Trials often exclude elderly patients, those with renal/hepatic impairment, and pregnant women. Special populations may have very different risk-benefit profiles.
- Is the comparator relevant? A trial comparing to placebo or an outdated standard of care may not answer whether the drug is better than current practice.
- What is my patient's baseline risk? RRR is constant across risk levels but ARR (and NNT) varies. A 30% RRR at a 1% baseline event rate yields NNT ≈ 333; at 20% baseline risk, NNT ≈ 17. Tools like the NNT calculator help tailor estimates.
- What are the patient's values and preferences? Shared decision-making tools such as those at Ottawa Hospital Research Institute translate trial data into formats patients can use.
Special trial designs
Adaptive trials
Pre-planned modifications to design or statistical analysis (e.g. sample size re-estimation, arm dropping) based on interim data. Governed by the FDA Adaptive Design Guidance (2019). Can accelerate development but require pre-specified adaptation rules to prevent inflation of Type I error.
Cluster RCTs
Whole units (clinics, wards, schools) are randomised rather than individuals. Common in pragmatic trials of healthcare interventions. Require analysis accounting for intracluster correlation (ICC) — ignoring clustering inflates precision. See BMJ cluster RCT reporting guide.
Crossover trials
Each participant receives both treatments in sequence, serving as their own control. More efficient but susceptible to carryover effects. Wash-out period must exceed ≥5 half-lives of the drug. Appropriate only for stable chronic conditions with reversible endpoints — not for curative treatments.
Non-inferiority trials
Designed to show a new treatment is not unacceptably worse than an active comparator by more than a pre-specified margin (Δ). Common for generics, biosimilars, and safer alternatives. Key pitfall: a generous margin can make an ineffective drug appear non-inferior. Review the margin justification against the FDA NI guidance (2016).
Frequently asked questions
What is the difference between a Phase 2 and Phase 3 trial?
Phase 2 trials (typically 100–500 participants) assess preliminary efficacy and dose-finding, whereas Phase 3 trials (hundreds to thousands of participants) are the pivotal confirmatory studies that form the basis of regulatory submissions. Phase 3 trials are powered to detect a pre-specified effect size with statistical certainty.
What does a p-value of 0.05 actually mean?
A p-value of 0.05 means that, if the null hypothesis were true (no real difference), there is a 5% probability of observing a result at least as extreme as the one obtained by chance alone. It does NOT mean there is a 95% probability that the treatment works, nor that the effect is clinically meaningful.
Why is randomisation important?
Randomisation distributes both known and unknown confounders evenly between treatment arms, so that differences in outcome can be attributed to the intervention rather than baseline imbalances. Without randomisation, observed effects may reflect selection bias rather than treatment efficacy.
What is a number needed to treat (NNT)?
NNT = 1 / absolute risk reduction. It represents how many patients need to receive the treatment instead of control for one additional patient to benefit. An NNT of 20 means treating 20 patients prevents one additional event. Smaller NNTs indicate greater clinical impact.
How do I assess risk of bias in a trial?
Use tools such as Cochrane RoB 2 (for RCTs) which evaluates five domains: randomisation process, deviations from intervention, missing outcome data, outcome measurement, and selection of reported results. Trials with "high risk" in key domains should be interpreted cautiously.
What is the difference between superiority, non-inferiority, and equivalence trials?
Superiority trials test whether a new treatment is better than control. Non-inferiority (NI) trials test whether a new treatment is not worse by more than a pre-defined margin (Δ). Equivalence trials test whether treatments fall within a symmetric margin in both directions. NI and equivalence trials require careful scrutiny of the chosen margin and underlying assumptions.