TL;DR — Not all evidence is equal. Sort study designs into a pyramid: at the top, systematic reviews / meta-analyses that pool many studies; below them, the RCT (randomize people to treatment vs control, follow forward — the strongest proof of cause); then observational designs — cohort (split by exposure, follow forward), case-control (start from the outcome, look backward, best for rare diseases), and cross-sectional (a snapshot, prevalence only); at the bottom, case series, case reports, and expert opinion (no comparison group, weakest). The design decides three downstream things: which effect measure is valid (exposure-sampled designs support RR/RD/NNT; outcome-sampled case-control supports only the odds ratio), where you start on GRADE (RCT starts high, observational starts low, because randomizing removes confounding), and which risk-of-bias tool you reach for (RoB2 for RCTs, ROBINS-I for non-randomized, AMSTAR-2 for systematic reviews). Learn to read a study's design off its abstract and the rest follows.
1. Simple explanation
Evidence comes in grades. If a friend says "this pill cured my headache," that is one person's story — interesting, but it could be luck, the pill, or the fact that headaches fade on their own. If instead a thousand people were split by coin flip into "pill" and "sugar tablet," followed for a month, and the pill group did clearly better, that is far harder to dismiss. Same claim, wildly different strength. The evidence hierarchy is just a way of ranking how much we should trust a result based on how the study was built.
Analogy — a court of law. Think of building a legal case.
- Expert opinion is a character witness saying "I think he's guilty." No evidence, just judgment.
- A case report is one eyewitness: "I saw one thing happen once." Vivid, but a single account.
- A case series is several eyewitnesses who all saw the defendant — but nobody watched anyone else, so there is no comparison.
- A cross-sectional survey is a snapshot of the whole town on one day: who was where, right now.
- A case-control study starts from the crime (the outcome) and works backward to who was near the scene.
- A cohort study follows suspects forward in time to see who commits the crime.
- An RCT is the gold standard: you could randomly assign who gets exposed, so the two groups are otherwise identical — the cleanest possible proof of cause.
- A systematic review is the appeals court reading every trial transcript and weighing them together.
The higher you climb, the more the design protects you from being fooled by luck, bias, and hidden differences between groups. This article teaches you to name a design on sight and to know what that name lets you do next.
The three questions it answers: What is each design and when is it used? Which way does time run — forward or backward? And what does the design let you compute, claim, and score afterward?
2. Diagram
THE EVIDENCE PYRAMID
(top = strongest cause proof, bottom = weakest)
╱╲
╱ ╲ SYSTEMATIC REVIEW / META-ANALYSIS
╱ SR ╲ pool many studies into one estimate
╱──────╲
╱ ╲ RANDOMIZED CONTROLLED TRIAL (RCT)
╱ RCT ╲ coin-flip to treat vs control → forward
╱────────────╲
╱ ╲ COHORT
╱ cohort ╲ split by EXPOSURE → follow forward
╱──────────────────╲
╱ ╲ CASE-CONTROL
╱ case-control ╲ start from OUTCOME → look backward
╱────────────────────────╲
╱ ╲ CROSS-SECTIONAL
╱ cross-sectional ╲ one snapshot in time (prevalence)
╱──────────────────────────────╲
╱ case series ╲ a group treated, NO control
╱──────────────────────────────────╲
╱ case report ╲ a single patient
╱──────────────────────────────────────╲
╱ expert opinion ╲ no data, just judgment
╱────────────────────────────────────────────╲
TIME DIRECTION ANALYTIC vs DESCRIPTIVE
───────────── ────────────────────────
forward → RCT, cohort has a comparison group:
(cause → effect, prospective) RCT, cohort, case-control
backward → case-control no comparison group:
(effect → cause, look back) case series, case report
3. How it works
3.1 The designs, one at a time
| Design | What you do | Time direction | Sampled by | Best for |
|---|---|---|---|---|
| RCT | Randomize people to treatment vs control, follow forward | Forward (prospective) | Exposure (assigned) | Proving a treatment causes an effect |
| Cohort | Observe a group split by exposure, follow forward to outcome | Forward | Exposure | Harms/prognosis; when randomizing is impossible |
| Case-control | Start from cases (have disease) + controls (don't), look back at exposure | Backward | Outcome | Rare diseases; cheap, fast |
| Cross-sectional | Measure exposure + outcome at one moment | None (snapshot) | The population now | Prevalence, associations |
| Case series | Describe a group who got a treatment, no control | Usually forward | — | Hypothesis-generating |
| Case report | Describe a single patient | — | — | Flagging something new/rare |
| Expert opinion | Judgment, no collected data | — | — | Filling gaps when nothing else exists |
RCT (randomized controlled trial). The researcher decides who is exposed by a random draw — a coin flip. Because the assignment is random, the two groups are, on average, identical in everything else (age, habits, hidden illnesses). So if outcomes differ, the treatment is the reason. Example: 200 patients with high blood pressure are coin-flipped to a new BP pill or a placebo, then followed forward for a year to compare heart attacks.
Cohort. You observe rather than assign. You find people already split by an exposure and follow them forward. Prospective cohorts follow in real time; retrospective cohorts reconstruct the follow-up from old records. Example: follow 1,000 smokers and 1,000 non-smokers for 10 years and count lung cancers. You couldn't ethically assign people to smoke, so you observe.
Case-control. You start from the outcome. Gather people who already have the disease (cases) and a comparable group who don't (controls), then look backward at their past exposures. This is the trick for rare diseases — you don't have to follow a million people hoping a few get a rare cancer; you just start with the ones who have it. Example: 100 patients with a rare cancer + 100 without, asking each about past asbestos exposure.
Cross-sectional. A single snapshot. You measure exposure and outcome at the same moment, so you get prevalence ("how many have it right now") and associations — but no sense of what came first. Example: survey a city on one day for both obesity and knee pain; you'll see they travel together, but not which caused which.
Case series / case report / expert opinion. These have no comparison group. A case series describes several patients who got a treatment; a case report describes one; expert opinion is judgment with no collected data. They can raise a hypothesis ("five patients on this drug all developed the same rash") but can never prove cause, because there is nothing to compare against.
3.2 Forward vs backward — the arrow of time
This single distinction organizes half the pyramid.
FORWARD (start from CAUSE → follow to EFFECT) = RCT, cohort
exposure known first ───────────────► outcome later
BACKWARD (start from EFFECT → look back at CAUSE) = case-control
outcome known first ◄─────────────── exposure in the past
Forward designs are prospective: you fix the exposure groups, then wait to see outcomes. Backward designs fix the outcome groups (cases vs controls), then reconstruct exposure. A retrospective cohort is a hybrid — the events already happened, but you still reason forward from exposure to outcome using old records.
3.3 What the sampling direction lets you measure
How you sampled people decides whether "risk" even means anything.
| You sampled by... | Designs | Are row totals real? | Valid measures |
|---|---|---|---|
| Exposure | RCT, cohort | Yes — you know everyone exposed and not | RR, RD, NNT (and OR) |
| Outcome | Case-control | No — you chose the case:control ratio | OR only |
When you sample by exposure, the number of exposed and unexposed people is a real feature of your study, so the fraction who get the disease — the risk — is meaningful. That unlocks the relative risk (RR), risk difference (RD), and number needed to treat (NNT). (The odds ratio still works too.)
When you sample by outcome (case-control), you picked how many cases and controls to enroll — maybe 100 of each, maybe 1 case per 4 controls. That ratio is an arbitrary design choice, so any "risk" you compute from it is a fiction. The one measure that survives this is the odds ratio (OR), because of its symmetry — the OR is the same whether you sample forward or backward. See 2×2 Tables & Effect Measures (OR, RR, RD, NNT, HR) for the algebra.
3.4 Randomized vs observational — and why GRADE cares
The deepest split in the pyramid is randomized vs observational, and it exists because of one word: confounding.
Confounding is a hidden difference between the groups that isn't the exposure you care about. Example: smokers also tend to drink more alcohol. If smokers get more of some disease, is it the smoke or the drink? In an observational study you can't be sure. Randomization fixes this — a coin flip makes the groups balanced on everything, even things you never measured.
Because of that, the GRADE framework (see GRADE: Rating Certainty of Evidence) gives each design a starting certainty:
RCT ──► start HIGH (randomizing removes confounding)
Cohort ──► start LOW (observation → possible confounding)
Case-control ──► start LOW (observation → possible confounding)
An RCT begins with the benefit of the doubt and can be downgraded for flaws; observational evidence begins low and must earn upgrades (e.g. a very large effect). This is why "the study was observational" is not an insult — it's a starting position.
3.5 The risk-of-bias tool follows the design
Every design has its own checklist for how it could have gone wrong (see Risk of Bias: RoB2, ROBINS-I, AMSTAR-2).
| Design | Risk-of-bias tool | What it checks |
|---|---|---|
| RCT | RoB2 | Randomization, deviations, missing data, measurement, selective reporting |
| Cohort / case-control (non-randomized) | ROBINS-I | Confounding, selection, classification of exposure, and more |
| Systematic review | AMSTAR-2 | Was the review itself done rigorously — search, selection, pooling |
Pick the wrong tool and your appraisal is meaningless — you can't ask RoB2's "was randomization concealed?" of a cohort study that never randomized anything.
4. The rules
The design of a study is a key that unlocks three downstream decisions. Memorize the mapping.
STEP 1 — name the design (read the abstract)
STEP 2 — apply the three rules:
RULE A (effect measure): sampled by EXPOSURE → RR, RD, NNT (+OR)
sampled by OUTCOME → OR only
RULE B (GRADE start): RANDOMIZED → start HIGH
OBSERVATIONAL → start LOW
RULE C (risk-of-bias): RCT → RoB2
non-randomized → ROBINS-I
systematic review → AMSTAR-2
A worked mapping. You are handed this abstract: "We enrolled 1,000 smokers and 1,000 non-smokers and followed them for 10 years, comparing lung-cancer incidence."
- Name it. People are split by an exposure (smoking) and followed forward → this is a cohort study.
- Rule A — measure. Sampled by exposure, so risk is real → you may report RR, RD, NNT (and OR). "Smokers had 5× the risk (RR = 5)" is a valid sentence here.
- Rule B — GRADE. Observational → start LOW. (But smoking→lung-cancer has such a huge effect that GRADE would likely upgrade it.)
- Rule C — RoB tool. Non-randomized → appraise with ROBINS-I, paying special attention to confounding (do smokers differ in other ways?).
Now a second abstract: "200 hypertensive patients were randomly assigned to a new pill or placebo and followed for one year." → RCT → risk is real, all measures valid → GRADE starts HIGH → appraise with RoB2. Same three rules, different answers, because the design changed.
5. Real code
A single deterministic function maps a study design to its risk-of-bias tool, GRADE starting certainty, and valid effect measures. No external libraries — pure Python, so it runs anywhere.
"""Map a study design to: its risk-of-bias tool, GRADE starting certainty,
and which effect measures are valid. Deterministic lookup, no dependencies."""
# Knowledge table. `sampled_by` drives which effect measures are valid:
# EXPOSURE-sampled (rct, cohort) -> risk is real -> RR/RD/NNT valid (+OR)
# OUTCOME-sampled (case_control) -> risk is a design artifact -> OR only
DESIGNS = {
"rct": {
"randomized": True, "sampled_by": "exposure",
"rob_tool": "RoB2", "grade_start": "high",
},
"cohort": {
"randomized": False, "sampled_by": "exposure",
"rob_tool": "ROBINS-I", "grade_start": "low",
},
"case_control": {
"randomized": False, "sampled_by": "outcome",
"rob_tool": "ROBINS-I", "grade_start": "low",
},
"systematic_review": {
"randomized": None, "sampled_by": None,
"rob_tool": "AMSTAR-2", "grade_start": "depends-on-included-studies",
},
}
def assess_design(study_design):
"""Return {rob_tool, grade_start, valid_measures} for a design name.
valid_measures rule:
exposure-sampled -> ["RR", "RD", "NNT", "OR"] (risk is meaningful)
outcome-sampled -> ["OR"] (risk is meaningless)
systematic review -> inherits from pooled studies (["OR", "RR", "RD"])
"""
key = study_design.strip().lower().replace("-", "_").replace(" ", "_")
if key not in DESIGNS:
raise ValueError(f"unknown design: {study_design!r}")
info = DESIGNS[key]
if info["sampled_by"] == "exposure":
measures = ["RR", "RD", "NNT", "OR"]
elif info["sampled_by"] == "outcome":
measures = ["OR"] # only the odds ratio survives
else: # systematic review
measures = ["OR", "RR", "RD"] # whatever its trials reported
return {
"rob_tool": info["rob_tool"],
"grade_start": info["grade_start"],
"valid_measures": measures,
}
if __name__ == "__main__":
header = f"{'design':18s} {'RoB tool':10s} {'GRADE start':28s} valid measures"
print(header)
print("-" * len(header))
for design in ("rct", "cohort", "case_control", "systematic_review"):
r = assess_design(design)
print(f"{design:18s} {r['rob_tool']:10s} "
f"{r['grade_start']:28s} {', '.join(r['valid_measures'])}")
Expected output:
design RoB tool GRADE start valid measures
--------------------------------------------------------------------------
rct RoB2 high RR, RD, NNT, OR
cohort ROBINS-I low RR, RD, NNT, OR
case_control ROBINS-I low OR
systematic_review AMSTAR-2 depends-on-included-studies OR, RR, RD
Read the table top to bottom and you can see all three rules at once: the RoB tool tracks randomization, the GRADE start tracks randomized-vs-observational, and the valid-measure list collapses to OR only for the outcome-sampled case-control design.
6. Real-world example
One clinical question, answered by four designs — with small numbers.
-
RCT — does a new blood-pressure pill prevent heart attacks? Coin-flip 200 patients: 100 to the pill, 100 to placebo. After a year, 15 heart attacks in the pill group, 30 in placebo. Because the coin flip balanced the groups, the difference is the drug's doing. Risk is real, so report RR = 0.50 ("halves the risk"), RD = −0.15, NNT = 7. GRADE starts high; appraise with RoB2.
-
Cohort — does smoking cause lung cancer? You can't ethically assign people to smoke, so you observe: follow 1,000 smokers and 1,000 non-smokers for 10 years. Say 80 smokers and 16 non-smokers develop lung cancer. Sampled by exposure → RR = (80/1000)/(16/1000) = 5.0. GRADE starts low (observational), but the effect is so large it would likely be upgraded. Appraise with ROBINS-I — watch for confounding (do smokers also differ in diet, alcohol, occupation?).
-
Case-control — what causes a rare cancer? The cancer is too rare to follow a cohort, so start from the outcome: 100 patients with the cancer (cases) + 100 without (controls), and ask about past asbestos exposure. Suppose 40 cases vs 10 controls were exposed. You chose 100:100, so risk is meaningless — report the odds ratio only: OR = (40·90)/(60·10) = 6.0. GRADE starts low; appraise with ROBINS-I.
-
A new surgery, described as a case series. A surgeon reports 12 patients who got a novel operation and all recovered well. Encouraging — but there is no control group, so we can't know how they'd have done without it. This sits low on the pyramid; it generates a hypothesis that an RCT should later test. No comparison → no valid effect measure, no formal RoB tool beyond descriptive appraisal.
-
On top of it all — the systematic review. Later, a review gathers every RCT of that blood-pressure pill, appraises each with RoB2, pools them with meta-analysis, and appraises itself with AMSTAR-2. One pooled estimate, built from the strongest rungs of the pyramid, sits at the very top. See Meta-Analysis: Pooling Studies (Fixed vs Random Effects).
7. Interview questions companies actually ask
Q [Pfizer / clinical epidemiology] "Walk me up the evidence pyramid, top to bottom."
A Systematic review / meta-analysis at the top (pools many studies), then RCT (randomized,
strongest single-study proof of cause), then cohort (observe by exposure, follow forward),
case-control (start from outcome, look backward, good for rare disease), cross-sectional
(snapshot, prevalence), then case series, case report, and expert opinion at the bottom.
Higher = better protected from luck and bias, especially from confounding.
Q [a health-tech startup] "When would you choose a case-control study over a cohort?"
A When the outcome is RARE. A cohort would need to follow a huge number of people for years
to catch a few cases — slow and expensive. A case-control study starts with the cases you
already have plus matched controls and looks backward at exposure, so it's fast and cheap.
The price: you sampled by outcome, so you can only report the odds ratio, and you must guard
against recall and selection bias.
Q [Genentech / biostatistics] "Why can't you compute a relative risk from a case-control study?"
A Because you fixed the case:control ratio by design — say 100 cases to 100 controls. That
ratio isn't a real feature of the population, so any 'risk' derived from it is an artifact.
Only the odds ratio is valid, because the OR is symmetric: it's the same whether you sample
forward (by exposure) or backward (by outcome).
Q [a CRO / medical affairs] "Why does an RCT start HIGH on GRADE but a cohort starts LOW?"
A Randomization. A coin flip makes the treatment and control groups balanced on everything,
even unmeasured factors, so a difference in outcome is attributable to the treatment.
Observational designs leave room for confounding — hidden differences between the exposed and
unexposed. GRADE reflects that by starting RCTs high and observational studies low.
Q [UnitedHealth / clinical analytics] "What is confounding? Give an example."
A A confounder is a variable linked to both the exposure and the outcome that isn't the causal
path you care about. Classic example: smokers also tend to drink more, so if smokers show
more of some disease, alcohol could be the real driver. Randomization removes confounding by
balancing groups; in observational studies you must adjust for it statistically and can never
be fully sure you caught it all.
Q [Moderna / vaccine epidemiology] "Prospective vs retrospective cohort — what's the difference?"
A Both split people by exposure and reason forward to outcome. A prospective cohort enrolls
people now and follows them into the future in real time. A retrospective cohort uses records
where the outcomes have already occurred, reconstructing the follow-up from the past. Same
logic, different timing; retrospective is faster but limited by whatever the old records
captured.
Q [a diagnostics company] "Which risk-of-bias tool goes with which design?"
A RoB2 for randomized trials, ROBINS-I for non-randomized studies (cohort and case-control),
and AMSTAR-2 for systematic reviews. You must match the tool to the design — RoB2's questions
about randomization make no sense for a cohort that never randomized anyone.
Q [Bristol Myers Squibb] "A case series shows all 12 patients on a new drug improved. Is that
proof it works?"
A No. A case series has no control group, so we don't know how those patients would have done
without the drug — they might have improved anyway. It's low on the pyramid: useful for
generating a hypothesis, not for proving cause. The proper next step is an RCT.
Q [a payer / HEOR] "You want to study a rare adverse event of a widely used drug. Which design?"
A Case-control. The event is rare, so a cohort or trial would be impractical. Start with people
who had the event (cases) and comparable people who didn't (controls), then look back at drug
exposure and report the odds ratio. Watch for recall bias and confounding by indication.
Q [Novartis / evidence synthesis] "Where does a cross-sectional study fit, and what's its main
limitation?"
A It's a one-moment snapshot, good for measuring prevalence and associations. Its main
limitation is no time direction: you measure exposure and outcome simultaneously, so you
can't tell which came first — it can't establish temporality, and therefore can't establish
cause.
8. When to use / tradeoffs
MATCH THE DESIGN TO THE QUESTION:
✓ RCT — "does this treatment CAUSE benefit?" strongest, but costly + must be ethical
✓ Cohort — harms/prognosis; when randomizing is impossible or unethical (e.g. smoking)
✓ Case-control — a RARE outcome; fast + cheap; start from the disease and look back
✓ Cross-section — prevalence + associations RIGHT NOW; a quick lay of the land
✓ Case series — a brand-new treatment or an early signal; hypothesis-generating only
✓ Systematic review — pull the whole body of evidence into ONE pooled answer
HONEST LIMITS:
✗ RCTs can be unethical (can't assign people to smoke) or too slow/costly for rare outcomes
✗ observational designs risk CONFOUNDING — a difference you didn't randomize away
✗ case-control gives ONLY the odds ratio; no risk, no RR/RD/NNT
✗ cross-sectional can't order events in time → can't prove cause
✗ case series / report / expert opinion have NO control → never prove cause
✓ a systematic review is only as strong as the studies it pools (garbage in, garbage out)
THE RANKING RULE:
higher on the pyramid = better protection from luck and confounding — but the RIGHT design
is the one that fits the question, not always the highest one you can afford.
The pyramid ranks internal validity — how sure we are the result is real. But the best study is the one that fits the question and is feasible and ethical: you don't run an RCT to study a poison, and you don't run a case series when you could run a trial. Design first, then let the design tell you which effect measure, which GRADE start, and which appraisal tool follow.
9. Summary + related articles
- Evidence is ranked in a pyramid: systematic review / meta-analysis → RCT → cohort → case-control → cross-sectional → case series → case report → expert opinion.
- RCT randomizes and follows forward — the strongest single-study proof of cause; cohort observes by exposure and follows forward; case-control starts from the outcome and looks backward (best for rare disease); cross-sectional is a one-moment snapshot.
- Forward designs (RCT, cohort) go cause → effect; backward designs (case-control) go effect → cause.
- Sampling by exposure (RCT, cohort) makes risk real → RR, RD, NNT valid (plus OR); sampling by outcome (case-control) makes risk a fiction → only the OR is valid.
- RCT starts HIGH on GRADE, observational starts LOW, because randomizing removes confounding while observing does not.
- Match the risk-of-bias tool to the design: RoB2 (RCT), ROBINS-I (non-randomized), AMSTAR-2 (systematic review).
- Case series, case reports, and expert opinion have no control group — they raise hypotheses, they never prove cause.
Related: 2×2 Tables & Effect Measures (OR, RR, RD, NNT, HR) · Meta-Analysis: Pooling Studies (Fixed vs Random Effects) · GRADE: Rating Certainty of Evidence · Risk of Bias: RoB2, ROBINS-I, AMSTAR-2
Resources
- Cochrane Handbook, Ch. 3 "Defining the criteria for including studies" — https://training.cochrane.org/handbook/current/chapter-03
- GRADE Handbook — https://gdt.gradepro.org/app/handbook/handbook.html
- Sterne et al., "RoB 2: a revised tool for assessing risk of bias in randomised trials" — https://www.bmj.com/content/366/bmj.l4898
- Sterne et al., "ROBINS-I: a tool for assessing risk of bias in non-randomised studies" — https://www.bmj.com/content/355/bmj.i4919
- Shea et al., "AMSTAR 2: a critical appraisal tool for systematic reviews" — https://www.bmj.com/content/358/bmj.j4008
- CEBM, "Study Designs" — https://www.cebm.ox.ac.uk/resources/ebm-tools/study-designs