← Back to Learning Hub

Study Designs & the Evidence Hierarchy

Study designs2×2Intermediate22 min

By: Anacodic Team

TL;DR — Not all evidence is equal. Sort study designs into a pyramid: at the top, systematic reviews / meta-analyses that pool many studies; below them, the RCT (randomize people to treatment vs control, follow forward — the strongest proof of cause); then observational designs — cohort (split by exposure, follow forward), case-control (start from the outcome, look backward, best for rare diseases), and cross-sectional (a snapshot, prevalence only); at the bottom, case series, case reports, and expert opinion (no comparison group, weakest). The design decides three downstream things: which effect measure is valid (exposure-sampled designs support RR/RD/NNT; outcome-sampled case-control supports only the odds ratio), where you start on GRADE (RCT starts high, observational starts low, because randomizing removes confounding), and which risk-of-bias tool you reach for (RoB2 for RCTs, ROBINS-I for non-randomized, AMSTAR-2 for systematic reviews). Learn to read a study's design off its abstract and the rest follows.


1. Simple explanation

Evidence comes in grades. If a friend says "this pill cured my headache," that is one person's story — interesting, but it could be luck, the pill, or the fact that headaches fade on their own. If instead a thousand people were split by coin flip into "pill" and "sugar tablet," followed for a month, and the pill group did clearly better, that is far harder to dismiss. Same claim, wildly different strength. The evidence hierarchy is just a way of ranking how much we should trust a result based on how the study was built.

Analogy — a court of law. Think of building a legal case.

  • Expert opinion is a character witness saying "I think he's guilty." No evidence, just judgment.
  • A case report is one eyewitness: "I saw one thing happen once." Vivid, but a single account.
  • A case series is several eyewitnesses who all saw the defendant — but nobody watched anyone else, so there is no comparison.
  • A cross-sectional survey is a snapshot of the whole town on one day: who was where, right now.
  • A case-control study starts from the crime (the outcome) and works backward to who was near the scene.
  • A cohort study follows suspects forward in time to see who commits the crime.
  • An RCT is the gold standard: you could randomly assign who gets exposed, so the two groups are otherwise identical — the cleanest possible proof of cause.
  • A systematic review is the appeals court reading every trial transcript and weighing them together.

The higher you climb, the more the design protects you from being fooled by luck, bias, and hidden differences between groups. This article teaches you to name a design on sight and to know what that name lets you do next.

The three questions it answers: What is each design and when is it used? Which way does time run — forward or backward? And what does the design let you compute, claim, and score afterward?


2. Diagram

                     THE EVIDENCE PYRAMID
             (top = strongest cause proof, bottom = weakest)

                       ╱╲
                      ╱  ╲     SYSTEMATIC REVIEW / META-ANALYSIS
                     ╱ SR ╲     pool many studies into one estimate
                    ╱──────╲
                   ╱        ╲   RANDOMIZED CONTROLLED TRIAL (RCT)
                  ╱   RCT    ╲   coin-flip to treat vs control → forward
                 ╱────────────╲
                ╱              ╲ COHORT
               ╱    cohort      ╲ split by EXPOSURE → follow forward
              ╱──────────────────╲
             ╱                    ╲ CASE-CONTROL
            ╱     case-control     ╲ start from OUTCOME → look backward
           ╱────────────────────────╲
          ╱                          ╲ CROSS-SECTIONAL
         ╱      cross-sectional       ╲ one snapshot in time (prevalence)
        ╱──────────────────────────────╲
       ╱        case series             ╲ a group treated, NO control
      ╱──────────────────────────────────╲
     ╱          case report               ╲ a single patient
    ╱──────────────────────────────────────╲
   ╱            expert opinion               ╲ no data, just judgment
  ╱────────────────────────────────────────────╲

   TIME DIRECTION                     ANALYTIC vs DESCRIPTIVE
   ─────────────                      ────────────────────────
   forward  → RCT, cohort              has a comparison group:
   (cause → effect, prospective)         RCT, cohort, case-control
   backward → case-control              no comparison group:
   (effect → cause, look back)            case series, case report

3. How it works

3.1 The designs, one at a time

DesignWhat you doTime directionSampled byBest for
RCTRandomize people to treatment vs control, follow forwardForward (prospective)Exposure (assigned)Proving a treatment causes an effect
CohortObserve a group split by exposure, follow forward to outcomeForwardExposureHarms/prognosis; when randomizing is impossible
Case-controlStart from cases (have disease) + controls (don't), look back at exposureBackwardOutcomeRare diseases; cheap, fast
Cross-sectionalMeasure exposure + outcome at one momentNone (snapshot)The population nowPrevalence, associations
Case seriesDescribe a group who got a treatment, no controlUsually forwardHypothesis-generating
Case reportDescribe a single patientFlagging something new/rare
Expert opinionJudgment, no collected dataFilling gaps when nothing else exists

RCT (randomized controlled trial). The researcher decides who is exposed by a random draw — a coin flip. Because the assignment is random, the two groups are, on average, identical in everything else (age, habits, hidden illnesses). So if outcomes differ, the treatment is the reason. Example: 200 patients with high blood pressure are coin-flipped to a new BP pill or a placebo, then followed forward for a year to compare heart attacks.

Cohort. You observe rather than assign. You find people already split by an exposure and follow them forward. Prospective cohorts follow in real time; retrospective cohorts reconstruct the follow-up from old records. Example: follow 1,000 smokers and 1,000 non-smokers for 10 years and count lung cancers. You couldn't ethically assign people to smoke, so you observe.

Case-control. You start from the outcome. Gather people who already have the disease (cases) and a comparable group who don't (controls), then look backward at their past exposures. This is the trick for rare diseases — you don't have to follow a million people hoping a few get a rare cancer; you just start with the ones who have it. Example: 100 patients with a rare cancer + 100 without, asking each about past asbestos exposure.

Cross-sectional. A single snapshot. You measure exposure and outcome at the same moment, so you get prevalence ("how many have it right now") and associations — but no sense of what came first. Example: survey a city on one day for both obesity and knee pain; you'll see they travel together, but not which caused which.

Case series / case report / expert opinion. These have no comparison group. A case series describes several patients who got a treatment; a case report describes one; expert opinion is judgment with no collected data. They can raise a hypothesis ("five patients on this drug all developed the same rash") but can never prove cause, because there is nothing to compare against.

3.2 Forward vs backward — the arrow of time

This single distinction organizes half the pyramid.

   FORWARD  (start from CAUSE → follow to EFFECT)      = RCT, cohort
       exposure known first ───────────────► outcome later

   BACKWARD (start from EFFECT → look back at CAUSE)   = case-control
       outcome known first ◄─────────────── exposure in the past

Forward designs are prospective: you fix the exposure groups, then wait to see outcomes. Backward designs fix the outcome groups (cases vs controls), then reconstruct exposure. A retrospective cohort is a hybrid — the events already happened, but you still reason forward from exposure to outcome using old records.

3.3 What the sampling direction lets you measure

How you sampled people decides whether "risk" even means anything.

You sampled by...DesignsAre row totals real?Valid measures
ExposureRCT, cohortYes — you know everyone exposed and notRR, RD, NNT (and OR)
OutcomeCase-controlNo — you chose the case:control ratioOR only

When you sample by exposure, the number of exposed and unexposed people is a real feature of your study, so the fraction who get the disease — the risk — is meaningful. That unlocks the relative risk (RR), risk difference (RD), and number needed to treat (NNT). (The odds ratio still works too.)

When you sample by outcome (case-control), you picked how many cases and controls to enroll — maybe 100 of each, maybe 1 case per 4 controls. That ratio is an arbitrary design choice, so any "risk" you compute from it is a fiction. The one measure that survives this is the odds ratio (OR), because of its symmetry — the OR is the same whether you sample forward or backward. See 2×2 Tables & Effect Measures (OR, RR, RD, NNT, HR) for the algebra.

3.4 Randomized vs observational — and why GRADE cares

The deepest split in the pyramid is randomized vs observational, and it exists because of one word: confounding.

Confounding is a hidden difference between the groups that isn't the exposure you care about. Example: smokers also tend to drink more alcohol. If smokers get more of some disease, is it the smoke or the drink? In an observational study you can't be sure. Randomization fixes this — a coin flip makes the groups balanced on everything, even things you never measured.

Because of that, the GRADE framework (see GRADE: Rating Certainty of Evidence) gives each design a starting certainty:

   RCT               ──► start HIGH   (randomizing removes confounding)
   Cohort            ──► start LOW    (observation → possible confounding)
   Case-control      ──► start LOW    (observation → possible confounding)

An RCT begins with the benefit of the doubt and can be downgraded for flaws; observational evidence begins low and must earn upgrades (e.g. a very large effect). This is why "the study was observational" is not an insult — it's a starting position.

3.5 The risk-of-bias tool follows the design

Every design has its own checklist for how it could have gone wrong (see Risk of Bias: RoB2, ROBINS-I, AMSTAR-2).

DesignRisk-of-bias toolWhat it checks
RCTRoB2Randomization, deviations, missing data, measurement, selective reporting
Cohort / case-control (non-randomized)ROBINS-IConfounding, selection, classification of exposure, and more
Systematic reviewAMSTAR-2Was the review itself done rigorously — search, selection, pooling

Pick the wrong tool and your appraisal is meaningless — you can't ask RoB2's "was randomization concealed?" of a cohort study that never randomized anything.


4. The rules

The design of a study is a key that unlocks three downstream decisions. Memorize the mapping.

   STEP 1 — name the design (read the abstract)
   STEP 2 — apply the three rules:

     RULE A (effect measure):  sampled by EXPOSURE  → RR, RD, NNT (+OR)
                               sampled by OUTCOME    → OR only
     RULE B (GRADE start):     RANDOMIZED            → start HIGH
                               OBSERVATIONAL         → start LOW
     RULE C (risk-of-bias):    RCT                   → RoB2
                               non-randomized        → ROBINS-I
                               systematic review     → AMSTAR-2

A worked mapping. You are handed this abstract: "We enrolled 1,000 smokers and 1,000 non-smokers and followed them for 10 years, comparing lung-cancer incidence."

  1. Name it. People are split by an exposure (smoking) and followed forward → this is a cohort study.
  2. Rule A — measure. Sampled by exposure, so risk is real → you may report RR, RD, NNT (and OR). "Smokers had 5× the risk (RR = 5)" is a valid sentence here.
  3. Rule B — GRADE. Observational → start LOW. (But smoking→lung-cancer has such a huge effect that GRADE would likely upgrade it.)
  4. Rule C — RoB tool. Non-randomized → appraise with ROBINS-I, paying special attention to confounding (do smokers differ in other ways?).

Now a second abstract: "200 hypertensive patients were randomly assigned to a new pill or placebo and followed for one year."RCT → risk is real, all measures valid → GRADE starts HIGH → appraise with RoB2. Same three rules, different answers, because the design changed.


5. Real code

A single deterministic function maps a study design to its risk-of-bias tool, GRADE starting certainty, and valid effect measures. No external libraries — pure Python, so it runs anywhere.

"""Map a study design to: its risk-of-bias tool, GRADE starting certainty,
and which effect measures are valid. Deterministic lookup, no dependencies."""

# Knowledge table. `sampled_by` drives which effect measures are valid:
#   EXPOSURE-sampled (rct, cohort) -> risk is real -> RR/RD/NNT valid (+OR)
#   OUTCOME-sampled  (case_control) -> risk is a design artifact -> OR only
DESIGNS = {
    "rct": {
        "randomized": True,  "sampled_by": "exposure",
        "rob_tool": "RoB2",  "grade_start": "high",
    },
    "cohort": {
        "randomized": False, "sampled_by": "exposure",
        "rob_tool": "ROBINS-I", "grade_start": "low",
    },
    "case_control": {
        "randomized": False, "sampled_by": "outcome",
        "rob_tool": "ROBINS-I", "grade_start": "low",
    },
    "systematic_review": {
        "randomized": None,  "sampled_by": None,
        "rob_tool": "AMSTAR-2", "grade_start": "depends-on-included-studies",
    },
}

def assess_design(study_design):
    """Return {rob_tool, grade_start, valid_measures} for a design name.

    valid_measures rule:
      exposure-sampled -> ["RR", "RD", "NNT", "OR"]  (risk is meaningful)
      outcome-sampled  -> ["OR"]                     (risk is meaningless)
      systematic review -> inherits from pooled studies (["OR", "RR", "RD"])
    """
    key = study_design.strip().lower().replace("-", "_").replace(" ", "_")
    if key not in DESIGNS:
        raise ValueError(f"unknown design: {study_design!r}")
    info = DESIGNS[key]

    if info["sampled_by"] == "exposure":
        measures = ["RR", "RD", "NNT", "OR"]
    elif info["sampled_by"] == "outcome":
        measures = ["OR"]                       # only the odds ratio survives
    else:                                       # systematic review
        measures = ["OR", "RR", "RD"]           # whatever its trials reported

    return {
        "rob_tool": info["rob_tool"],
        "grade_start": info["grade_start"],
        "valid_measures": measures,
    }

if __name__ == "__main__":
    header = f"{'design':18s} {'RoB tool':10s} {'GRADE start':28s} valid measures"
    print(header)
    print("-" * len(header))
    for design in ("rct", "cohort", "case_control", "systematic_review"):
        r = assess_design(design)
        print(f"{design:18s} {r['rob_tool']:10s} "
              f"{r['grade_start']:28s} {', '.join(r['valid_measures'])}")

Expected output:

design             RoB tool   GRADE start                  valid measures
--------------------------------------------------------------------------
rct                RoB2       high                         RR, RD, NNT, OR
cohort             ROBINS-I   low                          RR, RD, NNT, OR
case_control       ROBINS-I   low                          OR
systematic_review  AMSTAR-2   depends-on-included-studies  OR, RR, RD

Read the table top to bottom and you can see all three rules at once: the RoB tool tracks randomization, the GRADE start tracks randomized-vs-observational, and the valid-measure list collapses to OR only for the outcome-sampled case-control design.


6. Real-world example

One clinical question, answered by four designs — with small numbers.

  • RCT — does a new blood-pressure pill prevent heart attacks? Coin-flip 200 patients: 100 to the pill, 100 to placebo. After a year, 15 heart attacks in the pill group, 30 in placebo. Because the coin flip balanced the groups, the difference is the drug's doing. Risk is real, so report RR = 0.50 ("halves the risk"), RD = −0.15, NNT = 7. GRADE starts high; appraise with RoB2.

  • Cohort — does smoking cause lung cancer? You can't ethically assign people to smoke, so you observe: follow 1,000 smokers and 1,000 non-smokers for 10 years. Say 80 smokers and 16 non-smokers develop lung cancer. Sampled by exposure → RR = (80/1000)/(16/1000) = 5.0. GRADE starts low (observational), but the effect is so large it would likely be upgraded. Appraise with ROBINS-I — watch for confounding (do smokers also differ in diet, alcohol, occupation?).

  • Case-control — what causes a rare cancer? The cancer is too rare to follow a cohort, so start from the outcome: 100 patients with the cancer (cases) + 100 without (controls), and ask about past asbestos exposure. Suppose 40 cases vs 10 controls were exposed. You chose 100:100, so risk is meaningless — report the odds ratio only: OR = (40·90)/(60·10) = 6.0. GRADE starts low; appraise with ROBINS-I.

  • A new surgery, described as a case series. A surgeon reports 12 patients who got a novel operation and all recovered well. Encouraging — but there is no control group, so we can't know how they'd have done without it. This sits low on the pyramid; it generates a hypothesis that an RCT should later test. No comparison → no valid effect measure, no formal RoB tool beyond descriptive appraisal.

  • On top of it all — the systematic review. Later, a review gathers every RCT of that blood-pressure pill, appraises each with RoB2, pools them with meta-analysis, and appraises itself with AMSTAR-2. One pooled estimate, built from the strongest rungs of the pyramid, sits at the very top. See Meta-Analysis: Pooling Studies (Fixed vs Random Effects).


7. Interview questions companies actually ask

Q [Pfizer / clinical epidemiology] "Walk me up the evidence pyramid, top to bottom."
  A Systematic review / meta-analysis at the top (pools many studies), then RCT (randomized,
    strongest single-study proof of cause), then cohort (observe by exposure, follow forward),
    case-control (start from outcome, look backward, good for rare disease), cross-sectional
    (snapshot, prevalence), then case series, case report, and expert opinion at the bottom.
    Higher = better protected from luck and bias, especially from confounding.

Q [a health-tech startup] "When would you choose a case-control study over a cohort?"
  A When the outcome is RARE. A cohort would need to follow a huge number of people for years
    to catch a few cases — slow and expensive. A case-control study starts with the cases you
    already have plus matched controls and looks backward at exposure, so it's fast and cheap.
    The price: you sampled by outcome, so you can only report the odds ratio, and you must guard
    against recall and selection bias.

Q [Genentech / biostatistics] "Why can't you compute a relative risk from a case-control study?"
  A Because you fixed the case:control ratio by design — say 100 cases to 100 controls. That
    ratio isn't a real feature of the population, so any 'risk' derived from it is an artifact.
    Only the odds ratio is valid, because the OR is symmetric: it's the same whether you sample
    forward (by exposure) or backward (by outcome).

Q [a CRO / medical affairs] "Why does an RCT start HIGH on GRADE but a cohort starts LOW?"
  A Randomization. A coin flip makes the treatment and control groups balanced on everything,
    even unmeasured factors, so a difference in outcome is attributable to the treatment.
    Observational designs leave room for confounding — hidden differences between the exposed and
    unexposed. GRADE reflects that by starting RCTs high and observational studies low.

Q [UnitedHealth / clinical analytics] "What is confounding? Give an example."
  A A confounder is a variable linked to both the exposure and the outcome that isn't the causal
    path you care about. Classic example: smokers also tend to drink more, so if smokers show
    more of some disease, alcohol could be the real driver. Randomization removes confounding by
    balancing groups; in observational studies you must adjust for it statistically and can never
    be fully sure you caught it all.

Q [Moderna / vaccine epidemiology] "Prospective vs retrospective cohort — what's the difference?"
  A Both split people by exposure and reason forward to outcome. A prospective cohort enrolls
    people now and follows them into the future in real time. A retrospective cohort uses records
    where the outcomes have already occurred, reconstructing the follow-up from the past. Same
    logic, different timing; retrospective is faster but limited by whatever the old records
    captured.

Q [a diagnostics company] "Which risk-of-bias tool goes with which design?"
  A RoB2 for randomized trials, ROBINS-I for non-randomized studies (cohort and case-control),
    and AMSTAR-2 for systematic reviews. You must match the tool to the design — RoB2's questions
    about randomization make no sense for a cohort that never randomized anyone.

Q [Bristol Myers Squibb] "A case series shows all 12 patients on a new drug improved. Is that
   proof it works?"
  A No. A case series has no control group, so we don't know how those patients would have done
    without the drug — they might have improved anyway. It's low on the pyramid: useful for
    generating a hypothesis, not for proving cause. The proper next step is an RCT.

Q [a payer / HEOR] "You want to study a rare adverse event of a widely used drug. Which design?"
  A Case-control. The event is rare, so a cohort or trial would be impractical. Start with people
    who had the event (cases) and comparable people who didn't (controls), then look back at drug
    exposure and report the odds ratio. Watch for recall bias and confounding by indication.

Q [Novartis / evidence synthesis] "Where does a cross-sectional study fit, and what's its main
   limitation?"
  A It's a one-moment snapshot, good for measuring prevalence and associations. Its main
    limitation is no time direction: you measure exposure and outcome simultaneously, so you
    can't tell which came first — it can't establish temporality, and therefore can't establish
    cause.

8. When to use / tradeoffs

   MATCH THE DESIGN TO THE QUESTION:
     ✓ RCT           — "does this treatment CAUSE benefit?"  strongest, but costly + must be ethical
     ✓ Cohort        — harms/prognosis; when randomizing is impossible or unethical (e.g. smoking)
     ✓ Case-control  — a RARE outcome; fast + cheap; start from the disease and look back
     ✓ Cross-section — prevalence + associations RIGHT NOW; a quick lay of the land
     ✓ Case series   — a brand-new treatment or an early signal; hypothesis-generating only
     ✓ Systematic review — pull the whole body of evidence into ONE pooled answer
   HONEST LIMITS:
     ✗ RCTs can be unethical (can't assign people to smoke) or too slow/costly for rare outcomes
     ✗ observational designs risk CONFOUNDING — a difference you didn't randomize away
     ✗ case-control gives ONLY the odds ratio; no risk, no RR/RD/NNT
     ✗ cross-sectional can't order events in time → can't prove cause
     ✗ case series / report / expert opinion have NO control → never prove cause
     ✓ a systematic review is only as strong as the studies it pools (garbage in, garbage out)
   THE RANKING RULE:
     higher on the pyramid = better protection from luck and confounding — but the RIGHT design
     is the one that fits the question, not always the highest one you can afford.

The pyramid ranks internal validity — how sure we are the result is real. But the best study is the one that fits the question and is feasible and ethical: you don't run an RCT to study a poison, and you don't run a case series when you could run a trial. Design first, then let the design tell you which effect measure, which GRADE start, and which appraisal tool follow.


  • Evidence is ranked in a pyramid: systematic review / meta-analysis → RCT → cohort → case-control → cross-sectional → case series → case report → expert opinion.
  • RCT randomizes and follows forward — the strongest single-study proof of cause; cohort observes by exposure and follows forward; case-control starts from the outcome and looks backward (best for rare disease); cross-sectional is a one-moment snapshot.
  • Forward designs (RCT, cohort) go cause → effect; backward designs (case-control) go effect → cause.
  • Sampling by exposure (RCT, cohort) makes risk real → RR, RD, NNT valid (plus OR); sampling by outcome (case-control) makes risk a fiction → only the OR is valid.
  • RCT starts HIGH on GRADE, observational starts LOW, because randomizing removes confounding while observing does not.
  • Match the risk-of-bias tool to the design: RoB2 (RCT), ROBINS-I (non-randomized), AMSTAR-2 (systematic review).
  • Case series, case reports, and expert opinion have no control group — they raise hypotheses, they never prove cause.

Related: 2×2 Tables & Effect Measures (OR, RR, RD, NNT, HR) · Meta-Analysis: Pooling Studies (Fixed vs Random Effects) · GRADE: Rating Certainty of Evidence · Risk of Bias: RoB2, ROBINS-I, AMSTAR-2

Resources