← Back to Learning Hub

Summary of Findings & Evidence-to-Decision

GRADE certaintyRisk of biasAdvanced23 min

By: Anacodic Team

TL;DR — Once you've pooled the evidence and rated its certainty, two structured artifacts turn numbers into a decision. A Summary of Findings (SoF) table is the one-page scorecard: one row per critical outcome, columns for № of participants (studies), certainty (GRADE), relative effect (RR/OR with CI), and absolute effect (risk with vs without + the difference). Then the Evidence-to-Decision (EtD) framework walks a panel from that evidence to a recommendation through explicit criteria — problem priority, benefits vs harms, certainty, values/preferences, resources, equity, acceptability, feasibility — and outputs a direction (for/against) and a strength: STRONG (benefits clearly outweigh harms, most patients would choose it) or CONDITIONAL/weak (trade-offs are close, depends on values). Panels reach agreement by Delphi-style voting against a threshold (e.g., ≥70–80%); when evidence and panel disagree, or the threshold isn't met, that's discordance — flag it and revote. High certainty + clear net benefit → strong; low certainty or a close call → conditional.


1. Simple explanation

You've done the hard analytical work: you pooled the trials (a Meta-Analysis: Pooling Studies (Fixed vs Random Effects)) and rated how much to trust each result (GRADE: Rating Certainty of Evidence). Now someone has to actually decide: should the guideline tell doctors to use this treatment or not — and how forcefully?

Two things bridge that gap.

First, the Summary of Findings (SoF) table. Think of it as the nutrition label on the side of the evidence box. Every critical outcome gets one row, and the columns tell you, at a glance: how much evidence there is, how much to trust it, the relative effect (a ratio like "30% fewer events"), and the absolute effect ("that's 3 fewer per 100 people"). The relative number sounds dramatic; the absolute number tells you if it actually matters for a real patient.

Second, the Evidence-to-Decision (EtD) framework. A single strong result doesn't automatically become a strong recommendation. A panel has to weigh benefits against harms, factor in cost, equity, what patients actually value, and how certain the evidence is. EtD is a checklist that forces every one of those considerations into the open, then produces a recommendation with two parts: a direction (for or against) and a strength (strong or conditional).

Analogy — buying a car for a family. The SoF table is the spec sheet: mileage, safety rating (with a confidence stamp), price, cargo space — one row per thing you care about. The EtD is the family meeting where you decide. A great spec sheet isn't enough: you weigh benefits vs cost, whether it fits everyone (equity), whether the family actually wants an SUV (values), and how sure you are the reviews are trustworthy (certainty). If it's clearly the best on every axis and everyone agrees, you make a strong call ("we're buying it"). If it's a close trade-off that depends on taste, you make a conditional one ("lean toward it, but it depends what you prioritize"). And if the spec sheet says one thing but the family strongly feels another — that discordance is exactly when you slow down and re-discuss.


2. Diagram

   FROM POOLED EVIDENCE  ->  A SCORECARD  ->  A DECISION

   [1] SUMMARY OF FINDINGS (SoF) TABLE — one ROW per critical outcome
   ---------------------------------------------------------------
   Outcome     | № participants | Certainty | Relative effect | Absolute effect
               | (studies)      | (GRADE)   | (RR/OR, 95% CI) | (with vs without, diff)
   -----------------------------------------------------------------------------
   Mortality   | 4200 (6 RCTs)  | HIGH      | RR 0.75         | 80 -> 60 /1000
               |                |           | (0.66-0.85)     | = 20 fewer /1000
   Serious AE  | 3800 (5 RCTs)  | MODERATE  | RR 1.20         | 30 -> 36 /1000
               |                |           | (0.95-1.52)     | = 6 more /1000

   [2] EVIDENCE-TO-DECISION (EtD) — weigh explicit criteria
   -------------------------------------------------------
       problem priority ....... is this an important problem?
       benefits vs harms ...... net benefit? (from the SoF rows)
       certainty of evidence .. how sure? (GRADE)
       values / preferences ... do patients want this outcome trade-off?
       resources / cost ....... affordable / cost-effective?
       equity ................. does it widen or narrow disparities?
       acceptability .......... will stakeholders accept it?
       feasibility ............ can it actually be implemented?
                 |
                 v
   [3] RECOMMENDATION = DIRECTION  x  STRENGTH
   -------------------------------------------
       DIRECTION : FOR  or  AGAINST
       STRENGTH  : STRONG      (benefits clearly outweigh harms; most patients would choose it)
                   CONDITIONAL (trade-offs close; depends on values)  [a.k.a. "weak"]

       high certainty + clear net benefit ........ -> STRONG
       low certainty OR close trade-off .......... -> CONDITIONAL

   [4] PANEL CONSENSUS (how the panel actually agrees)
   ---------------------------------------------------
       Delphi-style voting rounds -> threshold (e.g. >= 70-80% agreement) = consensus
       DISCORDANCE (evidence vs panel disagree, or threshold not met) -> FLAG + REVOTE

3. How it works

3.1 The Summary of Findings table: rows and columns

The SoF table is the standard hand-off from analysis to decision. The structure is fixed:

ColumnWhat it holdsWhy it's there
Outcomeone critical outcome per rowkeeps the decision focused on what matters
№ of participants (studies)e.g., "4200 (6 RCTs)"how much evidence backs this row
Certainty (GRADE)HIGH / MODERATE / LOW / VERY LOWhow much to trust this row (see GRADE: Rating Certainty of Evidence)
Relative effectRR or OR with 95% CIthe ratio — comparable across baseline risks
Absolute effectrisk with vs without, and the differencewhat it means for 1000 real patients

One row per critical outcome — not per study, not per subgroup. You choose the handful of outcomes that actually drive the decision (usually the key benefit and the key harm) before you look at results, so you can't cherry-pick.

3.2 Relative vs absolute effect — why you need both

This is the most-tested idea in the whole topic:

   RELATIVE effect (RR/OR): stable across populations, but hides baseline risk.
   ABSOLUTE effect: baseline-dependent, but tells you if it MATTERS.

   Example: RR 0.50 ("halves the risk!") sounds identical in two settings:
      high-risk group : 200/1000 -> 100/1000  = 100 FEWER per 1000   (huge)
      low-risk group  :   2/1000 ->   1/1000  =   1 FEWER per 1000   (tiny)
   Same relative effect, 100x difference in what a patient actually gains.

The absolute effect is baseline_risk × (1 − RR) fewer (or more) events. Always report both; the relative number is for comparability, the absolute number is for the decision.

3.3 The Evidence-to-Decision framework: criteria

EtD makes the panel address every relevant consideration, on the record:

CriterionThe questionFeeds…
Problem priorityIs the problem important/common enough to act on?whether to bother
Benefits vs harmsDoes net benefit favor the option?direction
Certainty of evidenceHow sure are we (GRADE)?strength
Values / preferencesDo patients value this outcome trade-off, and is that consistent?strength/direction
Resources / costAffordable? cost-effective?strength, feasibility
EquityDoes it widen or narrow health disparities?direction/strength
AcceptabilityWill clinicians, patients, payers accept it?feasibility
FeasibilityCan it actually be delivered at scale?feasibility

The panel judges each, then synthesizes them into one recommendation.

3.4 Direction and strength

The output has exactly two parts:

   DIRECTION : FOR the intervention  or  AGAINST it
   STRENGTH  : STRONG  or  CONDITIONAL (weak)
STRONGCONDITIONAL (weak)
Meaningbenefits clearly outweigh harms (or vice versa)trade-offs are close / uncertain
Patientsalmost all would want this optionchoices vary with personal values
Wording"We recommend…""We suggest…"
Clinician actiondo it (few exceptions)shared decision-making per patient
Typical drivershigh certainty and clear net benefitlow certainty or a close call

The mapping to remember: high certainty + clear net benefit → strong; low certainty or close trade-off → conditional. (There are recognized exceptions — e.g., a strong recommendation from low-certainty evidence in a life-or-death situation — but the default logic is this.)

3.5 Panel consensus: Delphi voting, thresholds, discordance

A panel is a group of humans who must actually agree. The standard mechanism is Delphi-style voting: structured, often anonymous, rounds of voting with feedback and re-discussion between rounds. A voting threshold defines consensus — for example, ≥70% or ≥80% of panelists agreeing.

   Round 1: vote  -> 62% agree  (below 80% threshold)  -> discuss, share rationales
   Round 2: vote  -> 74% agree  (still below)          -> discuss again
   Round 3: vote  -> 85% agree  (>= 80%)               -> CONSENSUS reached

Discordance is the flag you watch for: either (a) the evidence and the panel judgment disagree (e.g., strong evidence of benefit but the panel votes against), or (b) the panel fails to reach the threshold. Discordance isn't an error to hide — it's a signal to flag, examine why, and revote (or record a minority position). Good process makes discordance visible instead of averaging it away.

3.6 Where deterministic code fits

EtD is a judgment-heavy process — values, equity, and acceptability are genuinely human calls. But two pieces are pure rules and belong in code: building the SoF row (absolute effect = baseline × relative effect is arithmetic) and mapping (certainty, net benefit, values) → strength via a fixed lookup. A robust general principle: let the panel/LLM supply the judgments (net benefit sign, values alignment, certainty level), and let deterministic code compute the derived numbers and apply the strength rule, so the same inputs always yield the same recommendation strength and every SoF number is reproducible.


4. The rules / worked example

4.1 The strength decision rule, compactly

   Inputs (judgments):
     certainty        in {HIGH, MODERATE, LOW, VERY LOW}
     net_benefit      in {clear_benefit, close, clear_harm}   (from SoF benefit vs harm)
     values_alignment in {consistent, variable}               (do most patients agree?)

   Rule:
     direction = FOR      if net_benefit == clear_benefit
                 AGAINST  if net_benefit == clear_harm
                 (if 'close', direction leans by the balance but strength will be CONDITIONAL)

     strength  = STRONG      if certainty in {HIGH, MODERATE}
                             and net_benefit is clear
                             and values_alignment == consistent
                 CONDITIONAL otherwise   (low certainty OR close call OR variable values)

4.2 Worked example — build a 2-row SoF, then walk EtD

A guideline panel evaluates a drug for preventing a serious complication.

Step A — build the SoF rows from the meta-analysis output (baseline risk = control-group risk):

   Outcome 1: serious complication (the BENEFIT)
     6 RCTs, 4200 participants
     RR = 0.75 (95% CI 0.66 - 0.85)
     baseline (control) risk = 80 per 1000
     absolute:  80 * 0.75 = 60 per 1000 with treatment
                80 -> 60  = 20 FEWER per 1000  (95% CI ~ 12 to 27 fewer)
     certainty: HIGH (consistent, low risk of bias, precise)

   Outcome 2: serious adverse event (the HARM)
     5 RCTs, 3800 participants
     RR = 1.20 (95% CI 0.95 - 1.52)
     baseline risk = 30 per 1000
     absolute:  30 * 1.20 = 36 per 1000 with treatment
                30 -> 36  = 6 MORE per 1000
     certainty: MODERATE (CI crosses 1.0 -> imprecision downgrade)
Outcome№ (studies)CertaintyRelative (95% CI)Absolute (per 1000)
Serious complication4200 (6 RCTs)HIGHRR 0.75 (0.66–0.85)80 → 60 = 20 fewer
Serious adverse event3800 (5 RCTs)MODERATERR 1.20 (0.95–1.52)30 → 36 = 6 more

Step B — walk the EtD:

   Problem priority   : the complication is serious & common      -> important
   Benefits vs harms  : 20 fewer serious complications vs 6 more adverse events per 1000
                        -> net benefit exists BUT the harm CI crosses 1.0 (may be no harm,
                           may be real) -> net benefit is real but not overwhelming -> "close-ish"
   Certainty          : benefit HIGH, harm MODERATE -> overall moderate confidence in the balance
   Values/preferences : patients differ — some fear the adverse event more than the complication
                        -> values_alignment = VARIABLE
   Resources          : drug is moderately expensive              -> some concern
   Equity/accept/feas : neutral / acceptable / feasible

   SYNTHESIS: direction = FOR (net benefit favors treatment)
              strength: values are VARIABLE and the harm is uncertain (close trade-off)
                        -> CONDITIONAL, not strong
   RECOMMENDATION: "We SUGGEST (conditional) offering the drug; the decision should reflect
                    how much the individual patient weighs avoiding the complication against
                    the possible adverse event."

Even though the benefit is HIGH certainty, the variable values and the uncertain harm pull the strength down to conditional. That's the rule working: strong needs certainty and a clear balance and consistent values.

4.3 Panel vote

   Round 1: 66% support the conditional-FOR wording  (threshold 80%) -> below
            discussion: a subgroup worries about the adverse event -> add a caveat
   Round 2: 88% support the revised wording          -> CONSENSUS
   Discordance check: evidence (net benefit) and panel (conditional FOR) AGREE -> no discordance.

5. Real code

Two deterministic pieces: sof_row() builds a Summary-of-Findings row from meta-analysis output (the arithmetic that turns a relative effect + baseline risk into an absolute effect), and recommendation_strength() applies the fixed strength rule. The panel supplies the judgments; the code guarantees the numbers and the mapping are reproducible.

"""Build a Summary-of-Findings row and derive recommendation strength — deterministically.

Judgments (certainty, net-benefit sign, values alignment) come from the panel/analysis.
The code only does arithmetic (absolute effect) and applies a fixed strength rule, so the
same inputs always give the same SoF numbers and the same recommendation strength.
"""
from dataclasses import dataclass


@dataclass
class SoFRow:
    outcome: str
    n_participants: int
    n_studies: int
    rr: float                  # relative risk (or OR) point estimate
    ci_low: float
    ci_high: float
    baseline_per_1000: float   # control-group risk per 1000
    certainty: str             # HIGH / MODERATE / LOW / VERY LOW


def sof_row(row: SoFRow) -> dict:
    """Turn a relative effect + baseline risk into the absolute effect (per 1000)."""
    with_treat = row.baseline_per_1000 * row.rr
    diff = with_treat - row.baseline_per_1000          # + = more events, - = fewer
    # propagate the CI to the absolute scale (same baseline)
    abs_low = row.baseline_per_1000 * row.ci_low - row.baseline_per_1000
    abs_high = row.baseline_per_1000 * row.ci_high - row.baseline_per_1000
    direction = "fewer" if diff < 0 else "more"
    return {
        "outcome": row.outcome,
        "participants_studies": f"{row.n_participants} ({row.n_studies} studies)",
        "certainty": row.certainty,
        "relative": f"RR {row.rr:.2f} ({row.ci_low:.2f}-{row.ci_high:.2f})",
        "absolute": (f"{row.baseline_per_1000:.0f} -> {with_treat:.0f} per 1000 "
                     f"= {abs(diff):.0f} {direction} "
                     f"(95% CI {abs(abs_low):.0f} to {abs(abs_high):.0f})"),
        "abs_diff_per_1000": round(diff, 1),
    }


_STRONG_CERTAINTY = {"HIGH", "MODERATE"}

def recommendation_strength(certainty: str, net_benefit: str, values_alignment: str) -> dict:
    """Fixed rule: STRONG needs solid certainty AND a clear balance AND consistent values.

    certainty        : HIGH / MODERATE / LOW / VERY LOW
    net_benefit      : 'clear_benefit' / 'close' / 'clear_harm'
    values_alignment : 'consistent' / 'variable'
    """
    if net_benefit == "clear_benefit":
        direction = "FOR"
    elif net_benefit == "clear_harm":
        direction = "AGAINST"
    else:                                   # 'close' -> lean but never strong
        direction = "CONDITIONAL-either-way"

    strong = (certainty in _STRONG_CERTAINTY
              and net_benefit in ("clear_benefit", "clear_harm")
              and values_alignment == "consistent")
    strength = "STRONG" if strong else "CONDITIONAL"

    reasons = []
    if certainty not in _STRONG_CERTAINTY:
        reasons.append("low/very-low certainty")
    if net_benefit == "close":
        reasons.append("close benefit-harm trade-off")
    if values_alignment == "variable":
        reasons.append("variable patient values")
    return {
        "direction": direction,
        "strength": strength,
        "verb": "recommend" if strength == "STRONG" else "suggest",
        "downgraded_because": reasons or ["none (strong criteria met)"],
    }


if __name__ == "__main__":
    benefit = sof_row(SoFRow("Serious complication", 4200, 6, 0.75, 0.66, 0.85, 80, "HIGH"))
    harm = sof_row(SoFRow("Serious adverse event", 3800, 5, 1.20, 0.95, 1.52, 30, "MODERATE"))
    print(benefit["absolute"])   # 80 -> 60 per 1000 = 20 fewer ...
    print(harm["absolute"])      # 30 -> 36 per 1000 = 6 more ...

    # Worked EtD: real benefit but variable values + uncertain harm -> CONDITIONAL FOR
    rec = recommendation_strength(certainty="HIGH",
                                  net_benefit="clear_benefit",
                                  values_alignment="variable")
    print(rec["strength"], rec["direction"], "-> we", rec["verb"],
          "| because:", rec["downgraded_because"])
    # CONDITIONAL FOR -> we suggest | because: ['variable patient values']

    # Contrast: high certainty, clear benefit, consistent values -> STRONG
    rec2 = recommendation_strength("HIGH", "clear_benefit", "consistent")
    print(rec2["strength"], rec2["direction"], "-> we", rec2["verb"])   # STRONG FOR -> we recommend

The load-bearing details: sof_row() computes the absolute effect as baseline × RR − baseline (the number that actually drives decisions), and recommendation_strength() gates STRONG behind all three conditions — solid certainty and a clear (non-close) balance and consistent values — so a single soft input (variable values, here) deterministically downgrades to CONDITIONAL with a recorded reason.


6. Real-world example

Scenario: a national panel decides whether to recommend a screening test.

The panel pools the evidence and builds a two-row SoF, then runs the EtD.

Outcome№ (studies)CertaintyRelative (95% CI)Absolute (per 1000 screened)
Deaths from the disease210,000 (4 RCTs)MODERATERR 0.80 (0.70–0.92)5.0 → 4.0 = 1 fewer
False-positive → invasive follow-up210,000 (4 RCTs)HIGH120 more per 1000
   Reading the SoF:
     - Benefit: screening cuts disease deaths by 20% RELATIVE — but the disease is rare, so
       ABSOLUTE benefit is ~1 fewer death per 1000 screened.
     - Harm: 120 per 1000 get a false positive leading to an invasive, anxiety-inducing
       follow-up. This is HIGH certainty (easy to count) and LARGE in absolute terms.

   The relative headline ('20% fewer deaths!') looks strong; the absolute trade-off
   (1 life saved vs 120 harmed by follow-up per 1000) is genuinely close.

EtD walk:

   Problem priority   : the disease is serious              -> important
   Benefits vs harms  : 1 fewer death vs 120 invasive follow-ups per 1000 -> CLOSE
   Certainty          : benefit MODERATE, harm HIGH         -> mixed
   Values/preferences : people differ enormously — some accept 120 scares to avoid 1 death,
                        others don't                          -> VARIABLE
   Resources          : screening program is costly at scale -> concern
   Equity             : uptake may be lower in underserved groups -> could widen disparity

   SYNTHESIS: direction FOR (small net benefit), but strength CONDITIONAL — close balance,
              variable values, cost and equity concerns.
   RECOMMENDATION: "We SUGGEST offering screening to eligible adults; the choice should be
                    a shared decision reflecting how the individual weighs a small mortality
                    benefit against a high chance of a false-positive workup."

   PANEL VOTE (Delphi, threshold 75%):
     Round 1: 58% -> below; a bloc argues the absolute benefit is too small to recommend at all
     Round 2: 79% -> consensus on the CONDITIONAL wording, with a documented minority who would
                     not recommend screening.
     DISCORDANCE: the numeric net benefit (favor) and the near-split panel are in tension ->
                  FLAGGED and recorded as a minority position rather than hidden.

The lesson: a dramatic relative effect collapsed into a close decision once the absolute harm and variable values were on the table — and the process surfaced the disagreement (discordance) instead of pretending unanimity. That transparency is the entire point of doing SoF + EtD instead of an expert just declaring an answer.


7. Interview questions companies actually ask

Q1 [easy] (HTA agencies, guideline developers) "What goes in a Summary of Findings table?"
  A One row per CRITICAL outcome, with columns: № of participants (studies), certainty (GRADE),
    relative effect (RR/OR with 95% CI), and absolute effect (risk with vs without + the
    difference). It's the one-page scorecard handed from analysis to decision.

Q2 [easy] (any EBM role) "Why report BOTH relative and absolute effect?"
  A The relative effect (RR/OR) is stable across populations but hides baseline risk; the absolute
    effect tells you if it MATTERS. RR 0.5 is '100 fewer per 1000' in a high-risk group but '1
    fewer per 1000' in a low-risk one. Same ratio, very different decision.

Q3 [medium] (guideline panels) "What are the two parts of a recommendation, and what do they
   mean?"
  A DIRECTION (for or against) and STRENGTH (strong vs conditional/weak). Strong = benefits clearly
    outweigh harms and almost all patients would choose it ('we recommend'). Conditional = close or
    uncertain trade-off, depends on values ('we suggest', shared decision-making).

Q4 [medium] (HTA, payers) "What pushes a recommendation from strong to conditional?"
  A Low/very-low certainty, a close benefit-harm balance, variable patient values, high cost, or
    equity/feasibility concerns. Default rule: high certainty + clear net benefit -> strong;
    low certainty OR a close trade-off -> conditional.

Q5 [medium] (evidence-to-decision roles) "Name the EtD criteria."
  A Problem priority, benefits vs harms, certainty of evidence, values/preferences, resources/cost,
    equity, acceptability, feasibility. The panel judges each explicitly, then synthesizes them
    into direction + strength.

Q6 [hard] (methodologists) "Can you have high-certainty evidence but only a conditional
   recommendation?"
  A Yes. Certainty is just one EtD criterion. A high-certainty benefit can still yield a conditional
    recommendation if the benefit-harm balance is close, patient values vary a lot, or cost/equity
    weigh against it. Certainty gates strength but doesn't determine it alone.

Q7 [hard] (guideline organizations) "How does a panel reach consensus, and what is discordance?"
  A Delphi-style voting: structured, often anonymous rounds with feedback and re-discussion between
    them, against a threshold (e.g., >=70-80% agreement) that defines consensus. DISCORDANCE is when
    the evidence and the panel judgment disagree, or the panel can't hit the threshold — you FLAG it,
    examine why, and revote or record a minority position rather than averaging it away.

Q8 [hard] (HTA, screening programs) "A screening test has RR 0.80 for mortality but you only
   'suggest' it. Explain."
  A The 20% relative reduction is small in ABSOLUTE terms when the disease is rare (~1 fewer death
    per 1000), while harms (false positives -> invasive follow-up) are large and certain (e.g., 120
    more per 1000). That's a close trade-off with variable values -> conditional, shared-decision.

Q9 [medium] (tooling / platform roles) "Which parts of EtD should be code vs human judgment?"
  A Human/panel: net-benefit sign, values alignment, equity, acceptability — genuine judgments.
    Code: computing the absolute effect from RR x baseline for the SoF, and applying the fixed
    (certainty, net-benefit, values) -> strength rule, so the numbers and the mapping are
    reproducible and auditable.

Q10 [medium] (any EBM interview) "Why one row per CRITICAL outcome, chosen in advance?"
  A To keep the decision focused on what matters (usually the key benefit and key harm) and to
    prevent cherry-picking outcomes after seeing results. Pre-specifying the critical outcomes is
    part of an honest, non-gameable process.

8. When to use / tradeoffs

   USE SoF + EtD when:
     ✓ you must turn pooled evidence + certainty into an actual recommendation
     ✓ a guideline/panel needs a transparent, auditable evidence-to-decision trail
     ✓ decisions hinge on trade-offs (benefit vs harm, cost, equity, values), not just a p-value
     ✓ you need to communicate both HOW BIG (absolute) and HOW SURE (certainty) an effect is

   STRENGTHS:
     • forces every decision factor into the open (no hidden reasoning)
     • separates direction from strength -> honest 'we suggest' vs 'we recommend'
     • absolute effects stop dramatic relative numbers from over-driving decisions
     • Delphi voting + discordance flags make disagreement visible instead of averaged

   HONEST LIMITS:
     ✗ EtD is judgment-heavy — values, equity, acceptability are genuinely subjective
     ✗ panel composition biases outcomes (who's in the room matters); manage conflicts of interest
     ✗ the strength rule has recognized EXCEPTIONS (e.g., strong recs from low-certainty evidence
       in life-or-death cases) — don't apply the default rule blindly
     ✗ voting thresholds are conventions, not laws; a bare-threshold 'consensus' can mask a deep split
     ✗ absolute effects depend on the assumed BASELINE risk — wrong baseline -> misleading row
     ✗ SoF/EtD structure the reasoning; they don't fix bad underlying evidence

   RULE OF THUMB: pre-specify critical outcomes; always show absolute alongside relative; let CODE
   compute the SoF numbers and the strength mapping; and record discordance and minority views
   instead of forcing false unanimity.

  • The Summary of Findings (SoF) table is the scorecard: one row per critical outcome; columns for № participants (studies), certainty (GRADE), relative effect (RR/OR, CI), and absolute effect (with vs without + difference).
  • Report relative and absolute effects together — relative is comparable, absolute tells you if it matters (absolute = baseline × RR − baseline).
  • The Evidence-to-Decision (EtD) framework weighs explicit criteria — problem priority, benefits vs harms, certainty, values, resources, equity, acceptability, feasibility.
  • A recommendation = direction (for/against) × strength (STRONG vs CONDITIONAL/weak). High certainty + clear net benefit → strong; low certainty or close trade-off → conditional.
  • Panels agree via Delphi-style voting to a threshold (e.g., ≥70–80%); discordance (evidence vs panel disagree, or threshold not met) is flagged and revoted, not hidden.
  • The SoF arithmetic and the strength mapping are fixed rules — put them in deterministic code; keep the values/equity/acceptability judgments with the humans.

Related: GRADE: Rating Certainty of Evidence · Risk of Bias: RoB2, ROBINS-I, AMSTAR-2 · Meta-Analysis: Pooling Studies (Fixed vs Random Effects)

Resources

Runnable notebook

Run it end to end — the mock model needs no API key; add your own key for the real Claude section.

Open In Colab