TL;DR — Once you've pooled the evidence and rated its certainty, two structured artifacts turn numbers into a decision. A Summary of Findings (SoF) table is the one-page scorecard: one row per critical outcome, columns for № of participants (studies), certainty (GRADE), relative effect (RR/OR with CI), and absolute effect (risk with vs without + the difference). Then the Evidence-to-Decision (EtD) framework walks a panel from that evidence to a recommendation through explicit criteria — problem priority, benefits vs harms, certainty, values/preferences, resources, equity, acceptability, feasibility — and outputs a direction (for/against) and a strength: STRONG (benefits clearly outweigh harms, most patients would choose it) or CONDITIONAL/weak (trade-offs are close, depends on values). Panels reach agreement by Delphi-style voting against a threshold (e.g., ≥70–80%); when evidence and panel disagree, or the threshold isn't met, that's discordance — flag it and revote. High certainty + clear net benefit → strong; low certainty or a close call → conditional.
1. Simple explanation
You've done the hard analytical work: you pooled the trials (a Meta-Analysis: Pooling Studies (Fixed vs Random Effects)) and rated how much to trust each result (GRADE: Rating Certainty of Evidence). Now someone has to actually decide: should the guideline tell doctors to use this treatment or not — and how forcefully?
Two things bridge that gap.
First, the Summary of Findings (SoF) table. Think of it as the nutrition label on the side of the evidence box. Every critical outcome gets one row, and the columns tell you, at a glance: how much evidence there is, how much to trust it, the relative effect (a ratio like "30% fewer events"), and the absolute effect ("that's 3 fewer per 100 people"). The relative number sounds dramatic; the absolute number tells you if it actually matters for a real patient.
Second, the Evidence-to-Decision (EtD) framework. A single strong result doesn't automatically become a strong recommendation. A panel has to weigh benefits against harms, factor in cost, equity, what patients actually value, and how certain the evidence is. EtD is a checklist that forces every one of those considerations into the open, then produces a recommendation with two parts: a direction (for or against) and a strength (strong or conditional).
Analogy — buying a car for a family. The SoF table is the spec sheet: mileage, safety rating (with a confidence stamp), price, cargo space — one row per thing you care about. The EtD is the family meeting where you decide. A great spec sheet isn't enough: you weigh benefits vs cost, whether it fits everyone (equity), whether the family actually wants an SUV (values), and how sure you are the reviews are trustworthy (certainty). If it's clearly the best on every axis and everyone agrees, you make a strong call ("we're buying it"). If it's a close trade-off that depends on taste, you make a conditional one ("lean toward it, but it depends what you prioritize"). And if the spec sheet says one thing but the family strongly feels another — that discordance is exactly when you slow down and re-discuss.
2. Diagram
FROM POOLED EVIDENCE -> A SCORECARD -> A DECISION
[1] SUMMARY OF FINDINGS (SoF) TABLE — one ROW per critical outcome
---------------------------------------------------------------
Outcome | № participants | Certainty | Relative effect | Absolute effect
| (studies) | (GRADE) | (RR/OR, 95% CI) | (with vs without, diff)
-----------------------------------------------------------------------------
Mortality | 4200 (6 RCTs) | HIGH | RR 0.75 | 80 -> 60 /1000
| | | (0.66-0.85) | = 20 fewer /1000
Serious AE | 3800 (5 RCTs) | MODERATE | RR 1.20 | 30 -> 36 /1000
| | | (0.95-1.52) | = 6 more /1000
[2] EVIDENCE-TO-DECISION (EtD) — weigh explicit criteria
-------------------------------------------------------
problem priority ....... is this an important problem?
benefits vs harms ...... net benefit? (from the SoF rows)
certainty of evidence .. how sure? (GRADE)
values / preferences ... do patients want this outcome trade-off?
resources / cost ....... affordable / cost-effective?
equity ................. does it widen or narrow disparities?
acceptability .......... will stakeholders accept it?
feasibility ............ can it actually be implemented?
|
v
[3] RECOMMENDATION = DIRECTION x STRENGTH
-------------------------------------------
DIRECTION : FOR or AGAINST
STRENGTH : STRONG (benefits clearly outweigh harms; most patients would choose it)
CONDITIONAL (trade-offs close; depends on values) [a.k.a. "weak"]
high certainty + clear net benefit ........ -> STRONG
low certainty OR close trade-off .......... -> CONDITIONAL
[4] PANEL CONSENSUS (how the panel actually agrees)
---------------------------------------------------
Delphi-style voting rounds -> threshold (e.g. >= 70-80% agreement) = consensus
DISCORDANCE (evidence vs panel disagree, or threshold not met) -> FLAG + REVOTE
3. How it works
3.1 The Summary of Findings table: rows and columns
The SoF table is the standard hand-off from analysis to decision. The structure is fixed:
| Column | What it holds | Why it's there |
|---|---|---|
| Outcome | one critical outcome per row | keeps the decision focused on what matters |
| № of participants (studies) | e.g., "4200 (6 RCTs)" | how much evidence backs this row |
| Certainty (GRADE) | HIGH / MODERATE / LOW / VERY LOW | how much to trust this row (see GRADE: Rating Certainty of Evidence) |
| Relative effect | RR or OR with 95% CI | the ratio — comparable across baseline risks |
| Absolute effect | risk with vs without, and the difference | what it means for 1000 real patients |
One row per critical outcome — not per study, not per subgroup. You choose the handful of outcomes that actually drive the decision (usually the key benefit and the key harm) before you look at results, so you can't cherry-pick.
3.2 Relative vs absolute effect — why you need both
This is the most-tested idea in the whole topic:
RELATIVE effect (RR/OR): stable across populations, but hides baseline risk.
ABSOLUTE effect: baseline-dependent, but tells you if it MATTERS.
Example: RR 0.50 ("halves the risk!") sounds identical in two settings:
high-risk group : 200/1000 -> 100/1000 = 100 FEWER per 1000 (huge)
low-risk group : 2/1000 -> 1/1000 = 1 FEWER per 1000 (tiny)
Same relative effect, 100x difference in what a patient actually gains.
The absolute effect is baseline_risk × (1 − RR) fewer (or more) events. Always report both; the relative number is for comparability, the absolute number is for the decision.
3.3 The Evidence-to-Decision framework: criteria
EtD makes the panel address every relevant consideration, on the record:
| Criterion | The question | Feeds… |
|---|---|---|
| Problem priority | Is the problem important/common enough to act on? | whether to bother |
| Benefits vs harms | Does net benefit favor the option? | direction |
| Certainty of evidence | How sure are we (GRADE)? | strength |
| Values / preferences | Do patients value this outcome trade-off, and is that consistent? | strength/direction |
| Resources / cost | Affordable? cost-effective? | strength, feasibility |
| Equity | Does it widen or narrow health disparities? | direction/strength |
| Acceptability | Will clinicians, patients, payers accept it? | feasibility |
| Feasibility | Can it actually be delivered at scale? | feasibility |
The panel judges each, then synthesizes them into one recommendation.
3.4 Direction and strength
The output has exactly two parts:
DIRECTION : FOR the intervention or AGAINST it
STRENGTH : STRONG or CONDITIONAL (weak)
| STRONG | CONDITIONAL (weak) | |
|---|---|---|
| Meaning | benefits clearly outweigh harms (or vice versa) | trade-offs are close / uncertain |
| Patients | almost all would want this option | choices vary with personal values |
| Wording | "We recommend…" | "We suggest…" |
| Clinician action | do it (few exceptions) | shared decision-making per patient |
| Typical drivers | high certainty and clear net benefit | low certainty or a close call |
The mapping to remember: high certainty + clear net benefit → strong; low certainty or close trade-off → conditional. (There are recognized exceptions — e.g., a strong recommendation from low-certainty evidence in a life-or-death situation — but the default logic is this.)
3.5 Panel consensus: Delphi voting, thresholds, discordance
A panel is a group of humans who must actually agree. The standard mechanism is Delphi-style voting: structured, often anonymous, rounds of voting with feedback and re-discussion between rounds. A voting threshold defines consensus — for example, ≥70% or ≥80% of panelists agreeing.
Round 1: vote -> 62% agree (below 80% threshold) -> discuss, share rationales
Round 2: vote -> 74% agree (still below) -> discuss again
Round 3: vote -> 85% agree (>= 80%) -> CONSENSUS reached
Discordance is the flag you watch for: either (a) the evidence and the panel judgment disagree (e.g., strong evidence of benefit but the panel votes against), or (b) the panel fails to reach the threshold. Discordance isn't an error to hide — it's a signal to flag, examine why, and revote (or record a minority position). Good process makes discordance visible instead of averaging it away.
3.6 Where deterministic code fits
EtD is a judgment-heavy process — values, equity, and acceptability are genuinely human calls. But two pieces are pure rules and belong in code: building the SoF row (absolute effect = baseline × relative effect is arithmetic) and mapping (certainty, net benefit, values) → strength via a fixed lookup. A robust general principle: let the panel/LLM supply the judgments (net benefit sign, values alignment, certainty level), and let deterministic code compute the derived numbers and apply the strength rule, so the same inputs always yield the same recommendation strength and every SoF number is reproducible.
4. The rules / worked example
4.1 The strength decision rule, compactly
Inputs (judgments):
certainty in {HIGH, MODERATE, LOW, VERY LOW}
net_benefit in {clear_benefit, close, clear_harm} (from SoF benefit vs harm)
values_alignment in {consistent, variable} (do most patients agree?)
Rule:
direction = FOR if net_benefit == clear_benefit
AGAINST if net_benefit == clear_harm
(if 'close', direction leans by the balance but strength will be CONDITIONAL)
strength = STRONG if certainty in {HIGH, MODERATE}
and net_benefit is clear
and values_alignment == consistent
CONDITIONAL otherwise (low certainty OR close call OR variable values)
4.2 Worked example — build a 2-row SoF, then walk EtD
A guideline panel evaluates a drug for preventing a serious complication.
Step A — build the SoF rows from the meta-analysis output (baseline risk = control-group risk):
Outcome 1: serious complication (the BENEFIT)
6 RCTs, 4200 participants
RR = 0.75 (95% CI 0.66 - 0.85)
baseline (control) risk = 80 per 1000
absolute: 80 * 0.75 = 60 per 1000 with treatment
80 -> 60 = 20 FEWER per 1000 (95% CI ~ 12 to 27 fewer)
certainty: HIGH (consistent, low risk of bias, precise)
Outcome 2: serious adverse event (the HARM)
5 RCTs, 3800 participants
RR = 1.20 (95% CI 0.95 - 1.52)
baseline risk = 30 per 1000
absolute: 30 * 1.20 = 36 per 1000 with treatment
30 -> 36 = 6 MORE per 1000
certainty: MODERATE (CI crosses 1.0 -> imprecision downgrade)
| Outcome | № (studies) | Certainty | Relative (95% CI) | Absolute (per 1000) |
|---|---|---|---|---|
| Serious complication | 4200 (6 RCTs) | HIGH | RR 0.75 (0.66–0.85) | 80 → 60 = 20 fewer |
| Serious adverse event | 3800 (5 RCTs) | MODERATE | RR 1.20 (0.95–1.52) | 30 → 36 = 6 more |
Step B — walk the EtD:
Problem priority : the complication is serious & common -> important
Benefits vs harms : 20 fewer serious complications vs 6 more adverse events per 1000
-> net benefit exists BUT the harm CI crosses 1.0 (may be no harm,
may be real) -> net benefit is real but not overwhelming -> "close-ish"
Certainty : benefit HIGH, harm MODERATE -> overall moderate confidence in the balance
Values/preferences : patients differ — some fear the adverse event more than the complication
-> values_alignment = VARIABLE
Resources : drug is moderately expensive -> some concern
Equity/accept/feas : neutral / acceptable / feasible
SYNTHESIS: direction = FOR (net benefit favors treatment)
strength: values are VARIABLE and the harm is uncertain (close trade-off)
-> CONDITIONAL, not strong
RECOMMENDATION: "We SUGGEST (conditional) offering the drug; the decision should reflect
how much the individual patient weighs avoiding the complication against
the possible adverse event."
Even though the benefit is HIGH certainty, the variable values and the uncertain harm pull the strength down to conditional. That's the rule working: strong needs certainty and a clear balance and consistent values.
4.3 Panel vote
Round 1: 66% support the conditional-FOR wording (threshold 80%) -> below
discussion: a subgroup worries about the adverse event -> add a caveat
Round 2: 88% support the revised wording -> CONSENSUS
Discordance check: evidence (net benefit) and panel (conditional FOR) AGREE -> no discordance.
5. Real code
Two deterministic pieces: sof_row() builds a Summary-of-Findings row from meta-analysis output (the arithmetic that turns a relative effect + baseline risk into an absolute effect), and recommendation_strength() applies the fixed strength rule. The panel supplies the judgments; the code guarantees the numbers and the mapping are reproducible.
"""Build a Summary-of-Findings row and derive recommendation strength — deterministically.
Judgments (certainty, net-benefit sign, values alignment) come from the panel/analysis.
The code only does arithmetic (absolute effect) and applies a fixed strength rule, so the
same inputs always give the same SoF numbers and the same recommendation strength.
"""
from dataclasses import dataclass
@dataclass
class SoFRow:
outcome: str
n_participants: int
n_studies: int
rr: float # relative risk (or OR) point estimate
ci_low: float
ci_high: float
baseline_per_1000: float # control-group risk per 1000
certainty: str # HIGH / MODERATE / LOW / VERY LOW
def sof_row(row: SoFRow) -> dict:
"""Turn a relative effect + baseline risk into the absolute effect (per 1000)."""
with_treat = row.baseline_per_1000 * row.rr
diff = with_treat - row.baseline_per_1000 # + = more events, - = fewer
# propagate the CI to the absolute scale (same baseline)
abs_low = row.baseline_per_1000 * row.ci_low - row.baseline_per_1000
abs_high = row.baseline_per_1000 * row.ci_high - row.baseline_per_1000
direction = "fewer" if diff < 0 else "more"
return {
"outcome": row.outcome,
"participants_studies": f"{row.n_participants} ({row.n_studies} studies)",
"certainty": row.certainty,
"relative": f"RR {row.rr:.2f} ({row.ci_low:.2f}-{row.ci_high:.2f})",
"absolute": (f"{row.baseline_per_1000:.0f} -> {with_treat:.0f} per 1000 "
f"= {abs(diff):.0f} {direction} "
f"(95% CI {abs(abs_low):.0f} to {abs(abs_high):.0f})"),
"abs_diff_per_1000": round(diff, 1),
}
_STRONG_CERTAINTY = {"HIGH", "MODERATE"}
def recommendation_strength(certainty: str, net_benefit: str, values_alignment: str) -> dict:
"""Fixed rule: STRONG needs solid certainty AND a clear balance AND consistent values.
certainty : HIGH / MODERATE / LOW / VERY LOW
net_benefit : 'clear_benefit' / 'close' / 'clear_harm'
values_alignment : 'consistent' / 'variable'
"""
if net_benefit == "clear_benefit":
direction = "FOR"
elif net_benefit == "clear_harm":
direction = "AGAINST"
else: # 'close' -> lean but never strong
direction = "CONDITIONAL-either-way"
strong = (certainty in _STRONG_CERTAINTY
and net_benefit in ("clear_benefit", "clear_harm")
and values_alignment == "consistent")
strength = "STRONG" if strong else "CONDITIONAL"
reasons = []
if certainty not in _STRONG_CERTAINTY:
reasons.append("low/very-low certainty")
if net_benefit == "close":
reasons.append("close benefit-harm trade-off")
if values_alignment == "variable":
reasons.append("variable patient values")
return {
"direction": direction,
"strength": strength,
"verb": "recommend" if strength == "STRONG" else "suggest",
"downgraded_because": reasons or ["none (strong criteria met)"],
}
if __name__ == "__main__":
benefit = sof_row(SoFRow("Serious complication", 4200, 6, 0.75, 0.66, 0.85, 80, "HIGH"))
harm = sof_row(SoFRow("Serious adverse event", 3800, 5, 1.20, 0.95, 1.52, 30, "MODERATE"))
print(benefit["absolute"]) # 80 -> 60 per 1000 = 20 fewer ...
print(harm["absolute"]) # 30 -> 36 per 1000 = 6 more ...
# Worked EtD: real benefit but variable values + uncertain harm -> CONDITIONAL FOR
rec = recommendation_strength(certainty="HIGH",
net_benefit="clear_benefit",
values_alignment="variable")
print(rec["strength"], rec["direction"], "-> we", rec["verb"],
"| because:", rec["downgraded_because"])
# CONDITIONAL FOR -> we suggest | because: ['variable patient values']
# Contrast: high certainty, clear benefit, consistent values -> STRONG
rec2 = recommendation_strength("HIGH", "clear_benefit", "consistent")
print(rec2["strength"], rec2["direction"], "-> we", rec2["verb"]) # STRONG FOR -> we recommend
The load-bearing details: sof_row() computes the absolute effect as baseline × RR − baseline (the number that actually drives decisions), and recommendation_strength() gates STRONG behind all three conditions — solid certainty and a clear (non-close) balance and consistent values — so a single soft input (variable values, here) deterministically downgrades to CONDITIONAL with a recorded reason.
6. Real-world example
Scenario: a national panel decides whether to recommend a screening test.
The panel pools the evidence and builds a two-row SoF, then runs the EtD.
| Outcome | № (studies) | Certainty | Relative (95% CI) | Absolute (per 1000 screened) |
|---|---|---|---|---|
| Deaths from the disease | 210,000 (4 RCTs) | MODERATE | RR 0.80 (0.70–0.92) | 5.0 → 4.0 = 1 fewer |
| False-positive → invasive follow-up | 210,000 (4 RCTs) | HIGH | — | 120 more per 1000 |
Reading the SoF:
- Benefit: screening cuts disease deaths by 20% RELATIVE — but the disease is rare, so
ABSOLUTE benefit is ~1 fewer death per 1000 screened.
- Harm: 120 per 1000 get a false positive leading to an invasive, anxiety-inducing
follow-up. This is HIGH certainty (easy to count) and LARGE in absolute terms.
The relative headline ('20% fewer deaths!') looks strong; the absolute trade-off
(1 life saved vs 120 harmed by follow-up per 1000) is genuinely close.
EtD walk:
Problem priority : the disease is serious -> important
Benefits vs harms : 1 fewer death vs 120 invasive follow-ups per 1000 -> CLOSE
Certainty : benefit MODERATE, harm HIGH -> mixed
Values/preferences : people differ enormously — some accept 120 scares to avoid 1 death,
others don't -> VARIABLE
Resources : screening program is costly at scale -> concern
Equity : uptake may be lower in underserved groups -> could widen disparity
SYNTHESIS: direction FOR (small net benefit), but strength CONDITIONAL — close balance,
variable values, cost and equity concerns.
RECOMMENDATION: "We SUGGEST offering screening to eligible adults; the choice should be
a shared decision reflecting how the individual weighs a small mortality
benefit against a high chance of a false-positive workup."
PANEL VOTE (Delphi, threshold 75%):
Round 1: 58% -> below; a bloc argues the absolute benefit is too small to recommend at all
Round 2: 79% -> consensus on the CONDITIONAL wording, with a documented minority who would
not recommend screening.
DISCORDANCE: the numeric net benefit (favor) and the near-split panel are in tension ->
FLAGGED and recorded as a minority position rather than hidden.
The lesson: a dramatic relative effect collapsed into a close decision once the absolute harm and variable values were on the table — and the process surfaced the disagreement (discordance) instead of pretending unanimity. That transparency is the entire point of doing SoF + EtD instead of an expert just declaring an answer.
7. Interview questions companies actually ask
Q1 [easy] (HTA agencies, guideline developers) "What goes in a Summary of Findings table?"
A One row per CRITICAL outcome, with columns: № of participants (studies), certainty (GRADE),
relative effect (RR/OR with 95% CI), and absolute effect (risk with vs without + the
difference). It's the one-page scorecard handed from analysis to decision.
Q2 [easy] (any EBM role) "Why report BOTH relative and absolute effect?"
A The relative effect (RR/OR) is stable across populations but hides baseline risk; the absolute
effect tells you if it MATTERS. RR 0.5 is '100 fewer per 1000' in a high-risk group but '1
fewer per 1000' in a low-risk one. Same ratio, very different decision.
Q3 [medium] (guideline panels) "What are the two parts of a recommendation, and what do they
mean?"
A DIRECTION (for or against) and STRENGTH (strong vs conditional/weak). Strong = benefits clearly
outweigh harms and almost all patients would choose it ('we recommend'). Conditional = close or
uncertain trade-off, depends on values ('we suggest', shared decision-making).
Q4 [medium] (HTA, payers) "What pushes a recommendation from strong to conditional?"
A Low/very-low certainty, a close benefit-harm balance, variable patient values, high cost, or
equity/feasibility concerns. Default rule: high certainty + clear net benefit -> strong;
low certainty OR a close trade-off -> conditional.
Q5 [medium] (evidence-to-decision roles) "Name the EtD criteria."
A Problem priority, benefits vs harms, certainty of evidence, values/preferences, resources/cost,
equity, acceptability, feasibility. The panel judges each explicitly, then synthesizes them
into direction + strength.
Q6 [hard] (methodologists) "Can you have high-certainty evidence but only a conditional
recommendation?"
A Yes. Certainty is just one EtD criterion. A high-certainty benefit can still yield a conditional
recommendation if the benefit-harm balance is close, patient values vary a lot, or cost/equity
weigh against it. Certainty gates strength but doesn't determine it alone.
Q7 [hard] (guideline organizations) "How does a panel reach consensus, and what is discordance?"
A Delphi-style voting: structured, often anonymous rounds with feedback and re-discussion between
them, against a threshold (e.g., >=70-80% agreement) that defines consensus. DISCORDANCE is when
the evidence and the panel judgment disagree, or the panel can't hit the threshold — you FLAG it,
examine why, and revote or record a minority position rather than averaging it away.
Q8 [hard] (HTA, screening programs) "A screening test has RR 0.80 for mortality but you only
'suggest' it. Explain."
A The 20% relative reduction is small in ABSOLUTE terms when the disease is rare (~1 fewer death
per 1000), while harms (false positives -> invasive follow-up) are large and certain (e.g., 120
more per 1000). That's a close trade-off with variable values -> conditional, shared-decision.
Q9 [medium] (tooling / platform roles) "Which parts of EtD should be code vs human judgment?"
A Human/panel: net-benefit sign, values alignment, equity, acceptability — genuine judgments.
Code: computing the absolute effect from RR x baseline for the SoF, and applying the fixed
(certainty, net-benefit, values) -> strength rule, so the numbers and the mapping are
reproducible and auditable.
Q10 [medium] (any EBM interview) "Why one row per CRITICAL outcome, chosen in advance?"
A To keep the decision focused on what matters (usually the key benefit and key harm) and to
prevent cherry-picking outcomes after seeing results. Pre-specifying the critical outcomes is
part of an honest, non-gameable process.
8. When to use / tradeoffs
USE SoF + EtD when:
✓ you must turn pooled evidence + certainty into an actual recommendation
✓ a guideline/panel needs a transparent, auditable evidence-to-decision trail
✓ decisions hinge on trade-offs (benefit vs harm, cost, equity, values), not just a p-value
✓ you need to communicate both HOW BIG (absolute) and HOW SURE (certainty) an effect is
STRENGTHS:
• forces every decision factor into the open (no hidden reasoning)
• separates direction from strength -> honest 'we suggest' vs 'we recommend'
• absolute effects stop dramatic relative numbers from over-driving decisions
• Delphi voting + discordance flags make disagreement visible instead of averaged
HONEST LIMITS:
✗ EtD is judgment-heavy — values, equity, acceptability are genuinely subjective
✗ panel composition biases outcomes (who's in the room matters); manage conflicts of interest
✗ the strength rule has recognized EXCEPTIONS (e.g., strong recs from low-certainty evidence
in life-or-death cases) — don't apply the default rule blindly
✗ voting thresholds are conventions, not laws; a bare-threshold 'consensus' can mask a deep split
✗ absolute effects depend on the assumed BASELINE risk — wrong baseline -> misleading row
✗ SoF/EtD structure the reasoning; they don't fix bad underlying evidence
RULE OF THUMB: pre-specify critical outcomes; always show absolute alongside relative; let CODE
compute the SoF numbers and the strength mapping; and record discordance and minority views
instead of forcing false unanimity.
9. Summary + related articles
- The Summary of Findings (SoF) table is the scorecard: one row per critical outcome; columns for № participants (studies), certainty (GRADE), relative effect (RR/OR, CI), and absolute effect (with vs without + difference).
- Report relative and absolute effects together — relative is comparable, absolute tells you if it matters (
absolute = baseline × RR − baseline). - The Evidence-to-Decision (EtD) framework weighs explicit criteria — problem priority, benefits vs harms, certainty, values, resources, equity, acceptability, feasibility.
- A recommendation = direction (for/against) × strength (STRONG vs CONDITIONAL/weak). High certainty + clear net benefit → strong; low certainty or close trade-off → conditional.
- Panels agree via Delphi-style voting to a threshold (e.g., ≥70–80%); discordance (evidence vs panel disagree, or threshold not met) is flagged and revoted, not hidden.
- The SoF arithmetic and the strength mapping are fixed rules — put them in deterministic code; keep the values/equity/acceptability judgments with the humans.
Related: GRADE: Rating Certainty of Evidence · Risk of Bias: RoB2, ROBINS-I, AMSTAR-2 · Meta-Analysis: Pooling Studies (Fixed vs Random Effects)
Resources
- GRADE Handbook — Summary of Findings & Evidence-to-Decision — https://gdt.gradepro.org/app/handbook/handbook.html
- Alonso-Coello et al. — GRADE Evidence to Decision (EtD) frameworks, parts 1 & 2 (BMJ 2016) — https://doi.org/10.1136/bmj.i2016
- Guyatt et al. — GRADE: going from evidence to recommendations — https://doi.org/10.1136/bmj.39489.470347.AD
- Cochrane Handbook, Chapter 14 (Summary of Findings tables) — https://training.cochrane.org/handbook
- Diamond et al. — Defining consensus: a systematic review of Delphi threshold criteria — https://doi.org/10.1016/j.jclinepi.2013.12.002
- GRADEpro GDT — building SoF tables and EtD frameworks — https://www.gradepro.org/