Language Access in Healthcare

How to review an extraction

We're turning papers into a discourse graph: a network of small, linked, quote-grounded nodes. Your job is to check that the AI's nodes are faithful to the source — not to write them from scratch. This guide says what each node type is and what a good one looks like.

The golden rule everything else follows from: every substantive statement is backed by a verbatim quote with a page number, and each node says exactly one thing (one finding, one claim, one limitation). If a node mixes two findings, or a quote doesn't actually say what the node claims, it needs an edit.

The node types

QUE

Question. An unknown we want to make known — e.g. “How does language concordance affect healthcare outcomes?” Claims answer questions.

CLM

Claim. An atomic, generalized assertion about the world that (proposes to) answer a question — e.g. “Professional interpreters reduce readmissions for LEP patients.” A claim transcends any single paper: many studies can support or oppose it. Claims are deliberately more lossy than evidence.

EVD

Evidence. A specific empirical observation from one study — a number, comparison, or qualitative finding, grounded in a verbatim quote (and a figure/table where relevant). Evidence supports or opposes a claim. This is what you review here.

SRC

Source. The paper itself (authors, year, journal, DOI). Evidence is derived from a source.

CVT

Caveat. A limitation that qualifies a piece of evidence (not a claim) — e.g. a single-site sample, a retrospective design. Shown under the evidence it qualifies.

What a correct & complete EVD looks like

  • Atomic. One finding. “LOS did not differ (IRR 0.94)” is one EVD; readmission is a separate EVD even from the same table.
  • Past tense. An EVD reports what a specific study observed (“LOS was1.5 days longer”) — signalling its situated, contextual nature. The timeless, present-tense version belongs in a claim. Tense mismatch usually means it's on the wrong side of the evidence/claim line.
  • Verbatim-grounded. The quote is copied exactly from the paper and actually states the finding— it's the right sentence, not a coincidental keyword match. Page number present.
  • Substantively faithful. The finding is faithful to the source — direction, magnitude, significance, and CIs for quantitative results; an accurate characterization for qualitative ones. A null result is reported as null, not spun (or vice-versa).
  • Grounded in the right object.If the finding lives in a table/figure, that object's crop is embedded — or it's correctly text-only.
  • Methods context (What / How / Who). What = the observable measured (the outcome itself, not the design). How = the design and procedure. Who = the sample / setting it generalizes to. Each backed by its own quote.
  • Linked to a claim with the right polarity. A null or contrary finding should oppose the claim it bears on, not support it.

What a correct & complete CLM looks like

  • One generalization. A claim that combines two distinct assertions should be split.
  • Stated as a generalization, in present tense. “LEP is associated with longer stays” is a claim (timeless); “LOS was 2.26 vs 2.12 days” is evidence (past, situated).
  • Backed by evidence on both sides where it exists. Supporting and opposing EVDs are wired in; a contested claim should show both.
  • Body-of-evidence appraisal is a human/clinician task. The certainty / GRADE judgment is authored by an expert, not the AI.

The first-pass checklist (3 per evidence node)

This pass checks only that each evidence node is faithful to its source. Claim polarity and methods context are deferred to a later pass.

Verbatim
Is the quote the right sentence, and does it match the PDF? (An audit already checked the string; you confirm the meaning.)
Substantive fidelity
Direction, magnitude, significance, and CI faithful to the source for quantitative results; an accurate characterization for qualitative ones.
Grounding
Correct figure/table embedded — or correctly none.

Deferred to a later pass: whether the evidence really supports / opposesthe claim it's wired to (claim polarity), and whether its What / How / Who methods context is accurate.

The five verdicts

  • CorrectFaithful as-is. The common case — one click.
  • Needs an editClose but off — type the corrected value. This becomes a proposed fix.
  • WrongNot salvageable as written — say what's wrong.
  • MissingThe element should exist but isn't there — e.g. a finding with no grounding quote, or a figure/table that should be embedded but isn't. Flags it for another extraction pass (say what's missing). Distinct from “wrong”: nothing was captured to be wrong.
  • N/AThe dimension genuinely doesn't apply (e.g. a text-only finding with no figure to ground).

Effort scales with disagreement: an all-correct node is a few clicks; only edits need typing. Your name is attached to every judgment — you're proposing corrections, the maintainer commits them.