About
What is a discourse graph, and why is this synthesis rendered as one?
A discourse graph is an alternative form of scientific communication. Instead of a single linear document, the argument is composed of typed nodes — Questions, Claims, Evidence, Sources — connected by typed edges: addresses, supports, opposes, derived from, qualifies. Every node is self-contained, addressable, and individually contributable.
The form was developed by Joel Chan, Matthew Akamatsu, and collaborators, and refined inside Roam Research, Protocol Labs, and adjacent research communities. The Q/C/E/S core schema is small enough to remember; this synthesis extends it with a few types specific to evidence work: Caveat (a limitation that qualifies a piece of Evidence) and Artifact (a concrete intervention or system — a tablet-on-wheels interpreting cart, a bilingual-provider program — that the evidence is about).
- Question
- Question. An unknown we want to make known — addressable by studies, experiments, or prototypes.
- Claim
- Claim. An atomic, generalized assertion about the world that (proposes to) answer a research question.
- Evidence
- Evidence. A specific empirical observation from a particular study.
- Source
- Source. A published research source — a journal article, conference paper, or book.
- Artifact
- Artifact. A concrete system (prototype, standard, intervention) that instantiates a pattern or method.
How this synthesis is built
This graph is extracted from the published literature on language access in healthcare, working from a curated corpus of research papers. An AI-assisted pipeline reads each paper and drafts the nodes — the Questions it asks, the Claims it makes, the Evidence behind them, and the Caveats that bound them — every quote grounded verbatim against the source. Domain experts then review and commit: the AI proposes, the human verifies. Every node therefore carries a curation status — Initial AI draft, In expert review, or Expert-verified — and you can filter the topology by it to separate what an expert has checked from what is still a first draft. Nothing here is a finished review; it is a living evidence base that gets stronger as claims accumulate supporting and opposing evidence over time.
Contributions become atomic. A paper bundles a question, methods, claims, and evidence together; none of it gets published until all of it does. To share one new observation, you write the surrounding apparatus — introduction, methods, related work, discussion — even when none of that is new. A discourse graph removes the bundle. One new observation is one Evidence node, with edges to the Claims it supports or opposes. One new assertion is one Claim node, addressing a Question and supporting or opposing other Claims. One new line of inquiry is one Question node. Each attaches to what it bears on, and that's the contribution.
Specialists become authors. A paper demands generalist scaffolding — introduction, methods, related work, framing, discussion — so the people who hold one sharp contribution often can't be authors on their own terms. The data curator who tracked down a hard-to-find Source, the clinician who can name the caveat that bounds a finding, the practitioner with one decisive field observation: each typically has to partner with a generalist who will wrap the piece in apparatus, or watch the contribution go uncredited. The graph removes the apparatus requirement. A Caveat, an Evidence item, a Source, a single Claim is itself a complete, citable, credited contribution. Authorship stops being gated on the ability to produce a whole paper, and the population of people who can author scientific work expands to anyone with one good node.
Credit becomes granular. Each node has its own ID and its own PID — citable independently. A Caveat, a Source, an Evidence item, a Claim can be cited (and tracked) on its own merit. The contributor who proposed C-0017 gets credit when C-0017 is invoked, even when the paper that introduced it isn't. Funders, hiring committees, and citation indexes can resolve attribution to the unit of contribution rather than rolling it up into “lead author of paper X.”
Review becomes a linter; validation becomes topological. Peer review of a paper bundles many things at once — gatekeeping, wording, framing, validating the work, signaling trust to the reader. The bundle dissolves at the node level. Reviewing a node is mostly form-checking: does this Evidence cite the Source it claims, is the Claim it points at really a Claim, is the prose self-contained. Most of that is lintable. The substantive work — what's true, what holds up, what matters — doesn't happen in a review pass; it happens in the graph itself, over time. A weak Claim accumulates opposing Evidence. A strong one accumulates supporting Evidence and Claims that build on it. The trust signal is the topology, not a stamp.
Publishing becomes continuous. A paper waits — for a journal slot, a conference deadline, a grant cycle, an annual report. By the time the work appears it is often eighteen months old, and a counter-finding discovered next week has nowhere to land until the next cycle opens. The graph has no cycle. A new Evidence node ships the day it is found; a counter-Claim ships the day it is formulated; a Question that opens up at midnight is addressable by morning. Publishing tracks the rhythm of inquiry instead of the rhythm of institutions.
Narratives become snapshots. A paper captures the state of the argument at the moment it was written, and that is the state it continues to assert long after the evidence has moved. A narrative composed from the graph is dated by construction. Today's telling reflects today's evidence; next year's telling, regenerated against a graph that has accumulated supporting and opposing evidence in the meantime, is a different telling. Nothing is rewritten — the underlying nodes have moved, and the rendering follows. The narrative is a view of the graph at a moment in time, and another view can be composed whenever it is useful.
This site and each composed narrative all derive from the same node files in graph/.
Why this form — a revisable intermediate representation
Ways of organizing a literature sit on a spectrum. At one end, literature graphs (citation networks, topic maps) cover almost everything but say little about what any of it means. At the other, knowledge graphs and meta-analyses are richly expressive — typed entities and relations, pooled effect sizes — but only over the narrow slice of a literature that has been forced into a fixed schema or a single shared construct. A discourse graph sits in the middle: expressive enough to reason over, broad enough to cover a messy literature, and — crucially — carrying granular provenance and uncertainty on every node.
That middle position is the point, not a compromise. The graph is an intermediate representation— like a compiler's IR between source code and machine code. Source papers compile into the graph once; the graph then compiles out to whatever you need — a narrative, a knowledge graph, a meta-analysis for the sub-question where the evidence is commensurable. The expensive, lossy step — reading the papers — happens once. When the model has to change (a construct splits in two, a schema is revised, a moderator turns out to matter) you re-wire the graph; you do not re-read the corpus. Revising a model built on the graph is cheap; going destructively back to the source texts to start over is not.
That makes the graph a resource people build on directly, not just a finished output. A natural extension is letting a reader pick a claim's body of evidence and run a living meta-analysis of their own choosing over it — assembling the commensurable Evidence, pooling it under assumptions they can see and contest, and having it re-run as new evidence lands — rather than inheriting one pooled estimate, frozen at publication, that someone else chose for them.
This is why Evidence and Claims are distinct node types. A Claim is a compressed, generalized assertion — modular and quotable, but lossy: it has abstracted away the particulars. An Evidence node is the balance point between compression and context — modular enough to reuse, yet grounded in a verbatim quote and linked to the specific methods that produced it (its What, How, and Who). That retained context is what lets the graph be synthesized responsibly: deciding whether two findings measure the same construct, reasoning about whether they are commensurable enough to pool, noticing a hidden moderator that explains why they disagree. Compile straight from text to one pooled number and those judgments are made silently and irreversibly; hold them in the graph and they stay explicit and contestable.
The longer-term aim is to make even this cheaper: if research were modular by construction — a finding published as a grounded, addressable unit in the first place — the extraction step that builds this graph would be less necessary, or unnecessary. This synthesis extracts the graph from conventional papers because that is the literature we have; the form points at a world where the graph is the literature.
How to read it
- By topology: /graph shows the whole argument at a glance. Nodes are colored by type; edges are colored by relation. Click any node to inspect its bundle — everything one hop away.
- As narratives: /narratives renders linear readings composed directly from the graph, with each citation linked to its Source node. A toggle at the top swaps between narratives written for different audiences and framings.
- By node: every node sits at
/node/<ID>. Each page shows the prose body, outbound edges, inbound backlinks, and a deep link to open a GitHub issue about that one node.
How to contribute
Discussion happens at node granularity. Open an issue with the node:<ID> label, or open a pull request that adds a counterclaim, counter-evidence, or a new question. The full contribution model lives in CONTRIBUTING.md.
Further reading
- discoursegraphs.com — the canonical Q/C/E framework and its provenance.
- Discourse graphs and the future of science — Protocol Labs' framing of the form's research-infrastructure implications.
- DiscourseGraphs/schemas — the underlying schema repository this work extends.
Acknowledgments
- Renderer & site code — built on Resilient Data Futures — Discourse Graph by the SciOS Resilient Data Futures Working Group (rdf.scios.tech, jring-o/rdf), used under CC BY 4.0 (content) / MIT (code) and adapted for this synthesis.
- Extraction methodology — the initial discourse-graph extraction skill was developed by Jay Patel.
- Project lead & contact — Joel Chan (joelchan@umd.edu).