Original research · Write Plain
The Lay Summary Gap
Federal regulation requires that every clinical trial registered in the United States carry a brief summary written in language intended for the lay public. It names no reading level, and neither does the NIH or the CDC. The guidance written for trial summaries specifically names 6th grade; the most generous bar anyone defends is 8th. We scored 294,566 summaries against the generous one. 99.3% miss.
- Trials scored
- 601,119
- Words analysed
- 182.5M
- Median grade level
- 15.8
- Clearing grade 8
- 0.72%
A field with a stated audience and a stated target
Our first three studies asked whether rules change writing. The Notice/Rule Gap and The Plain English Paradox asked it of regulation; What Simplification Is Made Of asked what simplification consists of when people genuinely do it. This one asks something narrower and more answerable: when writing is aimed at a lay audience by people who were told to aim it there, does it land?
ClinicalTrials.gov is the rare corpus where that question has a defined answer, because both halves of it are written down. The audience is codified — 42 CFR 11.10(b)(3) defines the brief summary as a description of the trial written in language intended for the lay public
. And unlike plain-English mandates, which say be clearer without saying clearer than what, patient-facing health material has published reading-grade targets to measure against — though not, it turns out, the single agreed one everybody cites.
So this is not a before-and-after and there is no control group. It is a census against an external standard: every registered trial is scored, and the question is what share of a legally lay-facing field clears a target that professional guidance has named for decades. Nothing on this page is a causal claim, and the vocabulary throughout is deliberately descriptive.
Nobody actually agrees on the number
"Write health material at 6th to 8th grade" is repeated so widely that we assumed we could cite it and move on. We could not. Checking it against the issuing bodies rather than the secondary literature turned up a genuine disagreement, and two of the largest US agencies declining to name a grade at all. Since the entire headline is measured against one of these numbers, here is exactly who says what.
| Source | Names |
|---|---|
| European Commission, lay summaries of clinical trials (2017) | Grade 6 — the only guidance found that is about trial lay summaries specifically, and it names this exact formula. Its own body text elsewhere says "age 12 upwards", roughly grade 7, so it disagrees with itself by about a grade. |
| AHRQ (HHS), Health Literacy Universal Precautions Toolkit | Grade 4–6 "as recommended", and separately puts the average US adult at grade 8–9 |
| NIH, Clear Communication | No grade — and cautions against the measure |
| CDC, Clear Communication Index | No grade — no reading-grade item among its 20 |
The agency most often cited second-hand as the source of "6th grade" is the NIH, which names no grade anywhere and says this instead:
Readability scores and grade levels are not the be-all and end-all of plain language. The most important consideration is whether your document communicates your message clearly.
AHRQ, in the same document that carries the 4th-to-6th-grade recommendation, is equally direct that the formula is not the thing being measured:
Readability formulas focus on the length of the words and sentences, not on meaning. A high score signals that materials are difficult to read, but a low score does not necessarily mean they are easy to understand.
So this page reports a ladder rather than a verdict against one line, and the headline is read against grade 8 deliberately — the most generous bar any of these bodies supports, two full grades above the only trial-specific recommendation, and the level AHRQ describes as where the average American adult already reads. Choosing the bar most favourable to sponsors is the point. If the summaries fail even that, the finding survives whichever authority a reader prefers.
Almost nothing clears the target
Of 294,566 brief summaries long enough to score, written by 33,493 distinct sponsor organisations, the median lands at grade 15.8 — university reading, for a field whose defined audience is the general public. The share clearing each candidate target:
| Target | Share clearing it | 95% interval |
|---|---|---|
| Grade 6 or below | 0.06% | 0.04% – 0.08% |
| Grade 8 or below | 0.72% | 0.48% – 1.01% |
| Grade 10 or below | 2.93% | 2.42% – 3.52% |
| Grade 12 or below | 10.20% | 9.33% – 11.16% |
At the strict end of the guidance, grade 6, the answer is 0.06% — roughly one trial in 1,667. At the lenient end, grade 8, it is 0.72%. Even moving the goalposts to grade 12, which no health-literacy guidance proposes and which is ordinary secondary-school reading, still only 10.2% clear it.
Every interval on this page is a cluster bootstrap over lead sponsor organisations, not trials. One pharmaceutical company or academic medical centre registers hundreds of trials written by the same communications team, and treating those as independent observations is the single easiest way to make a confident-looking number out of nothing.
| Grade band | Trials | Share |
|---|---|---|
| < 6 | 164 | 0.1% |
| 6 to < 8 | 1,963 | 0.7% |
| 8 to < 10 | 6,516 | 2.2% |
| 10 to < 12 | 21,400 | 7.3% |
| 12 to < 16 | 123,197 | 41.8% |
| 16 and over | 141,326 | 48.0% |
It is not an artefact of short summaries
The obvious objection is length. Readability formulas are unstable on very short text, and a 69-word summary is short. So the census was re-run at every word floor from none to 200, discarding more of the corpus each time:
| Word floor | Trials | Median grade | Clearing grade 8 |
|---|---|---|---|
| None | 294,566 | 15.84 | 0.72% |
| 30 words | 294,566 | 15.84 | 0.72% |
| 60 words | 294,566 | 15.84 | 0.72% |
| 100 words | 208,267 | 16.01 | 0.81% |
| 150 words | 124,875 | 15.92 | 1.11% |
| 200 words | 77,403 | 15.85 | 1.54% |
The pass rate moves from 0.72% to 1.54% across that whole range, while the corpus shrinks by 73.7%. The finding survives; if anything the longer summaries are marginally worse.
One sponsor class is doing something different
Sorted by median grade level, the sponsor classes are mostly indistinguishable from one another — and then there is the NIH.
| Lead sponsor class | Trials | Median grade | Clearing grade 8 |
|---|---|---|---|
| NIH | 8,144 | 12.55 | 12.89% [11.33%, 16.10%] |
| Industry | 32,606 | 14.94 | 1.61% [1.10%, 2.27%] |
| Research network | 2,682 | 14.94 | 0.22% [0.03%, 0.48%] |
| Other government | 9,153 | 16.00 | 0.15% [0.08%, 0.23%] |
| Academic and other | 238,889 | 16.03 | 0.22% [0.18%, 0.28%] |
| Other US federal | 2,834 | 16.18 | 0.11% [0.00%, 0.26%] |
NIH-sponsored trials clear grade 8 12.89% of the time against 0.11% for other us federal — about 117 times the rate, and their median summary sits 3.6 grade levels lower. Their intervals do not overlap.
This page will not tell you why, because the design cannot support an answer: sponsor class is not randomly assigned, and NIH trials differ from industry trials in funding, subject matter, and who physically writes the registration. What the number does establish is that the target is reachable at scale. Across 8,144 trials, one class of sponsor hits it far more often than the rest, which makes the 99.3% miss rate a fact about practice rather than about the limits of the genre.
The lay field is harder than the technical one
Two thirds of trials also carry a detailed description, which the registry defines as the technical account of the protocol, with no lay-audience requirement attached. That gives a within-record comparison: the same sponsor, the same trial, the same week, writing once for the public and once for professionals. Across 121,182 trials with both fields long enough to score:
| Measure | Brief − detailed | 95% interval | Which field is plainer |
|---|---|---|---|
| Flesch-Kincaid grade | +0.57 | [0.41, 0.70] | the technical one |
| Words per sentence | +1.26 | [1.05, 1.46] | the technical one |
| Passive sentences | -9.1% | [-9.5%, -8.6%] | the lay one |
| Nominalizations per 1,000 words | -1.95 | [-2.68, -1.34] | the lay one |
The lay-facing field scores 0.57 grade levels harder than the explicitly technical one [0.41, 0.70], and the interval excludes zero. Only 43.6% of trials have the brief summary come out easier — worse than a coin flip. Median brief grade 16.05 against 15.56 for the detailed description.
But the last two rows point the other way, and that contradiction is the interesting part.
The gap is sentence length, not vocabulary
Flesch-Kincaid is not a black box. It is exactly 0.39 × (words per sentence) + 11.8 × (syllables per word) − 15.59, and it is linear in both components, so the grade gap above splits into two terms that add to the whole with nothing left over:
| Term | Grade levels | Share of gap |
|---|---|---|
| Sentence length (1.26 more words per sentence) | 0.492 | 86.8% |
| Word length (0.0063 more syllables per word) | 0.075 | 13.2% |
| Total | 0.567 | 100% |
86.8% of the gap is sentence length. Brief summaries are not written in fancier words than protocol descriptions — the vocabulary difference is 0.0063 syllables per word, which is close to nothing. They are written in longer sentences.
That squares the contradiction in the previous table. On the measures that describe construction — passive voice, nominalizations — brief summaries are genuinely plainer, and significantly so. Somebody is clearly trying. But the summarising itself compresses: a protocol section can afford a short declarative sentence per step, while a summary packs purpose, population, intervention, and endpoint into single long sentences joined by commas and relative clauses. The result is prose with plain construction and unreadable rhythm.
The syllable term is derived as the residual: the scored data publishes sentence length but not syllables per word, so it is the one number here not read directly from the analysis. The sentence-length term and the total are both published columns, so the subtraction is checkable from the CSVs below.
And it has been getting harder
Median grade level of the brief summary, by year of first posting. This is a description of the corpus and nothing more — there is no treatment here, no control, and no design that would let anyone read a cause into the slope. It is drawn because it is a real feature of the data, and labelled plainly because our first study is the reason we no longer build the machinery that would tempt us to over-read it.
The early years are thin — 2000 carries 357 trials against 26,748 in 2026 — so the left of that line is noisier than it looks. What is not noise is that no year in 27 comes close to the target.
Every therapeutic area misses
| Condition area | Trials | Median grade | Clearing grade 8 |
|---|---|---|---|
| Healthy volunteers | 1,742 | 14.67 | 5.45% |
| Dermatology | 2,793 | 14.84 | 1.00% |
| Ophthalmology | 2,964 | 15.40 | 1.05% |
| Haematology | 1,923 | 15.59 | 1.87% |
| Oncology | 51,352 | 15.61 | 0.98% |
| Immunology & allergy | 1,576 | 15.61 | 1.65% |
| Musculoskeletal & rheumatology | 11,541 | 15.70 | 0.38% |
| Reproductive & maternal | 6,312 | 15.79 | 0.40% |
| Other | 87,982 | 15.79 | 0.65% |
| Metabolic & endocrine | 16,521 | 15.80 | 0.75% |
| Renal & urology | 6,157 | 15.87 | 0.29% |
| Respiratory | 9,901 | 15.91 | 0.61% |
| Gastroenterology & hepatology | 7,248 | 15.91 | 0.57% |
| Neurology | 12,066 | 15.94 | 0.69% |
| Infectious disease | 20,465 | 15.96 | 1.02% |
| Anaesthesia & pain | 9,511 | 16.05 | 0.32% |
| Psychiatry & behavioural | 16,808 | 16.20 | 0.45% |
| Cardiovascular | 27,699 | 16.25 | 0.45% |
The spread across 18 therapeutic areas is 1.6 grade levels — from healthy volunteers at 14.67 to cardiovascular at 16.25. Not one of them has a median within four grade levels of the target. 1 area with fewer than 500 trials is held out of this table; at that size the interval is wider than the effect and a sorted table would rank noise.
How this was measured
The corpus is a full census of the ClinicalTrials.gov registry taken through its v2 JSON API on September 9, 2026 — 602,104 records, of which 601,119 carry a brief summary that survives normalisation. There is no sampling anywhere in this study; the only exclusion is structural, and it removes 985 records that have no brief summary at all.
Scoring uses the same analyzers that run on every tool page of this site, unmodified. No measure was added for this study. Registry prose arrives line-broken rather than paragraph-wrapped, and how a line break becomes a sentence boundary materially changes the grade level, so that one transformation lives in a unit-tested module rather than in the pipeline.
Two inclusion floors are pre-committed rather than tuned. The census requires 60 words and 3 sentences, which is what makes 294,566 of 601,119 trials eligible; the paired comparison requires 100 words in both fields, which is why it runs on 20.2% of the corpus. The SEC study's 200-word floor was deliberately not copied here: on this corpus it would have retained about 6% of trials, and inheriting a threshold because it worked elsewhere is how a study quietly becomes about its own filter.
Intervals are cluster bootstraps over lead sponsor organisation — 1,000 replicates for the headline estimates, 300 for the breakdown tables, a split recorded here so nobody has to guess which number got which treatment.
What this does not establish
- That these summaries are unreadable. Flesch-Kincaid counts syllables and sentence length. It does not know that randomised is a word patients learn quickly, and it penalises unavoidable terminology — a trial of pembrolizumab cannot spend fewer syllables on its own drug. The decomposition above is the honest response to that objection: 86.8% of the measured gap is sentence length, which no terminology argument explains away.
- That grade 8 is the right line. It comes from health-literacy guidance, not from regulation — 42 CFR 11.10(b)(3) names the audience and no number — and as the table above shows, the issuing bodies disagree with each other and two of them decline to name a grade at all. Grade 8 is used because it is the most generous defensible bar, not because it is authoritative. A reader who rejects it can read the distribution and pick their own line; the pass rates at grades 6, 10 and 12 are all published.
- That anything caused anything. No treatment, no control, no counterfactual. The by-year series and the sponsor table are descriptions of a corpus.
Check our work
Everything above is the output of the analysis rather than a summary of it. The registry API is public and needs no key, the census does not sample, and the pipeline is deterministic, so re-running it against the same corpus date reproduces these numbers exactly.
Per-trial scores are published in full — one row per trial, both fields, every measure — split by year of first posting because the whole file is 151.4 MB and no single asset that size belongs on a web page. Concatenating them reproduces the complete 601,119-row dataset. Those 28 files are served from object storage rather than from this site, because 151.4 MB of CSV does not belong in a git repository. Every aggregate the page actually renders stays here, so nothing above needs a second host to verify.
- Shard index — every per-trial file with its year, row count, and size (28 shards, 151.4 MB total)
- Pass rates — the headline table, with intervals
- Grade distribution — the bands, in full
- Threshold sensitivity — the word-floor robustness check
- By sponsor class, by condition area, by phase, by year
- Paired selection — who qualifies for the within-record comparison, and who does not
Per-trial scores, 28 yearly files
| Year | Trials | Size |
|---|---|---|
| 1999 | 1,056 | 0.3 MB |
| 2000 | 1,063 | 0.3 MB |
| 2001 | 1,773 | 0.5 MB |
| 2002 | 1,378 | 0.4 MB |
| 2003 | 3,588 | 1.0 MB |
| 2004 | 3,166 | 0.8 MB |
| 2005 | 12,796 | 3.2 MB |
| 2006 | 10,916 | 2.7 MB |
| 2007 | 12,541 | 3.1 MB |
| 2008 | 17,501 | 4.2 MB |
| 2009 | 16,958 | 4.0 MB |
| 2010 | 17,309 | 4.1 MB |
| 2011 | 17,786 | 4.3 MB |
| 2012 | 19,421 | 4.7 MB |
| 2013 | 20,389 | 4.9 MB |
| 2014 | 23,273 | 5.7 MB |
| 2015 | 24,068 | 6.1 MB |
| 2016 | 27,741 | 7.0 MB |
| 2017 | 29,117 | 7.4 MB |
| 2018 | 30,911 | 7.9 MB |
| 2019 | 32,463 | 8.3 MB |
| 2020 | 36,678 | 9.4 MB |
| 2021 | 36,967 | 9.4 MB |
| 2022 | 37,959 | 9.7 MB |
| 2023 | 39,638 | 10.1 MB |
| 2024 | 43,596 | 11.2 MB |
| 2025 | 42,861 | 10.9 MB |
| 2026 | 38,206 | 9.8 MB |
Published September 9, 2026. Corpus dated September 9, 2026. Derived from ClinicalTrials.gov, a service of the U.S. National Library of Medicine; registry records are US government works in the public domain. No summary text is republished here or in the data files — identifiers and numbers only, and every row carries the NCT ID needed to fetch the original. Corrections are welcome — get in touch.