Original research · Write Plain

What Simplification Is Made Of

A quarter of a million subjects have been written up twice: once on English Wikipedia, and once on Simple English Wikipedia by people whose stated and only distinguishing goal was to make it easier to read. We scored 1,398 matched pairs with this site's own analyzers — and found that one of our own tools measures something that moves on its own.

Article pairs
1,398
Words analysed
3.1M
Readability gap
11.0 Flesch points
Diagnostics that fall
7 of 12
The corpus

The same subject, written twice, on purpose

Our first study asked whether a law changed how people write. Our second asked whether a rule with teeth did. Both were questions about regulation. This one is a question about writing: when text genuinely gets easier, what changed?

Simple English Wikipedia is an unusual thing to have lying around. It covers the same subjects as English Wikipedia, under the same licence, edited by overlapping communities, with one difference held constant across every article: it is supposed to be easier to read. Each pair is therefore a matched observation of human simplification — the same facts, the same subject, one of them rewritten by somebody trying to do exactly what every writing tool on this site tells you to do.

That makes it the one corpus that can put our own product under test. We wrote down what we expected to find before running it, including four diagnostics we predicted would not move at all. Three of those predictions turned out to be wrong, and the section below says which.

The result

The gap is mostly word length

Flesch Reading Ease is not a black box. It is exactly 206.835 − 1.015 × (words per sentence) − 84.6 × (syllables per word), so the readability gap between two paired articles splits arithmetically into a sentence-length term and a word-length term, with nothing left over. This is a decomposition, not an attribution: the two numbers add up to the whole gap by construction, and you can recompute them from two columns of the CSV at the bottom of this page.

Across 1,398 pairs, the Simple English article scores 11.0 Flesch points easier [9.8, 12.3]. 59% of that is word length; the rest is sentence length.

Sentence length · 4.5 points · 41%Word length · 6.5 points · 59%English: 21.0 words per sentence, 1.59 syllables per wordSimple: 16.6 words per sentence, 1.51 syllables per word
The 11.0-point Flesch gap between paired articles, split into the only two things Flesch Reading Ease is made of. The split is exact — the residual across all 1,398 pairs is zero to fifteen significant figures.

Sentence length gets almost all the attention in writing advice, including ours. It is the easier thing to notice, the easier thing to count, and the easier thing to build a tool around. But when human beings actually simplify something, the larger share of the measurable gain comes from choosing shorter words — 1.59 syllables per word down to 1.51, which sounds like nothing and is worth 6.5 Flesch points.

The audit

Which of our own diagnostics real simplifiers actually act on

Every tool on this site asserts a theory. The passive voice finder says passive constructions make text harder. The plain language checker says nominalizations, hedges, expletive constructions and wordy phrases do. The Throat-Clearing Score says stall phrases at the start of a sentence do. Those theories come from style guides, and style guides do not come with measurements.

Here is what happens to each of them when a human being sets out to make a text easier. Every row is a claim we make to our readers. We registered the predicted order, and the four we expected not to move, before this analysis was run against real data — both predictions ship as columns in the audit file, so the scoring can be checked rather than taken on trust.

English minus Simple · 1,398 pairs · topic-clustered 95% intervals
DiagnosticEnglishSimpleCut by95% interval of the changeVerdict
Wordy phrasesper 1,000 words0.630.3446%[0.24, 0.35]falls
Sentences over the long-sentence thresholdshare of sentences0.4480.26042%[0.178, 0.200]falls
Adverbsper 1,000 words8.546.0629%[2.09, 2.95]falls
Nominalizationsper 1,000 words16.7611.9129%[4.08, 5.75]falls
Average sentence lengthwords21.0016.5821%[4.15, 4.74]falls
Subject-verb distanceper 1,000 words2.562.0321%[0.43, 0.67]falls
Sentences containing a passiveshare of sentences0.3040.24320%[0.050, 0.071]falls
Passive constructionsper 1,000 words16.6016.391%[-0.60, 0.86]does not move
Hedgesper 1,000 words1.872.08-11%[-0.34, -0.08]rises
Stacked negativesper 1,000 words0.220.26-18%[-0.09, 0.00]does not move
Expletive constructionsper 1,000 words1.031.58-53%[-0.78, -0.33]rises
Stall phrases at sentence openingsper 1,000 words0.000.01—[-0.01, -0.00]too rare to measure

7 of the 12 diagnostics fall by more than a tenth of their English level with an interval that excludes zero — a threshold fixed in advance, because with 1,398 pairs an interval can exclude zero for a change no writer would ever notice. 2 do not move: passive constructions, stacked negatives. And 2 go the wrong way — the Simple English article has more of them: hedges, expletive constructions. Stall phrases at sentence openings occurs too rarely in encyclopedia prose to measure either way — about one instance per 610,740 words — so this corpus has nothing to say about it.

The catch

Simple English contains just as much passive voice

Two rows in that table come from the same tool and point in opposite directions, and the gap between them is the most useful thing this study found.

The share of sentences containing a passive falls from 30.4% to 24.3% — a 20% cut, exactly the kind of number a writing tool likes to show you. But the number of passive constructions per thousand words barely changes at all: 16.60 against 16.39, an interval of [-0.60, 0.86] that comfortably contains zero.

Both are true, and the arithmetic connecting them is not subtle. An English sentence averages 21.0 words and carries 0.35 passives. A Simple English sentence averages 16.6 words and carries 0.27. Shorter sentences hold fewer passives because they hold fewer of everything. The passive rate went down without anyone removing a passive.

This matters beyond Wikipedia, because "percentage of sentences containing a passive" is the number almost every writing tool reports, ours included. It is not a measure of how passive your writing is. It is a measure of how passive your writing is and how long your sentences are, and it will improve on its own every time you split a sentence in two.

We are not claiming passive voice does not matter. We are claiming the metric moves for reasons that have nothing to do with passive voice, and that this site was reporting it without saying so. The passive voice finder now shows both figures and explains the difference — that change ships with this study, not after it.

Corrections

What we got wrong

3 of 12 predictions failed:

  • Passive constructions — we predicted it would fall; it does not move.
  • Hedges — we predicted it would fall; it goes up instead.
  • Adverbs — we predicted it would not move; it falls.

We committed in advance to publishing this whichever way it came out, and to correcting the pages it implicates rather than shipping a study that quietly disagrees with the site it appears on. The pages affected: /passive-voice-finder/, /adverb-checker/, all edited in the same change that published this study.

The rises deserve a sentence of their own, because they are the least convenient thing here. Simple English uses more expletive constructions — "there is", "it is" — than English Wikipedia does, by 53%. Ourfiller words tool flags those, and an article here explains when to leave them alone. Read charitably, "There are three kinds of engine" is a genuinely easier opening than a participial phrase doing the same job, and people writing for a struggling reader reach for it on purpose — which is what that article already says, and now has evidence for. Read uncharitably, a rule we inherited from style guides points the opposite way from the people who do this work.

A diagnostic that does not move here is not necessarily wrong. The feature may genuinely not impede readers; simplifiers may not think about it; it may already be rare enough in encyclopedia prose that there was nothing to remove. This corpus separates none of those. What it does establish is that the feature is not part of what human simplification consists of — a narrower claim than the tools make, and the reason we are making it about our own tools first.

Caution

Simple English articles are shorter, not only simpler

Length correlates with almost everything, so a study comparing a 4,000-word English article with a 300-word Simple stub would be measuring truncation and calling it simplification. Two defences, and the second is the one that matters.

First, every rate above is per 1,000 words — which is precisely why the passive result above could be seen at all. Second, a pair joins the study only if the English article is at most 4× its Simple counterpart by word count. That ceiling is the threshold doing the most work anywhere here, so this is what it does:

The decomposition at different length-ratio ceilings
CeilingPairsFlesch gapSentence-length shareWords per sentence removed
1.25×3629.532%3.0
1.5×5079.436%3.3
2×75610.137%3.7
3×1,11010.441%4.2
4×1,39811.041%4.4

The strictest arm is the honest one. On the 507 pairs whose two articles are within 1.5× of each other — as close to the same article written twice as this corpus gets — the gap is 9.4 Flesch points, of which 36% is sentence length. The headline direction survives and the word-length share grows: the more you control for length, the more of simplification turns out to be word choice.

Confidence intervals throughout come from a bootstrap that resamples whole topic areas, not articles. Wikipedia's simplification effort is uneven across subjects, and articles about the same subject share editors, WikiProjects and house habits. Treating them as independent would have understated the standard error on the headline by a factor of 1.15 — which is how our first study talked itself into two null results that were not significant at all.

Subjects

Which subjects get simplified hardest

Separately from everything above, and making no causal claim: here is how far apart the two encyclopedias sit, subject by subject. This is a description of the corpus, not an estimate of anything — but it is a striking one. The gap is widest exactly where the English article is hardest to begin with.

Topic areas with 25+ pairs · sorted by readability gap
Topic areaPairsFlesch gapEN words/sentenceSimple words/sentence
Biology and medicine11716.120.015.8
Physical sciences3914.520.916.8
History5614.323.216.7
Technology and engineering2613.720.917.1
unclassified19012.821.216.7
Mathematics and computing2712.721.417.0
Military5912.421.415.8
Arts and literature11311.120.515.5
Religion and philosophy5911.021.817.0
Transport2510.821.316.3
Geography and places1199.620.316.1
Sports and games1499.421.116.7
Society and culture809.320.717.0
Music898.620.416.8
Politics and law638.222.417.5
Film and television1237.921.016.6
Economics and business357.721.519.1

Note how little the sentence-length columns move across the table compared with the Flesch gap. Simple English writes at roughly the same sentence length whatever the subject; what changes between subjects is how much vocabulary there was to simplify. That is the same finding as the headline, arriving from a different direction.

Method

How this was measured

The population comes from two MediaWiki database dumps of Simple English Wikipedia (20260901): the page table, for titles and redirect flags, and the langlinks table, which records the English title each Simple article claims as its counterpart. That pairing is editorial — stated by Wikipedia's own editors — rather than a title-matching heuristic we invented, and every mistake such a heuristic made would be two articles about different subjects contributing a difference to the headline. It yields 269,003 pairs, of which 3,254 were fetched and 1,398 qualified.

Both sides' prose comes from the same action=query&prop=extracts&explaintext=1 endpoint and is cleaned by the same function. This is the single most important design decision in the study. The whole measurement is a difference between two cleaned texts, so anything the cleaner removes from one side more than the other is measured as simplification. Taking both sides from one endpoint makes symmetric cleaning true by construction rather than by care.

Reference lists, "See also", "Related pages", "External links", "Other websites" and their subsections are removed from both sides using both wikis' vocabularies. Heading lines are removed, because Simple English uses more and shorter sections and a heading left in the prose counts as a very short sentence. Pronunciation glosses and foreign-script renderings are removed, because English Wikipedia opens articles with them and Simple English does not — left in, an IPA transcription is counted as a word with a wild syllable count on one side of the pair only. Each of those, unfixed, would have pushed the result in the direction we were hoping for. On the pairs that qualified, cleaning removed an average of 2.8 apparatus sections per English article against 1.7 per Simple one, and 1.8 glosses against 1.7.

A pair joins the study only if both sides clear 250 words and 8 sentences after cleaning, and the length ratio clears the ceiling above. Redirects, disambiguation pages, and list, index and timeline pages are excluded — including when an English langlink now redirects to one, which is a real case we found by reading pairs by hand rather than by testing. Nothing is excluded for being badly written: Simple English's unevenness is the study's subject, and a filter keyed to the thing being measured would delete the signal.

Every measure comes from an analyzer that already runs on this site — the same code behind the passive voice finder, the sentence length visualizer, the plain language checker and the Throat-Clearing Score. Nothing was adjusted for this study, because the analyzers are what is under test and tuning them for it would have made the test circular.

What this cannot tell you

  • It is not causal, and it is not about readers. Nobody was tested on comprehension. What is measured is what simplifiers did, not what worked.
  • The surviving corpus is not Wikipedia. Most of Simple English Wikipedia is stubs and partial translations, and 53% of the fetched pairs failed the length-ratio ceiling alone. What passes is the subset where somebody did the work on both sides — the right corpus for the question, and a narrower one than the encyclopedia.
  • The sampling frame skips very short Simple articles. To avoid spending five requests in six on articles that could not qualify, the sample was drawn from pairs whose Simple article carries at least 2,000 bytes of wikitext. We measured what that discards rather than asserting it: of 400 pairs fetched from below the threshold, 2 would have qualified — a false-exclusion rate of 0.5%.
  • Readability formulas are crude. Flesch Reading Ease is used here because it decomposes exactly, not because it is a good measure of anything — and this study is partly a demonstration of why a composite score hides more than it shows. See how readability formulas fail.
  • Simple English is written for a specific audience — largely people reading English as an additional language, and children. What simplifiers do for that audience is not necessarily what a general reader needs. See writing for non-native English readers.
  • 14% of pairs are unclassified by the topic classifier and form their own cluster rather than being forced into a subject bucket.
Data

Check our work

Everything above is the actual output of the analysis, not a summary of it. This is the most reproducible of our three studies: the dumps are public, the API needs no key, and the sample is selected by a deterministic hash rather than a random draw, so anyone re-running the pipeline gets the same 1,398 pairs we did.

  • Per-pair scores — 1,398 rows, both wikis, every measure, plus both titles so you can fetch the originals
  • The tool audit — the table above, with the predictions we registered in advance
  • The decomposition — the headline split and its intervals
  • Topic areas — the subject table, in full

Published September 8, 2026. Derived from English Wikipedia and Simple English Wikipedia, whose text is CC BY-SA 4.0 © Wikipedia contributors. No article text is republished here or in the data files — titles and numbers only. Corrections are welcome — get in touch.