Original research · Write Plain
What a Regulator Calls Unclear
When SEC staff review a filing, they quote the phrase they could not follow and tell the company to rewrite it in plain English. That is a regulator pointing at particular words, in public, with the words attached. We collected 1,546 of those phrases from 585 letters, found each one in the filing it came from, and asked whether this site's own diagnostics flag what the regulator flagged. Mostly, they do not — because the regulator is flagging jargon, and the diagnostics are looking for sentence structure.
- Flagged phrases
- 1,546
- Letters
- 585
- Period
- 2004–2025
- Beat chance
- 2 of 10 diagnostics
The third rung
Our first study found that Congress asking federal agencies to write plainly produced no measurable change. Our second found that the SEC writing a rule left the regulated sections of a prospectus the hardest part of the document. This is the next step on that ladder: not a law, not a rule, but a named reviewer telling a named company that a named phrase is unclear.
SEC staff review registration statements before they go effective, and the review is a letter — filed publicly on EDGAR as form UPLOAD — that walks through the filing comment by comment. Some of those comments look like this:
Comment letter to TCW Private Asset Income Fund · 5 December 2024 · accession 0000000000-24-013422
"In the third sentence, disclosure refers to diversified portfolios of receivables across consumer credit, mortgage credit, small business loans, and other asset classes. Please explain diversified portfolios of receivables using plain English, and clarify what such other asset classes are."
The quoted phrase is a label this project did not create and cannot tune. That makes it the closest thing available to an external audit of the diagnostics this site runs on every text pasted into it: do they fire on the words a securities regulator singled out, more often than on other words from the same document?
Two things about the corpus need saying before any number. It is historical. The phrase "plain English" appears in 1,488 staff letters between 2004 and 2025 and in none dated 2026, against more than two thousand staff letters filed that year; the last is dated 2025-05-09. And most of those letters do not quote a phrase at all. Of 2,249 comments that mention plain English, 1,270 ask for a section rewritten, a footnote moved or a graphic redrawn, and 425 cite the Plain English Handbook or Rule 421. Only about 25% point at particular words. The study is built on those.
2 of 10 diagnostics fire more on flagged phrases than on the filing's other prose
Each flagged phrase was scored for whether each rule fired at all — presence, never a rate, because a rate per thousand words on a four-word phrase measures the denominator. It was then compared with five runs of prose of exactly the same length, cut from the same filing and, where the section could be recognised, the same section. A rule that fires on everything scores no better than chance.
| Diagnostic | Fires on flagged | Fires on matched prose | Difference | 95% interval | Verdict | Predicted |
|---|---|---|---|---|---|---|
| Nominalizations | 30.7% | 23.2% | +7.4 pp | [+4.4, +10.6] | beats chance | beats |
| Hedges | 0.9% | 0.8% | +0.1 pp | [-0.5, +0.8] | no better than chance | beats |
| Wordy phrases | 2.1% | 2.1% | -0.1 pp | [-0.8, +0.8] | no better than chance | beats |
| Passive voice | 8.2% | 10.3% | -2.1 pp | [-3.8, -0.6] | fires less | no |
| Adverbs | 10.4% | 8.2% | +2.2 pp | [+0.3, +4.2] | beats chance | no |
| Expletive openers (there is / it is) | 0.5% | 0.4% | +0.1 pp | [-0.4, +0.9] | no better than chance | no |
| Subject–verb distance | 1.8% | 2.5% | -0.7 pp | [-1.4, +0.1] | no better than chance | no |
| Throat-clearing openers | 0.0% | 0.0% | 0.0 pp | [0.0, 0.0] | no better than chance | no |
| Stacked negatives | 0.3% | 0.6% | -0.3 pp | [-0.6, +0.1] | no better than chance | no |
| Long sentence | 11.6% | 12.3% | -0.7 pp | [-1.2, -0.3] | fires less | no |
nominalizations and adverbs are the rules whose interval clears zero. Nominalizations fire on 31% of flagged phrases against 23% of matched prose — the largest gap in the table, and the one predicted to be largest before any phrase was scored. hedges, wordy phrases, expletive openers (there is / it is), subject–verb distance, throat-clearing openers and stacked negatives do no better than chance. passive voice and long sentence fire less on the flagged phrases than on the prose around them.
Passive voice — the rule behind the busiest tool on this site — fires on 8% of the phrases a regulator called unclear and on 10% of matched prose from the same filings. The regulator was not, on this evidence, objecting to passives.
The reason is visible in the phrases themselves. The median flagged phrase is 4 words long; nine in ten are under 24. They are things like return of capital, pursuant to and deep/machine learning — undefined terms, not badly built clauses. 5 of the 5 structural rules (passive voice, expletive openers (there is / it is), subject–verb distance, throat-clearing openers and stacked negatives) need a subject and a verb to fire, and a noun phrase has neither. A tool built to find sentence-level habits cannot see the problem the regulator saw, because the problem is a word.
Widen the lens to the whole sentence, and the picture barely changes
The regulator flagged a phrase, but the phrase sits inside a sentence and the diagnostics work at sentence level. So the fairer test for the structural rules is the sentence the phrase was found in, against five whole sentences of similar length from the same filing. 1,000 phrases were located in a prose sentence of their filing.
| Diagnostic | Fires on flagged sentence | Fires on matched sentences | Difference | 95% interval | Verdict | Predicted |
|---|---|---|---|---|---|---|
| Nominalizations | 67.6% | 61.2% | +6.4 pp | [+3.0, +9.9] | beats chance | beats |
| Hedges | 2.8% | 3.6% | -0.8 pp | [-2.3, +0.7] | no better than chance | no |
| Wordy phrases | 7.6% | 7.5% | +0.1 pp | [-2.0, +2.4] | no better than chance | no |
| Passive voice | 32.5% | 30.7% | +1.8 pp | [-2.0, +5.4] | no better than chance | no |
| Adverbs | 30.9% | 23.7% | +7.2 pp | [+3.4, +11.2] | beats chance | no |
| Expletive openers (there is / it is) | 1.3% | 2.9% | -1.6 pp | [-2.5, -0.5] | fires less | no |
| Subject–verb distance | 7.3% | 7.2% | +0.1 pp | [-1.8, +2.6] | no better than chance | beats |
| Throat-clearing openers | 0.0% | 0.0% | -0.0 pp | [-0.1, 0.0] | no better than chance | no |
| Stacked negatives | 1.7% | 2.3% | -0.6 pp | [-1.6, +0.4] | no better than chance | no |
| Long sentence | 76.3% | 68.3% | +8.0 pp | [+5.9, +10.2] | beats chance | beats |
nominalizations, adverbs and long sentence clear zero here. The sentences the SEC quoted from run 35.2 words on average against 29.3 for their controls — a gap of 5.88 words [5.13, 6.61] that survives the ±50% length band the controls were drawn within. Restricting the controls to within ±20% of the flagged sentence's length (865 sentences, 32.8 against 29.8 words) leaves long sentence at +1.8 pp [+0.3, +3.1] and nominalizations at +2.0 pp [-1.8, +6.0]. expletive openers (there is / it is) fires less in the flagged sentences than in their neighbours.
Passive voice fires in 33% of the flagged sentences and 31% of the matched ones. It is the house style of the whole document, not a property of the parts the regulator objected to. 91% of flagged sentences trip at least one structural rule — and so do 88% of the sentences around them.
What the regulator flagged, and what the tools said
Twelve flagged phrases chosen by hash, not by hand — the only filter is that the phrase and its sentence fit in a table row — each with the filing sentence it was found in and every rule that fired on that sentence. Read them and the table above stops being abstract: the phrases are terms of art, and the rules are looking for something else.
| Flagged phrase | The sentence in the filing | Rules that fired on the sentence |
|---|---|---|
| return of capitalCommonwealth Credit Partners BDC I, Inc. · 10-12G · 2021 | Distributions in excess of our current and accumulated earnings and profits would be treated first as a return of capital to the extent of the Stockholders tax basis, and any remaining distributions would be treated as a capital gain. | passive voice, long sentence |
| pursuant toFCB Bancorp · S-4/A · 2005 | If this Form is filed to register additional securities for an offering pursuant to Rule 462(b) under the Securities Act, check the following box and list the Securities Act registration statement number of earlier effective registration statement for the same offering. o | nominalizations, passive voice, long sentence |
| deep/machine learningSmartTrust 649 · S-6/A · 2024 | These functional areas include, but are not limited to, (1) deep/machine learning models, applications and platforms, (2) programming languages, (3) hardware, and (4) data centers. | passive voice, long sentence |
| fee titleMVP REIT, Inc. · S-11 · 2012 | We will generally hold fee title or a long-term leasehold estate in the real property we acquire. | adverbs |
| other instrumentsACAP Strategic Fund · N-2 · 2009 | Transactions in these and other instruments may be used in seeking long-term capital appreciation or for hedging purposes. | nominalizations, passive voice |
| risk-adjusted returnsMonachil Credit Income Fund · N-2 · 2021 | The Fund’s primary investment objective is to provide investors with current income and attractive risk-adjusted returns with low correlation to the equity and fixed income markets. | nominalizations, long sentence |
| funds of fundsGOTTEX TRUST · N-1A · 2013 | Private equity and special opportunity investments include interests in private equity or venture capital funds, or private equity “funds of funds,” the interests of which are expected to be publicly traded, or in other pooled vehicles that seek direct or indirect exposure to this Asset Class. | passive voice, adverbs, long sentence |
| accretive returnsNEXPOINT HEALTHCARE OPPORTUNITIES FUND · N-2 · 2016 | The intent of these borrowings will be to provide accretive returns to invested Fund equity. | none |
| yield standard deviationGUGGENHEIM DEFINED PORTFOLIOS, SERIES 1704 · S-6/A · 2018 | o Yield Standard Deviation: Calculate each remaining security's indicative yield by taking the most recently announced board-approved monthly or quarterly dividend and multiplying it by 12 or 4, respectively, and dividing by its closing price. | nominalizations, adverbs, subject–verb distance, long sentence |
| reasonable effortsSEPARATE ACCOUNT NO. 49 · N-4/A · 2010 | We will use reasonable efforts to allocate your Segment Maturity Value in accordance with your instructions, which may include holding amounts in Segment Type Holding Accounts until the next Segment Start Date. | nominalizations, long sentence |
| oral suspensionACCENTIA BIOPHARMACEUTICALS INC · S-1 · 2005 | We anticipate that the new drug application, or NDA, for SinuNase will be filed as a 505(b)(2) application, which will enable us to rely in part on the FDAs previous findings of safety and efficacy for an oral suspension of amphotericin B. | nominalizations, passive voice, long sentence |
| may include, but are not limited toPursuit Asset-Based Income Fund · N-2 · 2025 | These assets may include, but are not limited to: | passive voice |
Written down before the first phrase was scored
A tool audit run by the tool's own authors is a foregone conclusion unless the authors say in advance what they expect. Before any span was scored we recorded, for each arm, the rank order the diagnostics were expected to fall in and which were expected not to beat their control rate. The prediction was that most would not, because the regulator flags undefined jargon and the diagnostics look for sentence structure.
- Phrase arm. Predicted to beat chance: nominalizations, hedges and wordy phrases. Observed: nominalizations and adverbs. 7 of 10 verdicts as predicted; rank correlation between predicted and observed order 0.36.
- Sentence arm. Predicted to beat chance: nominalizations, subject–verb distance and long sentence. Observed: nominalizations, adverbs and long sentence. 8 of 10 verdicts as predicted; rank correlation 0.90.
The full pre-commitment, including two changes made after reading extraction output and before any score, is recorded in the study's verification log alongside the pipeline (scripts/studies/sec-comment-letters/VERIFICATION.md).
How this was measured
The population is every staff comment letter in EDGAR full-text search matching the phrase "plain English" — 1,488 letters after deduplicating the index, which counts a modern letter twice. Each letter was read from the document the SEC filed: native text in 2004–2005, PDF from 2009, a mix between. Not from the SEC's own plain-text extract, which replaces quotation marks and apostrophes alike with runs of spaces and takes phrase recovery from 87% to 6.5% on the same letters.
A phrase is whatever the staff put inside quotation marks in a comment that mentions plain English and is not a citation to the Handbook or the rule. The exclusions, all on structure and never on the words: 98 byte-identical letters sent to two registrants, 57 blocks of the filer's own replies, 24 quotations of the phrase "plain English" itself, 136 quoted section headings ("the text under 'Arrangements to Minimize Exposure'"), and every single-word quotation, because a lone word carries no structure for any rule to find. Raising that floor to three words or five moves the nominalization row to +5.6 and +3.1 points (943 and 587 phrases).
The filing each letter reviews was found by enumerating the registrant's EDGAR filings in the 400 days before the letter and accepting the first that contains one of the quoted phrases verbatim. That test is the linkage — the header field meant for the purpose is present on a fifth of letters and names nothing usable. 482 of 585 letters were linked this way; 1043 of 1546 phrases were found verbatim, 1000 of them inside a prose sentence. Phrases from unlinked letters have no controls and are not in the tables. Their uncontrolled nominalization rate is 41%, against 31% for the linked ones.
Controls for a phrase are five runs of the same word count cut at a deterministic offset from prose sentences of the same filing — the same section where the section could be recognised and held enough candidates (71% of phrases), otherwise anywhere in the filing. The pre-committed rule was whole sentences within ±50% of the phrase's length; at two to four words that produced name fragments and page numbers, which no rule can fire on, and every rule would have "beaten" them. The change was made on reading the control draw and before any score, and the pre-committed variant is still in the per-phrase file. Controls for a sentence are five whole sentences within ±50% of its length, same filing, same section where possible.
Every rule is the one that runs on this site — the same code behind the plain language checker, the passive voice finder and the readability explainer — called sentence by sentence and asked whether it fired. Confidence intervals come from a filer-clustered bootstrap (1,000 replicates, seeded): the phrases in one letter share a reviewer, a filing and a drafting firm, and a comment reissued in a follow-up letter shares all three again.
What this cannot tell you
- It cannot say the phrases are unclear. It says a regulator called them unclear. The SEC's staff are one reader with one purpose, and some of what they quoted — marketing slogans, defined terms in capitals — was flagged for reasons other than comprehension.
- A comment that quotes several phrases contributes all of them. Where one comment quoted a phrase for plain English and another for substantiation, both are in the data. The direction of that error is toward the null.
- The linked letters are not a random subset. A phrase is found verbatim when the extraction was clean and the staff quoted exactly; a paraphrase, an ellipsis or a PDF whose text layer ran words together is not. The uncontrolled rate on the unlinked phrases is reported above so the gap is visible.
- Nothing here measures what the company did next. The brief for this study had a second arm — score the next amendment — and it is not built. The filing a letter reviews is a registration statement, not a final prospectus, and the parser this site has was validated only on the latter.
- The corpus ends in 2025. Whether the staff stopped invoking plain English or stopped using those words, this study cannot say.
Check our work
Everything below is the actual output of the analysis. The per-phrase file carries the phrase, the letter's accession number, the filing's accession number where linked, and every rule's verdict on the phrase, its sentence and their controls — enough to recompute any row of either table, or to fetch both originals from EDGAR and disagree with the extraction.
- Per-phrase scores — 1,546 rows, one per flagged phrase
- Audit table — both arms, every rule, with intervals and the pre-registered prediction
Published September 13, 2026. Comment letters and filings are public records; the analysis is free to reuse with attribution. Corrections are welcome — get in touch.