Skip to content

Commit efff605

Browse files
peopleworksclaude
andcommitted
Measure the writers this category harms, and let the boundary move
The calibration corpus gains 206 classroom essays by adult learners of English (PELIC, University of Pittsburgh, 2006-2012): first submitted version, writing classes, 662+ words, one text per student, selected by a fixed rule so `fetch --source pelic` yields the same ids and hashes on any machine. The affiliation proxy stands down; the population is here. At the boundary the corpus supported before they joined, 25/100, the tool flagged 9 of their 206 essays and none of the 90 published texts. The boundary is now 30/100 (2 of 296, interval 0.2-2.4%), English has its own threshold for the first time, and the report paragraph that said no group sits above the rest now computes who is flagged at the recommended boundary and says so. The learners also caught a defect: chat.eager-opener, admitted on zero hits in published prose, fired in eleven essays because its pattern took "Of course," with a comma. It now requires the exclamation mark, in both packs, with tests that the old pattern fails. Not done, deliberately: refitting the human-rate gates on the pooled corpus. With learners as the majority, twelve rules lose their gate and the connector gates inflate up to seven-fold. That is issue #75. Version 0.6.0: the first change since 0.5.0 that touches Core. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_015PEbbiYSNPw7jE3LrPNhyF
1 parent dde7f01 commit efff605

29 files changed

Lines changed: 2931 additions & 182 deletions

‎Docs/CALIBRATION.md‎

Lines changed: 59 additions & 58 deletions
Original file line numberDiff line numberDiff line change
@@ -8,20 +8,20 @@ It is **not an accuracy figure**. Accuracy needs machine-written text to measure
88

99
## What was measured
1010

11-
- **Corpus** `signsofai-human-baseline`, fingerprint `123fa5b9ebca3f29`
12-
- **Texts** 90 (280,221 words)
13-
- **Lengths measured** 662 – 9,328 words (median 2,772)
14-
- **Engine** SignsOfAI.Core 0.5.0
15-
- **Run** 2026-08-24
11+
- **Corpus** `signsofai-human-baseline`, fingerprint `78bda061bde3dc99`
12+
- **Texts** 296 (452,184 words)
13+
- **Lengths measured** 649 – 9,328 words (median 832)
14+
- **Engine** SignsOfAI.Core 0.6.0
15+
- **Run** 2026-09-01
1616
- **Target false-positive rate** 5%
1717

18-
Every text here was published before generative models could have written it. That is the whole basis for calling it human, and it is a stronger guarantee than any classifier offers about anything. The manifest names each source, its licence and its year, so the claim can be traced rather than trusted.
18+
Every text here was written before generative models could have written it — articles and encyclopedia revisions with a date, and classroom essays from a learner corpus collected years earlier. That is the whole basis for calling it human, and it is a stronger guarantee than any classifier offers about anything. The manifest names each source, its licence and its year, so the claim can be traced rather than trusted.
1919

2020
## The headline
2121

22-
**At a threshold of 25/100, this tool flags at most 5% of writing known to be human** — 0 of 90 texts in this corpus, an observed 0% with a 95% interval of 0% – 4.1%.
22+
**At a threshold of 30/100, this tool flags at most 5% of writing known to be human** — 2 of 296 texts in this corpus, an observed 0.7% with a 95% interval of 0.2% – 2.4%.
2323

24-
**It covers documents of 662 words and up, because that is what was measured.** Nothing shorter was: the corpus has no text below that length, so the boundary below is not supported there and the tool withholds its verdict rather than extrapolating. That is a statement about coverage, not about where the tool breaks — though the direction of the length effect *has* been measured, and it goes the wrong way: the same documents flagged 0 of 32 whole and 6 of 32 as 400-word excerpts of themselves (`Docs/PARAPHRASE.md`, section *Length*). Lowering this floor means measuring short writing people actually composed at that length, not slicing long documents into pieces.
24+
**It covers documents of 649 words and up, because that is what was measured.** Nothing shorter was: the corpus has no text below that length, so the boundary below is not supported there and the tool withholds its verdict rather than extrapolating. That is a statement about coverage, not about where the tool breaks — though the direction of the length effect *has* been measured, and it goes the wrong way: the same documents flagged 0 of 32 whole and 6 of 32 as 400-word excerpts of themselves (`Docs/PARAPHRASE.md`, section *Length*). Lowering this floor means measuring short writing people actually composed at that length, not slicing long documents into pieces.
2525

2626
Read the interval, not the percentage. On a small corpus an observed rate is compatible with a much wider range, and the recommendation below is made from the **upper** end of that range rather than the flattering one — so it stays cautious while the corpus is thin and tightens on its own as it grows.
2727

@@ -31,7 +31,7 @@ A rate that holds in English and fails in Spanish is not one number, and reporti
3131

3232
| Group | Texts | Median | 90th pct | Highest | Threshold for 5% | Best bound it can support |
3333
|---|---|---|---|---|---|---|
34-
| **en** | 65 | 5.8 | 11.8 | 23.4 | — | 5.6% |
34+
| **en** | 271 | 8.8 | 18.3 | 33.8 | 30 | 1.4% |
3535
| **es** | 25 | 7.2 | 15.1 | 18.4 | — | 13.3% |
3636

3737
A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.
@@ -44,41 +44,42 @@ The reason the whole exercise exists. If this project cannot show a rate for sec
4444
|---|---|---|---|---|---|---|
4545
| **en-anglophone-affiliation** | 21 | 5.9 | 9.0 | 14.2 | — | 15.5% |
4646
| **en-other-affiliation** | 19 | 6.2 | 13.6 | 18.0 | — | 16.8% |
47+
| **en-second-language-learner** | 206 | 9.6 | 19.7 | 33.8 | 30 | 1.8% |
4748
| **en-wikipedia** | 25 | 4.9 | 10.4 | 23.4 | — | 13.3% |
4849
| **es-wikipedia** | 25 | 7.2 | 15.1 | 18.4 | — | 13.3% |
4950

5051
A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.
5152

52-
Across these groups the median score runs from 7.2 (**es-wikipedia**) down to 4.9 (**en-wikipedia**), a spread of 2.3 points on a scale of a hundred. The longest tail belongs to **es-wikipedia** at 15.1 for the ninetieth percentile. A tool with the defect this project criticises would show one group sitting well above the rest; on this corpus none does. It is a first indication rather than a finding — these are tens of texts, not hundreds — and the numbers move as the corpus grows, in whichever direction they move.
53+
Across these groups the median score runs from 9.6 (**en-second-language-learner**) down to 4.9 (**en-wikipedia**), a spread of 4.7 points on a scale of a hundred. The longest tail belongs to **en-second-language-learner** at 19.7 for the ninetieth percentile. At the boundary this page recommends, 30/100, **en-second-language-learner** is flagged 2 of 206 (1%, interval 0.3% – 3.5%); every other group is flagged nothing at all. It also sits highest in median and ninetieth percentile. That is the shape of the defect this project criticises, and it is reported here rather than averaged away — smaller than the figures published for other tools, which is a comparison, not an excuse. The groups run from tens of texts to a couple of hundred, and the numbers move as the corpus grows, in whichever direction they move.
5354

5455
## Every threshold
5556

5657
| Score at or above | Human texts flagged | Rate | 95% interval |
5758
|---|---|---|---|
58-
| 5 | 61 / 90 | 67.8% | 57.6% – 76.5% |
59-
| 10 | 16 / 90 | 17.8% | 11.2% – 26.9% |
60-
| 15 | 6 / 90 | 6.7% | 3.1% – 13.8% |
61-
| 20 | 1 / 90 | 1.1% | 0.2% – 6% |
62-
| 25 | 0 / 90 | 0% | 0% – 4.1% |
63-
| 30 | 0 / 90 | 0% | 0% – 4.1% |
64-
| 35 | 0 / 90 | 0% | 0% – 4.1% |
65-
| 40 | 0 / 90 | 0% | 0% – 4.1% |
66-
| 45 | 0 / 90 | 0% | 0% – 4.1% |
67-
| 50 | 0 / 90 | 0% | 0% – 4.1% |
68-
| 55 | 0 / 90 | 0% | 0% – 4.1% |
69-
| 60 | 0 / 90 | 0% | 0% – 4.1% |
70-
| 65 | 0 / 90 | 0% | 0% – 4.1% |
71-
| 70 | 0 / 90 | 0% | 0% – 4.1% |
72-
| 75 | 0 / 90 | 0% | 0% – 4.1% |
73-
| 80 | 0 / 90 | 0% | 0% – 4.1% |
74-
| 85 | 0 / 90 | 0% | 0% – 4.1% |
75-
| 90 | 0 / 90 | 0% | 0% – 4.1% |
76-
| 95 | 0 / 90 | 0% | 0% – 4.1% |
77-
| 100 | 0 / 90 | 0% | 0% – 4.1% |
59+
| 5 | 253 / 296 | 85.5% | 81% – 89% |
60+
| 10 | 113 / 296 | 38.2% | 32.8% – 43.8% |
61+
| 15 | 49 / 296 | 16.6% | 12.8% – 21.2% |
62+
| 20 | 21 / 296 | 7.1% | 4.7% – 10.6% |
63+
| 25 | 9 / 296 | 3% | 1.6% – 5.7% |
64+
| 30 | 2 / 296 | 0.7% | 0.2% – 2.4% |
65+
| 35 | 0 / 296 | 0% | 0% – 1.3% |
66+
| 40 | 0 / 296 | 0% | 0% – 1.3% |
67+
| 45 | 0 / 296 | 0% | 0% – 1.3% |
68+
| 50 | 0 / 296 | 0% | 0% – 1.3% |
69+
| 55 | 0 / 296 | 0% | 0% – 1.3% |
70+
| 60 | 0 / 296 | 0% | 0% – 1.3% |
71+
| 65 | 0 / 296 | 0% | 0% – 1.3% |
72+
| 70 | 0 / 296 | 0% | 0% – 1.3% |
73+
| 75 | 0 / 296 | 0% | 0% – 1.3% |
74+
| 80 | 0 / 296 | 0% | 0% – 1.3% |
75+
| 85 | 0 / 296 | 0% | 0% – 1.3% |
76+
| 90 | 0 / 296 | 0% | 0% – 1.3% |
77+
| 95 | 0 / 296 | 0% | 0% – 1.3% |
78+
| 100 | 0 / 296 | 0% | 0% – 1.3% |
7879

7980
## What the product does with this number
8081

81-
The tool speaks at **25/100** and nowhere else, taking the boundary from the table above rather than from anybody's judgement. Below it a document gets its score and the reason it gets nothing more: a low score is not evidence that a person wrote something, since a detector that detects nothing also returns a low score, and this project has deliberately never measured how much machine writing it catches. The boundary moves when this page moves — including upward if a larger corpus turns out to be less flattering.
82+
The tool speaks at **30/100** and nowhere else, taking the boundary from the table above rather than from anybody's judgement. Below it a document gets its score and the reason it gets nothing more: a low score is not evidence that a person wrote something, since a detector that detects nothing also returns a low score, and this project has deliberately never measured how much machine writing it catches. The boundary moves when this page moves — including upward if a larger corpus turns out to be less flattering.
8283

8384
Above it there is **one** verdict, not a scale of them. This corpus can place a boundary and can say nothing whatever about how much further past it a score has travelled: no text known to be human came close to the upper reaches, and grading "moderate" against "strong" would need machine-written text, which the opening of this page argues against collecting. Interfaces do shade a high score more urgently than a low one, and those shades are a display convention — they are not on this page because nothing measured them.
8485

@@ -90,39 +91,39 @@ Every rule below fired on text no machine wrote, so each hit is a false positive
9091

9192
| Rule | Texts it fired on | Share | Total hits |
9293
|---|---|---|---|
93-
| `stat.burstiness` | 25 | 27.8% | 25 |
94-
| `rhet.in-terms-of` | 9 | 10% | 22 |
95-
| `rhet.not-only-but` | 8 | 8.9% | 16 |
96-
| `rhet.in-order-to` | 7 | 7.8% | 33 |
97-
| `lex.furthermore` | 7 | 7.8% | 24 |
98-
| `lex.robust` | 7 | 7.8% | 19 |
99-
| `lex.just` | 7 | 7.8% | 16 |
100-
| `lex.simply` | 7 | 7.8% | 16 |
101-
| `lex.utilizar` | 7 | 7.8% | 15 |
102-
| `rhet.in-this-article` | 7 | 7.8% | 15 |
103-
| `lex.notably` | 7 | 7.8% | 14 |
104-
| `syn.superficial-ing` | 7 | 7.8% | 12 |
105-
| `rhet.with-regard-to` | 7 | 7.8% | 10 |
106-
| `lex.crucial` | 7 | 7.8% | 8 |
107-
| `syn.serves-as` | 7 | 7.8% | 8 |
108-
| `lex.moreover` | 6 | 6.7% | 35 |
109-
| `lex.comprehensive` | 6 | 6.7% | 27 |
110-
| `lex.utilize` | 6 | 6.7% | 22 |
111-
| `lex.facilitate` | 6 | 6.7% | 16 |
112-
| `rhet.rule-of-three` | 6 | 6.7% | 16 |
113-
| `lex.importantly` | 6 | 6.7% | 9 |
114-
| `rhet.weasel-attribution` | 6 | 6.7% | 9 |
115-
| `lex.actually` | 6 | 6.7% | 8 |
116-
| `rhet.in-conclusion` | 6 | 6.7% | 8 |
117-
| `rhet.important-note` | 6 | 6.7% | 7 |
94+
| `lex.just` | 103 | 34.8% | 189 |
95+
| `stat.burstiness` | 95 | 32.1% | 95 |
96+
| `rhet.in-conclusion` | 89 | 30.1% | 99 |
97+
| `rhet.rule-of-three` | 77 | 26% | 143 |
98+
| `lex.moreover` | 66 | 22.3% | 126 |
99+
| `rhet.not-only-but` | 63 | 21.3% | 90 |
100+
| `lex.furthermore` | 55 | 18.6% | 92 |
101+
| `lex.crucial` | 50 | 16.9% | 61 |
102+
| `lex.actually` | 49 | 16.6% | 69 |
103+
| `rhet.in-order-to` | 41 | 13.9% | 87 |
104+
| `rhet.in-terms-of` | 22 | 7.4% | 38 |
105+
| `rhet.weasel-attribution` | 19 | 6.4% | 32 |
106+
| `rhet.in-this-article` | 16 | 5.4% | 24 |
107+
| `lex.simply` | 14 | 4.7% | 23 |
108+
| `lex.facilitate` | 12 | 4.1% | 22 |
109+
| `lex.utilize` | 10 | 3.4% | 30 |
110+
| `lex.truly` | 10 | 3.4% | 15 |
111+
| `lex.comprehensive` | 9 | 3% | 30 |
112+
| `rhet.when-it-comes` | 9 | 3% | 10 |
113+
| `syn.serves-as` | 9 | 3% | 10 |
114+
| `lex.notably` | 8 | 2.7% | 15 |
115+
| `rhet.with-regard-to` | 8 | 2.7% | 11 |
116+
| `rhet.on-one-hand` | 8 | 2.7% | 9 |
117+
| `lex.profound` | 8 | 2.7% | 8 |
118+
| `lex.robust` | 7 | 2.4% | 19 |
118119

119120
A rule near the top is not automatically wrong. Some tells genuinely appear in human academic prose and the catalog says so. But a rule firing on most human texts is measuring the genre rather than the machine, and should be reweighted or retired.
120121

121122
## What this does not tell you
122123

123124
- **Nothing about how much AI writing it catches.** That is the other half of the picture and it is not measured here, deliberately. A tool that flags nothing has a perfect false-positive rate.
124-
- **Nothing about text unlike this corpus.** These are published articles. A first-year essay is shorter, looser and differently edited, and the rate on one does not transfer to the other. Calibrating on your own students' pre-2022 work is the fix, and the same tool does it.
125-
- **The grouping of writers is a proxy, not a fact.** Nobody's first language is recorded in a DOI. The manifest states the reasoning per text so it can be argued with; where it is wrong, the number moves.
125+
- **Nothing about text unlike this corpus.** These are published articles and the essays of adult learners in a university English programme. A first-year essay by a native speaker is a different population again, and the rate on one does not transfer to the other. Calibrating on your own students' pre-2022 work is the fix, and the same tool does it.
126+
- **The affiliation groups are a proxy, not a fact.** Nobody's first language is recorded in a DOI. The learner group is the exception — its corpus records each writer's first language — which is why it exists. Elsewhere the manifest states the reasoning per text so it can be argued with; where it is wrong, the number moves.
126127
- **Hashes prove what *this* run measured**, not that another person extracting the same articles would get identical text. They would not: PDF and HTML extraction differ. Reproducing this needs the extracted texts, not just the manifest.
127128

128129
Re-run it yourself:

‎Docs/Calibration/README.md‎

Lines changed: 36 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -29,10 +29,11 @@ and cannot also be the thing doing the measuring.
2929

3030
## What the corpus does not cover, and what it costs
3131

32-
Every text here is **662 words or longer** — that is the shortest one, and the Wikipedia fetcher skips
33-
anything under 700 by design. The threshold is therefore supported over that range and nowhere else,
34-
so since #59 the engine **withholds its verdict below 662 words** rather than extrapolating onto a
35-
population it never sampled.
32+
Every text here is **649 words or longer** as the engine counts them — that is the shortest one; the
33+
Wikipedia fetcher skips anything under 700 by design and the learner selection anything under 662 by
34+
a naive split. The threshold is therefore supported over that range and nowhere else, so since #59
35+
the engine **withholds its verdict below that length** rather than extrapolating onto a population it
36+
never sampled.
3637

3738
That is not a small exclusion. It is most of how the tool is used: somebody pastes a paragraph. And
3839
the direction of the error is known — the same documents flag 0 of 32 whole and 6 of 32 as 400-word
@@ -65,6 +66,10 @@ dotnet run --project tools/SignsOfAI.Calibration -- fetch --source plos --count
6566
dotnet run --project tools/SignsOfAI.Calibration -- fetch --source wikipedia --lang en --count 25
6667
dotnet run --project tools/SignsOfAI.Calibration -- fetch --source wikipedia --lang es --count 25
6768

69+
# 206 essays by adult learners of English, 2006–2012, one per student — a fixed rule, no --count,
70+
# so anyone running it gets the same texts and the same hashes (about 180 MB downloaded once)
71+
dotnet run --project tools/SignsOfAI.Calibration -- fetch --source pelic
72+
6873
# measure, and rewrite Docs/CALIBRATION.md
6974
dotnet run --project tools/SignsOfAI.Calibration -- run
7075
```
@@ -78,7 +83,8 @@ having.
7883
| Group | What it is | What it is for |
7984
|---|---|---|
8085
| `en-anglophone-affiliation` | PLOS articles with at least one author affiliated in an anglophone country | The comparison baseline |
81-
| `en-other-affiliation` | PLOS articles with no anglophone affiliation | **Standing in for second-language English** — the population this whole category harms |
86+
| `en-other-affiliation` | PLOS articles with no anglophone affiliation | Was standing in for second-language English until the learner group arrived; kept, because it is the same question asked of professional writers |
87+
| `en-second-language-learner` | Classroom essays from PELIC — University of Pittsburgh's Intensive English Program, 2006–2012, first language recorded per writer | **The population this whole category harms**, measured directly rather than through a proxy |
8288
| `en-wikipedia` | Pre-2022 English Wikipedia revisions | Same register as the Spanish group, so a language effect can be told from a register effect |
8389
| `es-wikipedia` | Pre-2022 Spanish Wikipedia revisions | The half nobody else measures at all |
8490

@@ -88,11 +94,28 @@ direction: a paper with *any* anglophone affiliation counts as anglophone, which
8894
second-language group and makes any gap found an understatement rather than an exaggeration. Every
8995
entry records the affiliation used, so each classification can be argued with individually.
9096

91-
**None of this is a student essay.** Published articles are longer, more heavily edited and written by
92-
people who write for a living. A first-year essay is a different thing and the rate measured on one
93-
does not transfer to the other. The honest fix is for a school to calibrate on its own students'
94-
pre-2022 work — the same tool does it, and the result would be a false-positive rate for *its*
95-
population instead of somebody else's.
97+
**The learner group is the one that is not a proxy.** [PELIC](https://github.com/ELI-Data-Mining-Group/PELIC-dataset)
98+
records each writer's first language — Arabic, Korean, Chinese, Japanese, Spanish, Thai and Turkish
99+
make up most of it — and every essay was written years before generative models, in a classroom,
100+
under a prompt. The selection is a rule rather than a choice: first submitted version, writing
101+
classes only, at least 662 words so the group enters at the floor the corpus already had, one text
102+
per student so nobody prolific counts twice. Nothing is picked by score. The licence is
103+
CC BY-NC-ND 4.0, which permits measuring and publishing the numbers and forbids redistributing the
104+
texts — the same arrangement every other source here already has.
105+
106+
What it found is on `Docs/CALIBRATION.md` and it is the reason the boundary moved from 25 to 30: at
107+
25, 9 of the 206 essays were flagged and none of the 90 published texts. The rules doing it are the
108+
connectors an academic-English course teaches (`rhet.in-conclusion` fires in 40% of learner essays
109+
and 14% of published ones) — see issue #75. It also caught a defect: `chat.eager-opener` had been
110+
admitted on zero hits in the published texts and fired on eleven learner essays, because its pattern
111+
accepted *"Of course, …"* with a comma, an ordinary concession, alongside *"Certainly!"*. Zero on one
112+
register is not zero.
113+
114+
**Still no native-speaker student essay.** Published articles are written by people who write for a
115+
living; the learner essays are by adults in a university language programme. A first-year essay by a
116+
native speaker is a different population again, and the rate measured here does not transfer to it.
117+
The honest fix is for a school to calibrate on its own students' pre-2022 work — the same tool does
118+
it, and the result would be a false-positive rate for *its* population instead of somebody else's.
96119

97120
## Contributing texts
98121

@@ -109,8 +132,9 @@ What is most wanted, in order:
109132
- **Spanish academic writing.** SciELO and Redalyc are the obvious sources and neither was reachable
110133
from where this was first assembled. Spanish is the half of this project nobody else measures, and
111134
it currently rests on encyclopedia prose alone.
112-
- **Anything closer to a student essay.** Coursework released under an open licence, pre-2022 writing
113-
competition entries, open thesis repositories.
135+
- **Native-speaker student writing.** Coursework released under an open licence, pre-2022 writing
136+
competition entries, open thesis repositories. The learner group covers second-language writers;
137+
nothing yet covers a first-year student writing in their own language.
114138
- **More of everything.** With nothing flagged it still takes roughly seventy-five texts in a group
115139
before the interval alone can bound a 5% rate. Most groups here are half that.
116140

0 commit comments

Comments
 (0)