You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Measure the writers this category harms, and let the boundary move
The calibration corpus gains 206 classroom essays by adult learners of
English (PELIC, University of Pittsburgh, 2006-2012): first submitted
version, writing classes, 662+ words, one text per student, selected by
a fixed rule so `fetch --source pelic` yields the same ids and hashes on
any machine. The affiliation proxy stands down; the population is here.
At the boundary the corpus supported before they joined, 25/100, the
tool flagged 9 of their 206 essays and none of the 90 published texts.
The boundary is now 30/100 (2 of 296, interval 0.2-2.4%), English has
its own threshold for the first time, and the report paragraph that
said no group sits above the rest now computes who is flagged at the
recommended boundary and says so.
The learners also caught a defect: chat.eager-opener, admitted on zero
hits in published prose, fired in eleven essays because its pattern
took "Of course," with a comma. It now requires the exclamation mark,
in both packs, with tests that the old pattern fails.
Not done, deliberately: refitting the human-rate gates on the pooled
corpus. With learners as the majority, twelve rules lose their gate and
the connector gates inflate up to seven-fold. That is issue #75.
Version 0.6.0: the first change since 0.5.0 that touches Core.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_015PEbbiYSNPw7jE3LrPNhyF
-**Lengths measured**649 – 9,328 words (median 832)
14
+
-**Engine** SignsOfAI.Core 0.6.0
15
+
-**Run** 2026-09-01
16
16
-**Target false-positive rate** 5%
17
17
18
-
Every text here was published before generative models could have written it. That is the whole basis for calling it human, and it is a stronger guarantee than any classifier offers about anything. The manifest names each source, its licence and its year, so the claim can be traced rather than trusted.
18
+
Every text here was written before generative models could have written it — articles and encyclopedia revisions with a date, and classroom essays from a learner corpus collected years earlier. That is the whole basis for calling it human, and it is a stronger guarantee than any classifier offers about anything. The manifest names each source, its licence and its year, so the claim can be traced rather than trusted.
19
19
20
20
## The headline
21
21
22
-
**At a threshold of 25/100, this tool flags at most 5% of writing known to be human** — 0 of 90 texts in this corpus, an observed 0% with a 95% interval of 0% – 4.1%.
22
+
**At a threshold of 30/100, this tool flags at most 5% of writing known to be human** — 2 of 296 texts in this corpus, an observed 0.7% with a 95% interval of 0.2% – 2.4%.
23
23
24
-
**It covers documents of 662 words and up, because that is what was measured.** Nothing shorter was: the corpus has no text below that length, so the boundary below is not supported there and the tool withholds its verdict rather than extrapolating. That is a statement about coverage, not about where the tool breaks — though the direction of the length effect *has* been measured, and it goes the wrong way: the same documents flagged 0 of 32 whole and 6 of 32 as 400-word excerpts of themselves (`Docs/PARAPHRASE.md`, section *Length*). Lowering this floor means measuring short writing people actually composed at that length, not slicing long documents into pieces.
24
+
**It covers documents of 649 words and up, because that is what was measured.** Nothing shorter was: the corpus has no text below that length, so the boundary below is not supported there and the tool withholds its verdict rather than extrapolating. That is a statement about coverage, not about where the tool breaks — though the direction of the length effect *has* been measured, and it goes the wrong way: the same documents flagged 0 of 32 whole and 6 of 32 as 400-word excerpts of themselves (`Docs/PARAPHRASE.md`, section *Length*). Lowering this floor means measuring short writing people actually composed at that length, not slicing long documents into pieces.
25
25
26
26
Read the interval, not the percentage. On a small corpus an observed rate is compatible with a much wider range, and the recommendation below is made from the **upper** end of that range rather than the flattering one — so it stays cautious while the corpus is thin and tightens on its own as it grows.
27
27
@@ -31,7 +31,7 @@ A rate that holds in English and fails in Spanish is not one number, and reporti
31
31
32
32
| Group | Texts | Median | 90th pct | Highest | Threshold for 5% | Best bound it can support |
33
33
|---|---|---|---|---|---|---|
34
-
|**en**|65|5.8 |11.8|23.4|—|5.6% |
34
+
|**en**|271|8.8 |18.3|33.8|30|1.4% |
35
35
|**es**| 25 | 7.2 | 15.1 | 18.4 | — | 13.3% |
36
36
37
37
A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.
@@ -44,41 +44,42 @@ The reason the whole exercise exists. If this project cannot show a rate for sec
A dash means this group has too few texts to bound that rate at all — with nothing flagged it still takes roughly seventy-five before the interval alone gets under 5%. That is a statement about the corpus, not the tool.
51
52
52
-
Across these groups the median score runs from 7.2 (**es-wikipedia**) down to 4.9 (**en-wikipedia**), a spread of 2.3 points on a scale of a hundred. The longest tail belongs to **es-wikipedia** at 15.1 for the ninetieth percentile. A tool with the defect this project criticises would show one group sitting well above the rest; on this corpus none does. It is a first indication rather than a finding — these are tens of texts, not hundreds — and the numbers move as the corpus grows, in whichever direction they move.
53
+
Across these groups the median score runs from 9.6 (**en-second-language-learner**) down to 4.9 (**en-wikipedia**), a spread of 4.7 points on a scale of a hundred. The longest tail belongs to **en-second-language-learner** at 19.7 for the ninetieth percentile. At the boundary this page recommends, 30/100, **en-second-language-learner** is flagged 2 of 206 (1%, interval 0.3% – 3.5%); every other group is flagged nothing at all. It also sits highest in median and ninetieth percentile. That is the shape of the defect this project criticises, and it is reported here rather than averaged away — smaller than the figures published for other tools, which is a comparison, not an excuse. The groups run from tens of texts to a couple of hundred, and the numbers move as the corpus grows, in whichever direction they move.
53
54
54
55
## Every threshold
55
56
56
57
| Score at or above | Human texts flagged | Rate | 95% interval |
57
58
|---|---|---|---|
58
-
| 5 |61 / 90|67.8% |57.6% – 76.5% |
59
-
| 10 |16 / 90|17.8% |11.2% – 26.9% |
60
-
| 15 |6 / 90|6.7% |3.1% – 13.8% |
61
-
| 20 |1 / 90|1.1% |0.2% – 6% |
62
-
| 25 |0 / 90|0% |0% – 4.1% |
63
-
| 30 |0 / 90| 0% | 0% – 4.1% |
64
-
| 35 | 0 / 90| 0% | 0% – 4.1% |
65
-
| 40 | 0 / 90| 0% | 0% – 4.1% |
66
-
| 45 | 0 / 90| 0% | 0% – 4.1% |
67
-
| 50 | 0 / 90| 0% | 0% – 4.1% |
68
-
| 55 | 0 / 90| 0% | 0% – 4.1% |
69
-
| 60 | 0 / 90| 0% | 0% – 4.1% |
70
-
| 65 | 0 / 90| 0% | 0% – 4.1% |
71
-
| 70 | 0 / 90| 0% | 0% – 4.1% |
72
-
| 75 | 0 / 90| 0% | 0% – 4.1% |
73
-
| 80 | 0 / 90| 0% | 0% – 4.1% |
74
-
| 85 | 0 / 90| 0% | 0% – 4.1% |
75
-
| 90 | 0 / 90| 0% | 0% – 4.1% |
76
-
| 95 | 0 / 90| 0% | 0% – 4.1% |
77
-
| 100 | 0 / 90| 0% | 0% – 4.1% |
59
+
| 5 |253 / 296|85.5% |81% – 89% |
60
+
| 10 |113 / 296|38.2% |32.8% – 43.8% |
61
+
| 15 |49 / 296|16.6% |12.8% – 21.2% |
62
+
| 20 |21 / 296|7.1% |4.7% – 10.6% |
63
+
| 25 |9 / 296|3% |1.6% – 5.7% |
64
+
| 30 |2 / 296| 0.7% | 0.2% – 2.4% |
65
+
| 35 | 0 / 296| 0% | 0% – 1.3% |
66
+
| 40 | 0 / 296| 0% | 0% – 1.3% |
67
+
| 45 | 0 / 296| 0% | 0% – 1.3% |
68
+
| 50 | 0 / 296| 0% | 0% – 1.3% |
69
+
| 55 | 0 / 296| 0% | 0% – 1.3% |
70
+
| 60 | 0 / 296| 0% | 0% – 1.3% |
71
+
| 65 | 0 / 296| 0% | 0% – 1.3% |
72
+
| 70 | 0 / 296| 0% | 0% – 1.3% |
73
+
| 75 | 0 / 296| 0% | 0% – 1.3% |
74
+
| 80 | 0 / 296| 0% | 0% – 1.3% |
75
+
| 85 | 0 / 296| 0% | 0% – 1.3% |
76
+
| 90 | 0 / 296| 0% | 0% – 1.3% |
77
+
| 95 | 0 / 296| 0% | 0% – 1.3% |
78
+
| 100 | 0 / 296| 0% | 0% – 1.3% |
78
79
79
80
## What the product does with this number
80
81
81
-
The tool speaks at **25/100** and nowhere else, taking the boundary from the table above rather than from anybody's judgement. Below it a document gets its score and the reason it gets nothing more: a low score is not evidence that a person wrote something, since a detector that detects nothing also returns a low score, and this project has deliberately never measured how much machine writing it catches. The boundary moves when this page moves — including upward if a larger corpus turns out to be less flattering.
82
+
The tool speaks at **30/100** and nowhere else, taking the boundary from the table above rather than from anybody's judgement. Below it a document gets its score and the reason it gets nothing more: a low score is not evidence that a person wrote something, since a detector that detects nothing also returns a low score, and this project has deliberately never measured how much machine writing it catches. The boundary moves when this page moves — including upward if a larger corpus turns out to be less flattering.
82
83
83
84
Above it there is **one** verdict, not a scale of them. This corpus can place a boundary and can say nothing whatever about how much further past it a score has travelled: no text known to be human came close to the upper reaches, and grading "moderate" against "strong" would need machine-written text, which the opening of this page argues against collecting. Interfaces do shade a high score more urgently than a low one, and those shades are a display convention — they are not on this page because nothing measured them.
84
85
@@ -90,39 +91,39 @@ Every rule below fired on text no machine wrote, so each hit is a false positive
90
91
91
92
| Rule | Texts it fired on | Share | Total hits |
92
93
|---|---|---|---|
93
-
|`stat.burstiness`|25|27.8% |25|
94
-
|`rhet.in-terms-of`|9|10% |22|
95
-
|`rhet.not-only-but`|8|8.9% |16|
96
-
|`rhet.in-order-to`|7|7.8% |33|
97
-
|`lex.furthermore`|7|7.8% |24|
98
-
|`lex.robust`|7|7.8% |19|
99
-
|`lex.just`|7|7.8% |16|
100
-
|`lex.simply`|7|7.8% |16|
101
-
|`lex.utilizar`|7|7.8% |15|
102
-
|`rhet.in-this-article`|7|7.8% |15|
103
-
|`lex.notably`|7| 7.8% |14|
104
-
|`syn.superficial-ing`|7|7.8% |12|
105
-
|`rhet.with-regard-to`|7|7.8% |10|
106
-
|`lex.crucial`|7|7.8% |8|
107
-
|`syn.serves-as`|7|7.8% |8|
108
-
|`lex.moreover`|6|6.7% |35|
109
-
|`lex.comprehensive`|6|6.7% |27|
110
-
|`lex.utilize`|6|6.7% |22|
111
-
|`lex.facilitate`|6|6.7% |16|
112
-
|`rhet.rule-of-three`|6|6.7% |16|
113
-
|`lex.importantly`|6|6.7% |9|
114
-
|`rhet.weasel-attribution`|6|6.7% |9|
115
-
|`lex.actually`|6|6.7% |8|
116
-
|`rhet.in-conclusion`|6|6.7% | 8 |
117
-
|`rhet.important-note`|6|6.7% |7|
94
+
|`lex.just`|103|34.8% |189|
95
+
|`stat.burstiness`|95|32.1% |95|
96
+
|`rhet.in-conclusion`|89|30.1% |99|
97
+
|`rhet.rule-of-three`|77|26% |143|
98
+
|`lex.moreover`|66|22.3% |126|
99
+
|`rhet.not-only-but`|63|21.3% |90|
100
+
|`lex.furthermore`|55|18.6% |92|
101
+
|`lex.crucial`|50|16.9% |61|
102
+
|`lex.actually`|49|16.6% |69|
103
+
|`rhet.in-order-to`|41|13.9% |87|
104
+
|`rhet.in-terms-of`|22| 7.4% |38|
105
+
|`rhet.weasel-attribution`|19|6.4% |32|
106
+
|`rhet.in-this-article`|16|5.4% |24|
107
+
|`lex.simply`|14|4.7% |23|
108
+
|`lex.facilitate`|12|4.1% |22|
109
+
|`lex.utilize`|10|3.4% |30|
110
+
|`lex.truly`|10|3.4% |15|
111
+
|`lex.comprehensive`|9|3% |30|
112
+
|`rhet.when-it-comes`|9|3% |10|
113
+
|`syn.serves-as`|9|3% |10|
114
+
|`lex.notably`|8|2.7% |15|
115
+
|`rhet.with-regard-to`|8|2.7% |11|
116
+
|`rhet.on-one-hand`|8|2.7% |9|
117
+
|`lex.profound`|8|2.7% | 8 |
118
+
|`lex.robust`|7|2.4% |19|
118
119
119
120
A rule near the top is not automatically wrong. Some tells genuinely appear in human academic prose and the catalog says so. But a rule firing on most human texts is measuring the genre rather than the machine, and should be reweighted or retired.
120
121
121
122
## What this does not tell you
122
123
123
124
-**Nothing about how much AI writing it catches.** That is the other half of the picture and it is not measured here, deliberately. A tool that flags nothing has a perfect false-positive rate.
124
-
-**Nothing about text unlike this corpus.** These are published articles. A first-year essay is shorter, looser and differently edited, and the rate on one does not transfer to the other. Calibrating on your own students' pre-2022 work is the fix, and the same tool does it.
125
-
-**The grouping of writers is a proxy, not a fact.** Nobody's first language is recorded in a DOI. The manifest states the reasoning per text so it can be argued with; where it is wrong, the number moves.
125
+
-**Nothing about text unlike this corpus.** These are published articles and the essays of adult learners in a university English programme. A first-year essay by a native speaker is a different population again, and the rate on one does not transfer to the other. Calibrating on your own students' pre-2022 work is the fix, and the same tool does it.
126
+
-**The affiliation groups are a proxy, not a fact.** Nobody's first language is recorded in a DOI. The learner group is the exception — its corpus records each writer's first language — which is why it exists. Elsewhere the manifest states the reasoning per text so it can be argued with; where it is wrong, the number moves.
126
127
-**Hashes prove what *this* run measured**, not that another person extracting the same articles would get identical text. They would not: PDF and HTML extraction differ. Reproducing this needs the extracted texts, not just the manifest.
dotnet run --project tools/SignsOfAI.Calibration -- fetch --source wikipedia --lang en --count 25
66
67
dotnet run --project tools/SignsOfAI.Calibration -- fetch --source wikipedia --lang es --count 25
67
68
69
+
# 206 essays by adult learners of English, 2006–2012, one per student — a fixed rule, no --count,
70
+
# so anyone running it gets the same texts and the same hashes (about 180 MB downloaded once)
71
+
dotnet run --project tools/SignsOfAI.Calibration -- fetch --source pelic
72
+
68
73
# measure, and rewrite Docs/CALIBRATION.md
69
74
dotnet run --project tools/SignsOfAI.Calibration -- run
70
75
```
@@ -78,7 +83,8 @@ having.
78
83
| Group | What it is | What it is for |
79
84
|---|---|---|
80
85
|`en-anglophone-affiliation`| PLOS articles with at least one author affiliated in an anglophone country | The comparison baseline |
81
-
|`en-other-affiliation`| PLOS articles with no anglophone affiliation |**Standing in for second-language English** — the population this whole category harms |
86
+
|`en-other-affiliation`| PLOS articles with no anglophone affiliation | Was standing in for second-language English until the learner group arrived; kept, because it is the same question asked of professional writers |
87
+
|`en-second-language-learner`| Classroom essays from PELIC — University of Pittsburgh's Intensive English Program, 2006–2012, first language recorded per writer |**The population this whole category harms**, measured directly rather than through a proxy |
82
88
|`en-wikipedia`| Pre-2022 English Wikipedia revisions | Same register as the Spanish group, so a language effect can be told from a register effect |
83
89
|`es-wikipedia`| Pre-2022 Spanish Wikipedia revisions | The half nobody else measures at all |
84
90
@@ -88,11 +94,28 @@ direction: a paper with *any* anglophone affiliation counts as anglophone, which
88
94
second-language group and makes any gap found an understatement rather than an exaggeration. Every
89
95
entry records the affiliation used, so each classification can be argued with individually.
90
96
91
-
**None of this is a student essay.** Published articles are longer, more heavily edited and written by
92
-
people who write for a living. A first-year essay is a different thing and the rate measured on one
93
-
does not transfer to the other. The honest fix is for a school to calibrate on its own students'
94
-
pre-2022 work — the same tool does it, and the result would be a false-positive rate for *its*
95
-
population instead of somebody else's.
97
+
**The learner group is the one that is not a proxy.**[PELIC](https://github.com/ELI-Data-Mining-Group/PELIC-dataset)
98
+
records each writer's first language — Arabic, Korean, Chinese, Japanese, Spanish, Thai and Turkish
99
+
make up most of it — and every essay was written years before generative models, in a classroom,
100
+
under a prompt. The selection is a rule rather than a choice: first submitted version, writing
101
+
classes only, at least 662 words so the group enters at the floor the corpus already had, one text
102
+
per student so nobody prolific counts twice. Nothing is picked by score. The licence is
103
+
CC BY-NC-ND 4.0, which permits measuring and publishing the numbers and forbids redistributing the
104
+
texts — the same arrangement every other source here already has.
105
+
106
+
What it found is on `Docs/CALIBRATION.md` and it is the reason the boundary moved from 25 to 30: at
107
+
25, 9 of the 206 essays were flagged and none of the 90 published texts. The rules doing it are the
108
+
connectors an academic-English course teaches (`rhet.in-conclusion` fires in 40% of learner essays
109
+
and 14% of published ones) — see issue #75. It also caught a defect: `chat.eager-opener` had been
110
+
admitted on zero hits in the published texts and fired on eleven learner essays, because its pattern
111
+
accepted *"Of course, …"* with a comma, an ordinary concession, alongside *"Certainly!"*. Zero on one
112
+
register is not zero.
113
+
114
+
**Still no native-speaker student essay.** Published articles are written by people who write for a
115
+
living; the learner essays are by adults in a university language programme. A first-year essay by a
116
+
native speaker is a different population again, and the rate measured here does not transfer to it.
117
+
The honest fix is for a school to calibrate on its own students' pre-2022 work — the same tool does
118
+
it, and the result would be a false-positive rate for *its* population instead of somebody else's.
96
119
97
120
## Contributing texts
98
121
@@ -109,8 +132,9 @@ What is most wanted, in order:
109
132
-**Spanish academic writing.** SciELO and Redalyc are the obvious sources and neither was reachable
110
133
from where this was first assembled. Spanish is the half of this project nobody else measures, and
111
134
it currently rests on encyclopedia prose alone.
112
-
-**Anything closer to a student essay.** Coursework released under an open licence, pre-2022 writing
113
-
competition entries, open thesis repositories.
135
+
-**Native-speaker student writing.** Coursework released under an open licence, pre-2022 writing
136
+
competition entries, open thesis repositories. The learner group covers second-language writers;
137
+
nothing yet covers a first-year student writing in their own language.
114
138
-**More of everything.** With nothing flagged it still takes roughly seventy-five texts in a group
115
139
before the interval alone can bound a 5% rate. Most groups here are half that.
0 commit comments