Skip to content

A corpus of formal prose cannot admit a rule aimed at student essays #66

Description

@peopleworks

Six candidate rules were left out of #65 for a reason that will keep coming up, so it belongs in an
issue rather than a commit message.

The six

Candidate Why it looks like a tell
this essay will explore The assignment restated as an opening line
let's dive into Blog runway
here's what you need to know Listicle framing
without further ado Same
here's the kicker / the best part is False suspense
the possibilities are endless / only time will tell / the future looks bright An ending that says nothing

All of them scored zero against the calibration corpus, which is the bar #65's rules cleared.
They were still left out.

Why zero was not enough

The corpus is 65 open-access research articles and 25 encyclopedia revisions. It can prove that a
phrase is absent from formal published prose. It cannot say anything about how often a human
sixteen-year-old writes "This essay will explore the causes of the war" — which is a sentence that
is actively taught, or about how often a human blogger writes "the possibilities are endless".

Admitting them would mean measuring the false-positive budget on one population and spending it on
another. That is the same shape as #59, where a threshold measured on documents of median 3,241
words is applied today to a pasted 400-word paragraph with nothing warning anybody.

The chat.* rules in #65 escape this because they are not a register at all — no corpus of human
writing in any genre contains "as of my last training update".

What would settle it

A short-form human corpus, pre-2022, of the two registers we score and cannot see:

  • student coursework — the hard one: licensing, consent, and it is the population that matters most;
  • blog and journalistic prose — much easier, and fetch already filters by length.

The second is worth doing on its own even before the first, because it also gives #59 the thing
it actually needs: short texts somebody composed at that length, not long documents truncated to
it. The calibration README already warns that these are not the same population, and one corpus
answers both questions.

Until it exists, these six stay out and this issue is the record of why — so nobody adds them later
on the grounds that they obviously look like AI, which they do.

Related

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions