Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
104 changes: 104 additions & 0 deletions Docs/Teaching/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
# For teachers — the part that is not software

*[Español abajo.](#para-docentes--la-parte-que-no-es-software)*

Six pieces of software have been built for this project. None of them is what a teacher needs first.

What a teacher needs first is language for a syllabus, a page to hand students before anything goes
wrong, and a procedure for the day a question becomes formal. None of that requires this tool, or
any tool. It is here because the hardest part of AI writing in a classroom was never detection — it
is what you do on the morning you suspect something, and there is nobody to ask.

**Use these however you like. Copy them, edit them, put your institution's name on them, translate
them. No licence, no attribution, no permission.** If they end up useful, we would like to hear how
you changed them — an issue or a pull request improves them for the next person.

| | | |
|---|---|---|
| **[Syllabus language](syllabus.en.md)** | [ES](syllabus.es.md) | Three policies at different strengths, ready to paste. Pick one — mixing them is how these fail. |
| **[The student sheet](student-sheet.en.md)** | [ES](student-sheet.es.md) | One page, written for students: what a flag means, what it does not, and what protects them. |
| **[Committee procedure](committee.en.md)** | [ES](committee.es.md) | For when it becomes formal, assuming the hardest case — the student denies it. |

## The three things these documents agree on

**A score is never the reason for a decision about a student.** Not at any threshold, not with any
tool, not ours. A detector tells you where to read; it does not tell you what happened. Every
document here is built so that a decision made by following it would survive the tool being removed
from it entirely.

**A conversation about the work settles what no software can.** Someone who wrote a text can say why
it is that text — why this example, what was left out, what a phrase of theirs means. Someone who did
not can restate it and cannot say why. That asymmetry is free, it takes ten minutes, and it is
stronger than anything in this repository.

**The burden does not go on the student.** Asking somebody to prove they wrote something reverses it
onto the person with the least power in the room, and it lands hardest on students writing in a
second language — the group every published study finds these tools flag most often. Ask them to talk
about their work instead.

## Why this project is entitled to say any of it

Because it publishes how often it is wrong, which almost nothing in this category does:
[`Docs/CALIBRATION.md`](../CALIBRATION.md). Ninety texts published before generative models existed,
zero flagged at the recommended threshold, and the honest reading is the interval rather than the
zero — under 4.1%.

Read further down that page and it says something less flattering, which is the part that matters
here: **no threshold is supported for English or Spanish on its own.** The corpus is too thin per
language. A figure measured mostly in one language, quoted at a student writing in another, is the
exact error these documents tell a committee not to make.

Everything runs offline. No document, or any part of one, leaves the machine it is on.

---

# Para docentes — la parte que no es software

Este proyecto ha construido seis piezas de software. Ninguna es lo que un docente necesita primero.

Lo que un docente necesita primero es un texto para el programa de la asignatura, una página que
entregar a los estudiantes antes de que pase nada, y un procedimiento para el día en que una duda se
vuelve formal. Nada de eso requiere esta herramienta ni ninguna otra. Está aquí porque lo difícil de
la escritura con IA en un aula nunca fue detectarla: es qué hacer la mañana en que uno sospecha algo
y no tiene a quién preguntar.

**Úselos como quiera. Cópielos, edítelos, póngales el nombre de su institución, tradúzcalos. Sin
licencia, sin atribución, sin permiso.** Si le resultan útiles, nos gustaría saber cómo los cambió —
una incidencia o un *pull request* los mejora para la próxima persona.

| | |
|---|---|
| **[Texto para el programa](syllabus.es.md)** | Tres políticas de distinta intensidad, listas para pegar. Elija una: mezclarlas es como fracasan. |
| **[La hoja del estudiante](student-sheet.es.md)** | Una página escrita para el estudiante: qué significa que lo marquen, qué no, y qué lo protege. |
| **[Procedimiento del comité](committee.es.md)** | Para cuando se vuelve formal, asumiendo el caso más difícil: el estudiante lo niega. |

## Las tres cosas en que estos documentos coinciden

**Una puntuación nunca es el motivo de una decisión sobre un estudiante.** Con ningún umbral, con
ninguna herramienta, tampoco la nuestra. Un detector le dice dónde leer; no le dice qué pasó. Todos
estos documentos están construidos para que una decisión tomada siguiéndolos sobreviva a eliminar la
herramienta por completo.

**Una conversación sobre el trabajo resuelve lo que ningún programa puede.** Quien escribió un texto
puede decir por qué es ese texto: por qué ese ejemplo, qué dejó fuera, qué significa una expresión
suya. Quien no lo escribió puede reformularlo y no puede decir por qué. Esa asimetría es gratis,
lleva diez minutos, y es más fuerte que cualquier cosa de este repositorio.

**La carga no recae sobre el estudiante.** Pedirle a alguien que demuestre que escribió algo invierte
la carga sobre la persona con menos poder de la sala, y cae con más fuerza sobre quien escribe en una
segunda lengua — el grupo que todos los estudios publicados encuentran marcado con más frecuencia por
estas herramientas. Pídale que hable de su trabajo.

## Por qué este proyecto puede permitirse decir todo esto

Porque publica con qué frecuencia se equivoca, cosa que casi nadie hace en esta categoría:
[`Docs/CALIBRATION.md`](../CALIBRATION.md). Noventa textos publicados antes de que existieran los
modelos generativos, ninguno marcado en el umbral recomendado, y la lectura honesta es el intervalo y
no el cero: por debajo del 4,1%.

Siga leyendo esa página y dice algo menos halagador, que es justo lo que importa aquí: **ni el
español ni el inglés respaldan por sí solos un umbral propio.** El corpus es demasiado delgado por
idioma. Una cifra medida mayoritariamente en un idioma, citada frente a un estudiante que escribe en
otro, es exactamente el error que estos documentos le piden a un comité que no cometa.

Todo funciona sin conexión. Ningún documento, ni fragmento de él, sale del equipo donde está.
129 changes: 129 additions & 0 deletions Docs/Teaching/committee.en.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
# A procedure for when the question becomes formal

For an academic-integrity committee, a programme director, or a teacher who has to write something
down. It assumes the hardest case: the student denies it.

**The one rule the rest depends on:** a detector's output is never evidence of authorship, and no
part of this procedure treats it as such. What follows is how to reach a defensible decision anyway.

---

## Sort what you have into three kinds

Most flawed integrity cases collapse because these were mixed together in one paragraph.

**1. Checkable facts.** Statements about the document that are true or false and that anyone can
verify independently — this citation appears in the text and nowhere in the reference list; these two
references carry the same DOI, so at most one of them is right; this reference is dated next year;
this word contains a Cyrillic letter shaped like a Latin one.

These do not depend on trusting any tool. **Anyone can check them by hand**, and a committee should:
open the bibliography, look. They also carry no opinion about authorship — a real source can be
produced in seconds, and that request is the whole test.

**2. A pattern score.** An opinion about prose. Useful for deciding where to read carefully. **It is
not evidence and it belongs in no finding.** If your written decision would be weaker with this
paragraph removed, the decision is not sound.

**3. What the student says about their own work.** In practice this decides the case, and the rest of
this document is about getting it right.

## Before the meeting

**Verify the checkable facts yourself.** Look up the DOIs. Search the reference list. If a claimed
contradiction does not hold when you check it by hand, drop it — one wrong item in a written finding
is enough to void the whole document at appeal, and rightly.

**Write down what specifically raised the question**, in one or two sentences a student could
understand. If you cannot, there is no case yet. *"The score was high"* is not a reason. *"Four of
the six sources cited do not appear in the bibliography, and two share a DOI"* is.

**Check the tool's error rate for the language the work is written in.** This matters more than it
sounds. This tool's corpus supports **no threshold for Spanish or English separately** — its best
bounds are 5.6% and 13.3% by language, against an aggregate of 4.1%. Quoting an aggregate measured
mostly on one language, at a student writing in another, is the exact error the tool's own
documentation warns against. If you cite a figure in a finding, cite the one for that language, and
if there is none, say there is none.

**Decide who is in the room.** For the student: someone they choose. The power difference is the
main source of unfair outcomes in these meetings, and it is cheap to reduce.

## The meeting

**Open by saying what it is.** *"This is a conversation about your assignment, not a decision. Nothing
has been decided."* Say it even if you think it is obvious. It is not obvious from the other chair.

**Ask about the work, not about the accusation.** Take two or three specific passages and ask:

> *You wrote this phrase here — what does it mean, in your own words?*
>
> *Why this example and not another one?*
>
> *What did you leave out of this section, and why?*
>
> *Where did this source come from? What does it actually argue?*

Someone who wrote a text can talk about the choices in it — not perfectly, not fluently, but they can
say why. Someone who did not write it produces summary rather than intention: they can restate the
paragraph and cannot say why it is that paragraph. **This asymmetry is the strongest instrument in the
room, and it is free.**

**Ask for the sources.** Not the citations, the sources. A real one arrives in seconds. An invented
one cannot arrive at all, and the failure to produce it is a fact you may write down.

**Ask for the drafts, but weigh their absence carefully.** Drafts are strong evidence when present
and weak evidence when missing. Plenty of honest students write in one pass, in one file, and delete
nothing because there was nothing to delete. Absence of drafts is not evidence of anything, and a
policy that treats it as such punishes people for how they work.

**Take the answers seriously when they explain things.** *"I write formally because that is how I was
taught"* is a complete explanation for a high score, and it is a common one among students writing in
a second language. A committee that has decided in advance what the meeting will conclude is not
holding a meeting.

## Writing the decision

The document should be able to stand with the detector removed from it entirely. Test that literally:
delete every sentence about the tool and read what is left. If a decision remains, write it. If
nothing remains, there is no finding.

**State the facts you verified yourself** and how you verified them. **State what the student said.**
**State what you concluded and why.** If you mention the tool at all, say what it is and what it is
not, in the finding itself:

> A writing-analysis tool was used to decide which sections to examine. Its output is not evidence of
> authorship and was not treated as such. It publishes a false-positive rate on writing known to be
> human of under 4.1% overall, with no threshold supported for either language on its own — the
> findings below rest on the source verification and the interview, not on that tool.

A committee that writes this sentence is in a stronger position than one that omits it, because the
sentence is going to be raised at appeal whether or not it appears.

## Proportion

Most of what these processes catch is not fabrication. It is a student who was overwhelmed, used a
tool for a paragraph, and did not disclose it because nobody told them how. A first response of
*rewrite it and tell me what you used* keeps a student in the room and produces a better writer. The
severe outcomes should be reserved for what deserves them: invented sources, work bought or taken
from another person, and denial maintained against facts the student cannot explain.

---

## What this project will not give you

**A confidence percentage.** It does not have one, and neither does anyone else who prints one.

**A verdict below its supported threshold.** The report prints the number and no interpretation.
That is deliberate: a page that says *reads mostly human* above *treat this score as saying nothing*
lets the reader keep whichever half they came in wanting.

**A false-positive rate for your student population.** The published figure was measured on articles
published before generative models existed — not on first-year coursework, not on your institution,
not on your language mix. It is a floor, not a promise. Anyone quoting it as though it applied
directly to a nineteen-year-old's essay is overstating it, including us.

---

Related: [`syllabus.en.md`](syllabus.en.md) — policy language before any of this is needed.
[`student-sheet.en.md`](student-sheet.en.md) — what students should know beforehand.
[`../CALIBRATION.md`](../CALIBRATION.md) — the error rate and its method.
Loading
Loading