Time: about 12 minutes. Checkpoint: cp18-sort-the-claims.
You will need: the student page for this unit, and your evidence pack from
Unit 17.
This is the module's closing activity. Two tasks and a checkpoint. Task 1 takes about five minutes. Task 2 takes about six and is the one that transfers furthest, so protect it if you are short of time.
No prior logic or cybernetics is assumed. Use the unit reading's definitions of observer, context, record, formal system, theorem, and truth predicate; the task tests distinctions made there rather than background knowledge.
Task 1 — sort claims by what supports them
Below are six statements. Choose three to analyse, and include C and D. The others are available for discussion or extension. Every statement either appears in this unit or is a short step from something that does.
Put each statement into one of three buckets.
| bucket | test |
|---|---|
| proved theorem | A published result with a date and a standard citation. Name it. |
| established research programme | Published, influential, decades old — and still not a theorem. Name the tradition or the researchers. |
| overclaim | It does not follow. Name what it contradicts, or name the step that is missing. |
For each of your three, write one sentence saying why, and one line saying what would change your placement: for a theorem, what would unsettle it; for a research programme, what would settle it either way; for an overclaim, which specific piece of your own work or which missing step refutes it.
One further instruction, and it is not a hint. If a statement does not fit any of the three buckets, say so, and name the instrument that does fit. A historical claim, for instance, is sorted as established, contested or unsourced — the scheme Unit 16 works through, and one you can apply here whether or not you did that unit first. Being handed the wrong scheme and using it anyway is a real failure mode, and refusing it is a correct answer.
Do not look the statements up before sorting them. Sorting a claim you have not resolved is the actual skill, because in practice most claims are used, or not used, in whatever state you found them.
The statements
A. Any consistent formal system strong enough to express arithmetic contains true statements that it cannot prove.
B. Mind is better described as a property of a system of relationships than as something belonging to an isolated individual.
C. The Macy Conferences on Cybernetics were ten meetings held between 1946 and 1953, funded by the Josiah Macy Jr. Foundation and chaired by Warren McCulloch.
D. Gödel's incompleteness theorem shows that a language model cannot verify its own output.
E. Because an observing system is part of what it observes, a report an AI system produces about its own behaviour carries no information at all.
F. Gödel's and Tarski's results concern formal systems of a specific kind, so carrying them across to people or institutions requires a separate argument.
Write it up like this
Statement: <letter>
Bucket: proved theorem | established research programme | overclaim | none
Citation / whose / what it contradicts: <one line>
Why: <one sentence>
What would change it: <one line>
Three of these are harder than the rest, and they are hard in three different directions. If nothing gave you trouble, you have probably sorted by how the sentences sound rather than by what supports them, which is the failure mode this whole module is built to prevent.
Task 2 — now sort a claim of your own
Open the evidence pack you assembled in Unit 17. You are going to turn the same method on your own work, which is the only version of it that costs you anything.
Step 1 — pick the claim
Take item 8, your two-sentence correctness defence. If your pack has several claims, take the one you would most mind being wrong about.
Copy it out as it stands. Do not improve it first.
Step 2 — name its index
A claim made by an observer inside the system carries an index. For your own sentence, write down which of these it states, and which it merely assumes.
| index | for your defence, this means |
|---|---|
| observer | who ran the check, and who is accountable for the result |
| context | which convention or specification the result is correct under |
| record | what was written down, and where another person can find it now |
Most defences state one or two of these and assume the rest. That is normal, and it is worth knowing precisely which one you left implicit, because that is the one that will be read differently by someone who is not you.
Step 3 — sort the clauses
Do not force your own result into Task 1's buckets. A test result is neither a published theorem nor a research programme, and reusing the wrong instrument here is the same error Task 1 asks you to spot. Use the instrument that fits this kind of claim:
- Record-supported result — the clause says what check ran, under which specification and range, and where another person can inspect or repeat it.
- Reasoned judgement — the clause interprets the record, or makes a choice that the evidence does not settle. Name whose judgement it is.
- Overclaim — the clause claims more than the evidence supports. The usual form is a check that passed over one range being reported as correctness for all possible inputs.
A defence may contain all three. Mark the clauses rather than assigning the whole paragraph one flattering label. Then write one line: what evidence would strengthen the claim, or what qualification would make its scope accurate?
Step 4 — one honest sentence
Finish with this, in your pack:
The part of my defence that rests on a record is ..., and the part that rests on my judgement is ...
If those two turn out to be the same thing, you have found something worth knowing.
Task 3 — the checkpoint
Do this after Tasks 1 and 2. The checkpoint asks you to sort claims made by this unit, and it is a different exercise once you have sorted three statements and one of your own.
From course/lab/:
python3 selfcheck.py run cp18-sort-the-claims # or in the browser
Both routes ask the same question, accept the same answers, and give the same diagnostics. Neither is a fallback for the other. Pick whichever you have to hand.
Optional extension — where the question came from
Optional. Outside the core time budget. It does not gate anything.
Spend five minutes finding out one thing about the Macy Conferences that this unit did not tell you: a participant it does not name, a topic that recurred across the ten meetings, or a disagreement that was never resolved. Write two or three sentences on what you found and where you found it.
Then answer one question in one line: is the source you used a primary record of the meetings, or someone's later account of them? Both are useful and they are not the same thing, and saying which one you have is the habit this module has been building since Unit 7.
Before you finish
You should be leaving this unit with:
- three sorted statements, including C and D, each with a citation, an attribution, or a named missing step
- one defence of your own, split into record-supported result, judgement, and any overclaim, with its three index entries marked stated or assumed
- one line saying what evidence or qualification it needs
cp18-sort-the-claimspassed
Add the whole thing to your evidence pack as an appendix. It is the shortest demonstration you will produce that you can tell a theorem from a research programme from a sentence that merely sounds settled. That is the module.
Marking guidance — open this once you have done the activity, to check your own work
Read this after attempting the activity.
This key does not restate the answer to cp18-sort-the-claims, and it does not
paraphrase it either. Sorting claims by what supports them is the learning
objective, and the checkpoint gives targeted diagnostics if you go astray.
course/lab/checkpoints.py is the canonical machine-readable source of
checkpoint answers, and this key is not a second copy of one. Answers do appear
elsewhere in the materials — the guidance edition of the handbook carries the
checkpoint data and its explanations — so the point is that they are generated
from that one source rather than restated by hand here. The six statements discussed below are deliberately disjoint from the
three the checkpoint uses: no statement here is the checkpoint's, and no statement
here is a restatement of one of the checkpoint's.
Scope. This unit teaches established, published material only. There is no research framework in it to mark, and answers that reach for one have gone outside the unit.
Citation shorthand. Bracket labels resolve to course/references.md:
[C1] von Foerster, [C2] Mead 1968, [C3] Maturana and Varela, [C4] the Macy
Conferences, [C5] Bateson, [C6] Gödel 1931, [C7] Tarski 1936; [G1] to
[G4] support the register; [Q1], [Q2] and [Q3] are the three songs. Fuller treatment of
this unit's argument: course/second-order.md.
The part that counts is the reasoning, not the label. Placing a statement differently but naming a real citation, a real tradition or a real missing step is doing the exercise. Landing on the intended bucket and writing "it sounds like a theorem" is not.
Task 1 — the six statements
A. Incompleteness
Proved theorem. Gödel 1931 [C6]. Any consistent formal system strong enough to express arithmetic contains true statements it cannot prove.
The tell is that it has a date, a name, and a place in every logic textbook. Marking it "established research programme" because it is unfamiliar confuses unfamiliar with unsettled — the mirror image of the error waiting in statement B, and worth naming as such.
What would unsettle it: an error in the proof, or a change to the hypotheses. Note that the statement is conditional — consistent, sufficiently strong — and almost every misuse of it in the wild starts by dropping one of those conditions.
B. Mind as a property of a system of relationships
Established research programme. Bateson [C5], and the wider tradition [C1], [C3]. Published, developed over decades, widely built on, and still an account rather than a theorem.
The tell is the phrase "is better described as". The statement is not reporting a result; it is proposing that a particular way of describing something is the right one. Descriptions of that shape are argued for, adopted, extended, and sometimes displaced. They are not looked up and settled.
Say plainly in discussion that this category is not a polite way of saying "wrong" or "weak". It is a statement about what kind of support the claim has.
What would settle it either way: sustained independent development that makes the account do work nobody could do without it, or a demonstration that it fails on the cases it was built for. Note that "better described as" invites the question better for what purpose, and asking it is finding the seam.
C. The Macy Conferences
None of the three buckets fits, and that is the point. This is a historical claim, and historical claims are sorted as established, contested, or unsourced — the instrument Unit 16 works through, and named in the activity itself so that it is available whether or not you did that unit first.
Sorted that way it is established — course/references.md [C4] records the
series, the dates, the funder and the chair.
The best answers do two things: notice that the bucket scheme does not fit, and reach for the other instrument instead. Doing only the first still counts, because recognising that you are holding the wrong instrument is most of the skill.
The trap is marking it "established research programme" because cybernetics is a research programme. The statement is not about the programme; it is about ten meetings, their dates, their funder and their chair. Those are facts with a record, not a way of framing problems.
D. "Gödel shows a language model cannot verify its own output"
Overclaim, and the most interesting one on the list, because both halves of it are true and the join is not.
Gödel 1931 [C6] is a proved theorem. It is a result about formal systems of a specific kind. A language model is not that object. Carrying the result across is an argument, requiring premises about what corresponds to what, and no such argument appears in the sentence.
The failure here is not falsity. The conclusion might even be defensible on other grounds. The failure is a theorem being used as a licence for a claim about a different kind of system, with the transfer step unstated. That move is common, it reads extremely well, and it is precisely the thing this module has been training you to catch.
"The theorem is proved, the transfer is not, and the sentence hides the seam" is worth much more than "overclaim" with nothing behind it.
E. "An AI system's report about itself carries no information at all"
Overclaim. This is the second overclaim from the student page wearing practical clothing, pushed one step further.
Two separate faults. First, it does not follow: that an observer is inside the system says nothing about whether its reports correlate with anything. Second, it is operationally empty — if a self-report carries nothing, you have no way to decide which outputs deserve a check, which in practice means you check nothing, which is exactly where blanket trust also lands you.
What refutes it from your own work: Unit 8. Generated output that was correct, generated output that was wrong, and a check that told you which was which. You were not applying a presumption; you were comparing against a record.
Describe behaviour, not intent. "The system reported X" is the right form; "the system believed X" is not, and slipping into the second is drifting away from what you can actually evidence.
F. Scope of Gödel and Tarski
Proved theorem, in the sense that matters here: this is a statement about what [C6] and [C7] actually say, and you can check it by reading the theorems' hypotheses. Accept "established" with that reasoning.
The generous alternative reading is that F is a methodological remark rather than a result. Saying so, and then saying the remark is checkable against the theorem statements, understands it better than the bucket alone shows.
What must not happen is marking F an overclaim. F is the caution, not the leap. Reading the caution as itself an overreach inverts the section you most needed.
Task 2 — a claim of your own
Step 2, the index
Typical results, and none of them is a failure:
- Observer — almost always implicit. The pack has one author and it feels redundant to name them. It is not redundant to a reader who has three packs on their desk.
- Context — usually stated, because Unit 17's template forces it. This is the one the module worked hardest on, and it shows.
- Record — often half-stated: "the checker reported no divergence" names the check but not where the output now lives, or whether it can be produced again.
A good answer marks at least one of the three as assumed. Marking all three as stated almost certainly means you read your own sentence generously, which is the one reading a sceptical reader will not give it.
Step 3, the clauses
A well-built pack usually contains a record-supported result: there is a record, the reader can consult or reproduce it, and the sentence states the range, convention, and properties actually checked. Do not call this a "proved theorem" or an "established research programme". Reusing Task 1's instrument on this kind of claim is the same error statement C is there to provoke.
The instructive outcome is finding an overclaim in your own writing. The common form is a scope slip: a check that ran over a specific range, under a specific convention, reported as though it established correctness in general. That is not dishonesty and it does not need punishing. It is the same structure as the Note G table — a narrow supported claim and a broad readable one, and the broad one is easier to write.
Reasoned judgement is usually present beside the record: which discrepancy mattered, which convention was the sensible default, and what counted as resolved. It should be attributed to you by name, not dressed up as something more.
Step 4, the honest sentence
The intended discovery is that the record-backed part of a defence is usually smaller than it felt while writing it, and that the remainder is not worthless — it is judgement, which is exactly the thing the module says must be attributable to a person. Concluding "so my defence is weak" misreads the exercise. The point is to be able to say which part is which.
Common ways this activity goes wrong
Sorting by how the sentence sounds. Statements B, D and F are all written in the register of established results. That is not a trick; it is what most writing about this material looks like, including writing produced by a language model. The sorting has to be done on provenance, not on prose.
Treating "established research programme" as a downgrade. It is not. A programme can be decades old, published, influential and worth taking seriously while still not being a theorem. The unit says so explicitly and the checkpoint tests it.
Reading the unit as a case against AI systems. If you leave saying "so it cannot be trusted", the second overclaim has not been closed properly. Nothing here establishes a base rate for correct output. What it teaches is how to match the strength and independence of a check to the claim and its risk.
Collapsing the whole thing into "everything is relative". The index is the opposite of that. An indexed claim is more checkable than an unindexed one, because it says what would have to be compared with what. The Bernoulli convention is the proof: naming it turns an unresolvable disagreement into a resolved one.
Why this unit is written the way it is
- It comes last, and nothing before it depends on it. No checkpoint, activity or definition in Units 0, 2, 3, 6, 7, 8, 10, 16 and 17 needs this unit; those units carry a pointer to it and nothing more. You can work through the module without reaching here and lose none of the method. This unit is the part that asks what the method cannot reach — which is a different question, and one you are entitled to leave until you want it.
- It raises the issue; it does not settle it. The scope is fixed to established, published material. Where the argument could be extended — however interesting the extension — the unit stops, because past that line it would be doing the thing it warns about.
- Section 4's second half is the unit. "What they do not say" is the point, not a caveat attached to one. If you read only one part of this unit, read that.
- Statement D is the one to sit with. A real theorem, a real system, and an unstated transfer between them is the shape of a large fraction of confident writing about AI and formal limits.
- The atmosphere never replaces the citation. Every claim in this unit keeps its source or its uncertainty label, including the ones on the darkest pages. A unit about unwarranted certainty that asked for your trust on mood would be refuting itself.
- If the status of any claim in
course/second-order.mdchanges, that file is updated first and this unit follows. This unit holds no independent position that could drift out of step with it.
Your private activity record
Browser storage is not a permanent copy
Progress is kept only in this browser, profile and device. Private browsing, clearing site data, removing the profile, a browser reset, storage eviction or device loss can erase it. Keep important answers and contributions separately.
These notes stay in this browser unless you download a backup or activity log.