← Module hub · Unit 9 · Self-check

Unit 9 activity: Alles Lüge! First- and second-order reasoning

DEVELOPMENT REVIEW DEPLOYMENT - NOT READY FOR RELEASE

Time: about 12 minutes. Checkpoint: cp09-sort-the-claims. You will need: the student page for this unit, and your evidence pack from Unit 8.

This is the module's closing activity. Two tasks and a checkpoint. Task 1 takes about five minutes. Task 2 takes about six and is the one that transfers furthest, so protect it if you are short of time.

No prior logic or cybernetics is assumed. Use the unit reading's definitions of observer, context, record, formal system, theorem, and truth predicate; the task tests distinctions made there rather than background knowledge.


Task 1 — sort claims by what supports them

Below are six statements. Choose three to analyse, and include C and D. The others are available for discussion or extension. Every statement either appears in this unit or is a short step from something that does.

Put each statement into one of three buckets.

bucket test
proved theorem A published result with a date and a standard citation. Name it.
established research programme Published, influential, decades old — and still not a theorem. Name the tradition or the researchers.
overclaim It does not follow. Name what it contradicts, or name the step that is missing.

For each of your three, write one sentence saying why, and one line saying what would change your placement: for a theorem, what would unsettle it; for a research programme, what would settle it either way; for an overclaim, which specific piece of your own work or which missing step refutes it.

One further instruction, and it is not a hint. If a statement does not fit any of the three buckets, say so, and name the instrument that does fit. A historical claim, for instance, is sorted as established, contested or unsourced — the scheme the optional Unit 7 works through, and one you can apply here without having done it. Being handed the wrong scheme and using it anyway is a real failure mode, and refusing it is a correct answer.

Do not look the statements up before sorting them. Sorting a claim you have not resolved is the actual skill, because in practice most claims are used, or not used, in whatever state you found them.

The statements

A. Any consistent formal system strong enough to express arithmetic contains true statements that it cannot prove.

B. Mind is better described as a property of a system of relationships than as something belonging to an isolated individual.

C. The Macy Conferences on Cybernetics were ten meetings held between 1946 and 1953, funded by the Josiah Macy Jr. Foundation and chaired by Warren McCulloch.

D. Gödel's incompleteness theorem shows that a language model cannot verify its own output.

E. Because an observing system is part of what it observes, a report an AI system produces about its own behaviour carries no information at all.

F. Gödel's and Tarski's results concern formal systems of a specific kind, so carrying them across to people or institutions requires a separate argument.

Write it up like this

Statement:        <letter>
Bucket:           proved theorem | established research programme | overclaim | none of these
Citation / whose / what it contradicts:   <one line>
Why:              <one sentence>
What would change it:   <one line>

Three of these are harder than the rest, and they are hard in three different directions. If nothing gave you trouble, you have probably sorted by how the sentences sound rather than by what supports them, which is the failure mode this whole module is built to prevent.


Task 2 — now sort a claim of your own

Open the evidence pack you assembled in Unit 8. You are going to turn the same method on your own work, which is the only version of it that costs you anything.

Step 1 — pick the claim

Take item 8, your two-sentence correctness defence. If your pack has several claims, take the one you would most mind being wrong about.

Copy it out as it stands. Do not improve it first.

Step 2 — name its index

A claim made by an observer inside the system carries an index. For your own sentence, write down which of these it states, and which it merely assumes.

index for your defence, this means
observer who ran the check, and who is accountable for the result
context which convention or specification the result is correct under
record what was written down, and where another person can find it now

Most defences state one or two of these and assume the rest. That is normal, and it is worth knowing precisely which one you left implicit, because that is the one that will be read differently by someone who is not you.

Step 3 — sort the clauses

Do not force your own result into Task 1's buckets. A test result is neither a published theorem nor a research programme, and reusing the wrong instrument here is the same error Task 1 asks you to spot. Use the instrument that fits this kind of claim:

A defence may contain all three. Mark the clauses rather than assigning the whole paragraph one flattering label. Then write one line: what evidence would strengthen the claim, or what qualification would make its scope accurate?

Step 4 — one honest sentence

Finish with this, in your pack:

The part of my defence that rests on a record is ..., and the part that rests on my judgement is ...

If those two turn out to be the same thing, you have found something worth knowing.


Task 3 — the checkpoint

Do this after Tasks 1 and 2. The checkpoint asks you to sort claims made by this unit, and it is a different exercise once you have sorted three statements and one of your own.

From course/lab/:

python3 selfcheck.py run cp09-sort-the-claims     # or use the browser self-check

Both routes ask the same question, accept the same answers, and give the same diagnostics. Neither is a fallback for the other. Pick whichever you have to hand.


Optional extension — where the question came from

Optional. Outside the core time budget. It does not gate anything.

Spend five minutes finding out one thing about the Macy Conferences that this unit did not tell you: a participant it does not name, a topic that recurred across the ten meetings, or a disagreement that was never resolved. Write two or three sentences on what you found and where you found it.

Then answer one question in one line: is the source you used a primary record of the meetings, or someone's later account of them? Both are useful and they are not the same thing, and saying which one you have is the habit this module has been building since Unit 4.


Before you finish

You should be leaving this unit with:

Add the whole thing to your evidence pack as an appendix. It is the shortest demonstration you will produce that you can tell a theorem from a research programme from a sentence that merely sounds settled.

That is the module. Wider course Items 3, 5, and 7 go further than this material does; course/course-map.md shows where.

Marking guidance — open this once you have done the activity, to check your own work

For learners checking their own work after attempting the activity, and for instructors marking or discussing it.

This key does not restate the answer to cp09-sort-the-claims, and it does not paraphrase it either. Sorting claims by what supports them is the learning objective, and the checkpoint gives targeted diagnostics if you go astray. Checkpoint answers live in course/lab/checkpoints.py and nowhere else, by design. The six statements discussed below are deliberately disjoint from the three the checkpoint uses: no statement here is the checkpoint's, and no statement here is a restatement of one of the checkpoint's.

Scope. This unit teaches established, published material only. There is no research framework in it to mark, and answers that reach for one have gone outside the unit.

Citation shorthand. Bracket labels resolve to course/references.md: [C1] von Foerster, [C2] Mead 1968, [C3] Maturana and Varela, [C4] the Macy Conferences, [C5] Bateson, [C6] Gödel 1931, [C7] Tarski 1936; [G1] to [G4] support the register; [Q1] and [Q2] are the songs. Fuller treatment of this unit's argument: course/second-order.md.

The assessable part is the reasoning, not the label. A learner who places a statement differently but names a real citation, a real tradition, or a real missing step has done the exercise. A learner who lands on the intended bucket and writes "it sounds like a theorem" has not.


Task 1 — the six statements

A. Incompleteness

Proved theorem. Gödel 1931 [C6]. Any consistent formal system strong enough to express arithmetic contains true statements it cannot prove.

The tell is that it has a date, a name, and a place in every logic textbook. A learner who marks it "established research programme" because it is unfamiliar has confused unfamiliar with unsettled, and that confusion is worth naming out loud — it is the mirror image of the error waiting in statement B.

What would unsettle it: an error in the proof, or a change to the hypotheses. Note that the statement is conditional — consistent, sufficiently strong — and almost every misuse of it in the wild starts by dropping one of those conditions.

B. Mind as a property of a system of relationships

Established research programme. Bateson [C5], and the wider tradition [C1], [C3]. Published, developed over decades, widely built on, and still an account rather than a theorem.

The tell is the phrase "is better described as". The statement is not reporting a result; it is proposing that a particular way of describing something is the right one. Descriptions of that shape are argued for, adopted, extended, and sometimes displaced. They are not looked up and settled.

Say plainly in discussion that this category is not a polite way of saying "wrong" or "weak". It is a statement about what kind of support the claim has.

What would settle it either way: sustained independent development that makes the account do work nobody could do without it, or a demonstration that it fails on the cases it was built for. Note that "better described as" invites the question better for what purpose, and a learner who asks it has found the seam.

C. The Macy Conferences

None of the three buckets fits, and that is the point. This is a historical claim, and historical claims are sorted as established, contested, or unsourced — the instrument the optional Unit 7 works through, and named in the activity itself so a learner who skipped Unit 7 can still reach for it.

Sorted that way it is establishedcourse/references.md [C4] records the series, the dates, the funder and the chair.

The best answers do two things: notice that the bucket scheme does not fit, and reach for the other instrument instead. A learner who does only the first has still done something worth marking, because recognising that you are holding the wrong instrument is most of the skill. Do not treat having done Unit 7 as a requirement here.

Watch for a learner who marks it "established research programme" because cybernetics is a research programme. The statement is not about the programme; it is about ten meetings, their dates, their funder and their chair. Those are facts with a record, not a way of framing problems.

D. "Gödel shows a language model cannot verify its own output"

Overclaim, and the most interesting one on the list, because both halves of it are true and the join is not.

Gödel 1931 [C6] is a proved theorem. It is a result about formal systems of a specific kind. A language model is not that object. Carrying the result across is an argument, requiring premises about what corresponds to what, and no such argument appears in the sentence.

The failure here is not falsity. The conclusion might even be defensible on other grounds. The failure is a theorem being used as a licence for a claim about a different kind of system, with the transfer step unstated. That move is common, it reads extremely well, and it is precisely the thing this module has been training you to catch.

Accept a learner who says "the theorem is proved, the transfer is not, and the sentence hides the seam" over one who writes "overclaim" with nothing behind it.

E. "An AI system's report about itself carries no information at all"

Overclaim. This is the second overclaim from the student page wearing practical clothing, pushed one step further.

Two separate faults. First, it does not follow: that an observer is inside the system says nothing about whether its reports correlate with anything. Second, it is operationally empty — if a self-report carries nothing, you have no way to decide which outputs deserve a check, which in practice means you check nothing, which is exactly where blanket trust also lands you.

What refutes it from your own work: Unit 5. Generated output that was correct, generated output that was wrong, and a check that told you which was which. You were not applying a presumption; you were comparing against a record.

Note for discussion: describe behaviour, not intent. "The system reported X" is the right form. "The system believed X" is not, and a learner who slips into the second has drifted from the module's language rules.

F. Scope of Gödel and Tarski

Proved theorem, in the sense that matters here: this is a statement about what [C6] and [C7] actually say, and you can check it by reading the theorems' hypotheses. Accept "established" with that reasoning.

The generous alternative reading is that F is a methodological remark rather than a result. A learner who says so, and then says the remark is checkable against the theorem statements, has understood it better than the bucket alone shows.

What must not happen is a learner marking F an overclaim. F is the caution, not the leap. If someone has read the caution as itself an overreach, they have inverted the section they most needed.


Task 2 — a claim of your own

Step 2, the index

Typical results, and none of them is a failure:

A good answer marks at least one of the three as assumed. A learner who marks all three as stated has almost certainly read their own sentence generously, which is the one reading a sceptical reader will not give it.

Step 3, the clauses

A well-built pack usually contains a record-supported result: there is a record, the reader can consult or reproduce it, and the sentence states the range, convention, and properties actually checked. Do not call this a "proved theorem" or an "established research programme". Reusing Task 1's instrument on this kind of claim is the same error statement C is there to provoke.

The instructive answer is when a learner finds an overclaim in their own writing. The common form is a scope slip: a check that ran over a specific range, under a specific convention, reported as though it established correctness in general. That is not dishonesty and it does not need punishing. It is the same structure as the Note G table — a narrow supported claim and a broad readable one, and the broad one is easier to write.

Reasoned judgement is usually present beside the record: which discrepancy mattered, which convention was the sensible default, and what counted as resolved. It should be attributed to the learner by name, not dressed up as something more.

Step 4, the honest sentence

The intended discovery is that the record-backed part of a defence is usually smaller than it felt while writing it, and that the remainder is not worthless — it is judgement, which is exactly the thing the module says must be attributable to a person. A learner who concludes "so my defence is weak" has misread the exercise. The point is to be able to say which part is which.


Common ways this activity goes wrong

Sorting by how the sentence sounds. Statements B, D and F are all written in the register of established results. That is not a trick; it is what most writing about this material looks like, including writing produced by a language model. The sorting has to be done on provenance, not on prose.

Treating "established research programme" as a downgrade. It is not. A programme can be decades old, published, influential and worth taking seriously while still not being a theorem. The unit says so explicitly and the checkpoint tests it.

Reading the unit as a case against AI systems. If a learner leaves saying "so it cannot be trusted", the second overclaim has not been closed properly. The materials establish no base rate for correct output. They teach learners to match the strength and independence of a check to the claim and its risk.

Collapsing the whole thing into "everything is relative". The index is the opposite of that. An indexed claim is more checkable than an unindexed one, because it says what would have to be compared with what. The Bernoulli convention is the proof: naming it turns an unresolvable disagreement into a resolved one.


For instructors

Your private activity record

Browser storage is not a permanent copy

Progress is kept only in this browser, profile and device. Private browsing, clearing site data, removing the profile, a browser reset, storage eviction or device loss can erase it. Keep important answers and contributions separately.

These notes stay in this browser unless you download a backup or activity log.