Video
Media Pending: Unit Video
Intended content: Full narrated video presentation, including visual assets, caption file, and transcript.
Learning purpose: Give the learner a way of sorting AI tasks that does not date, by asking
Planned form & duration: Video, ~6 minutes.
Accessible text alternative: Five classes, distinguished only by where the reference lives; The written material below covers the same complete learning path.
Watch what changes
Reading time: about 10 minutes. Activity: about 12 minutes.
The question worth asking
"What is generative AI good at?" is a question that goes stale. Answer it with a list of products and the list is wrong within a year. Answer it with a list of tasks and you have described what the systems could do on the day you wrote it.
This unit answers a different question, one whose answer does not move:
For this task, where would the evidence come from that the output is right?
That question has only a few possible answers, and they are the useful way to sort the work. Two tasks that look nothing alike — translating a paragraph and reformatting a table — sit in the same class, because in both cases you already hold the thing you would check against. Two tasks that look almost identical — "summarise this document" and "summarise the research on this topic" — sit in different classes, because in one case the document is in front of you and in the other you would have to go and find it.
The class decides what checking costs. That is the whole of this unit.
Five classes, by where the reference lives
A reference is whatever you would compare the output against to find out whether it is right. In Unit 6 you will build one by hand; in Unit 8 you will use it. Here we are only asking where it comes from.
| Class | What you supply | Where the reference lives | What checking costs |
|---|---|---|---|
| Transformation | The content itself | In your hands already — it is your input | Low: compare output against input |
| Grounded answering | A source, plus a question about it | In the source you supplied | Low to moderate: read the source at the cited point |
| Generation to a specification | A statement of what would count as correct | In the specification you wrote | Moderate: test the output against the specification |
| Recall | A question only | Somewhere external, and you must go and find it | High: the check is a separate research task |
| Judgement | A question with no single determinate answer | No single right answer to compare against, but the premises and the criteria are checkable | Check the factual premises; assess the proposal against stated criteria |
Read the last column downwards. It is not a ranking of usefulness. It is a ranking of how much work you have left to do afterwards, and it is the thing most people do not price before they start.
Transformation
You supply the content and ask for it in a different form: translated, summarised, reformatted, restructured, rephrased for a different reader, turned from prose into a table.
The reference is your input. Every claim in the output should be traceable to something you handed over, so the check is a comparison you can make at your desk with no other material. The characteristic failure is not invention but omission: something in the input that did not survive the transformation. That is easy to miss precisely because the output reads as complete.
Grounded answering
You supply a source — a document, a page, a specification — and ask a question about it. Some systems fetch the source themselves rather than being handed it [L3].
The reference is that source. The check is to read it at the point the answer relies on. Note what has and has not changed: attaching a source changes where you aim the check, not whether one is needed. Two questions remain live. Does the source say what the answer says it says? And is the source the right one to settle this question?
Generation to a specification
You state what would count as correct, and ask for something that meets it. A routine that satisfies a stated contract; a document in a required structure; a test that fails under a named condition.
The reference is the specification. This is the class the middle of this module lives in: Unit 7 asks for a routine, Unit 8 checks it against a reference you built, and Unit 10 is about what happens to the output when you make the specification more precise. The characteristic failure is that the output meets the specification you wrote rather than the one you meant.
Recall
You ask a question with no source attached and no specification: a date, an attribution, a requirement, a figure, who first proved something.
The reference is external and you do not have it. The generation step does not consult a record of what is true — Unit 3 explains the mechanism — so a recalled claim arrives with no evidence attached and no marker distinguishing the ones that happen to be right. Checking means finding an authoritative source, which is a research task of its own. Unit 18 is about doing that properly.
This is the class where the gap between how much a claim costs to produce and how much it costs to check is widest. Eight referenced claims can be generated in seconds and take an afternoon to verify.
Judgement
You ask which of two designs is better, whether an argument is convincing, whether a piece of writing is any good.
There is no reference in the sense the other classes have one, because there is no single answer to be right about. That does not mean there is nothing to check. A judgement rests on factual premises, and those can be wrong: an argument that one design is cheaper depends on a claim about cost. It can also be assessed against things you can state — the objectives, the constraints, the evidence offered, the trade-offs, professional standards, whose interests are served. What you cannot do is settle it by comparison, because there is nothing to compare it against. So the output is a proposal, and treating a proposal as a finding is the error. A confident register here is a property of the text, not evidence about the question.
How this relates to the triage you meet next
Unit 3 asks a second question about the same task: not where the reference lives but what kind of evidence settles it — something you can execute and compare, something you must look up in a record, or neither because an error would be cheap and visible.
The two cuts are not rivals, and mostly the first decides the second. A task with its reference in a specification is one you can check by executing and comparing. A task whose reference is a source, whether you supplied it or must go and find it, is one you check by reading the record. Judgement is the class where neither applies, which is why its output is a proposal.
So this unit gives you the cheaper question — where would the evidence come from — and Unit 3 gives you the sharper one, which is what to do about it.
The move worth learning
The classes are not fixed properties of a task. Which one you are in is often your choice, and moving between them is the most useful thing in this unit.
A recall question becomes a grounded question the moment you supply the source. "What does the regulation require?" is recall, and expensive to check. "Here is the regulation — what does clause 4 require?" is grounded answering, and the check is reading clause 4.
The same move works forwards. A vague request becomes generation to a specification the moment you write down what would count as correct, which is exactly what Unit 10 asks you to do, and why it is the unit that changes how much the rest of your work costs.
So the practical question is not "can it do this?" but "which class am I in, and can I move myself into a cheaper one before I start?"
One task, four classes
Take a single question: does our coursework policy allow me to use an AI system to draft an essay plan?
Asked with nothing attached, it is recall. The answer will name a policy, may quote it, and will read as settled. Nothing in it is evidence: to check it you must find the policy, which is the whole task you were trying to avoid.
Paste the policy in and ask the same question, and it is grounded answering. The check is now bounded: find the clause the answer relies on and read it. You may still find that the answer stretched a clause about brainstorming into one about drafting, but you will find it in a minute rather than an afternoon.
Ask instead for a table of every clause that mentions AI, one row per clause with its number and its exact wording, and you are in transformation. Every row is checkable against the document you supplied, and the failure to look for is a clause that quietly did not make it into the table.
Ask whether the policy is a good policy and you are in judgement. There is no clause that settles it. What comes back may be worth reading and is not a finding.
Four classes, one question, and the difference between an afternoon of checking and a minute of it is which one you chose. The same applies to the case this module is built on: "what is the eighth Bernoulli number?" is recall, and "here is the recurrence — compute it" is generation to a specification. Unit 6 has you build the reference, which is what makes the difference payable.
What these systems are not for
Not a list of forbidden tasks. A question:
If this output is wrong, how would I find out, and what would that cost?
Three situations where the honest answer is bad.
No reference is obtainable at a price you can pay. Questions about the future, about private matters of fact, or at the edge of what is known. There is nothing to check against, so any output is a proposal however it is phrased.
You would not recognise a wrong answer. This is the sharpest one, and it is uncomfortable because it is about you rather than the system. Outside your own competence, every surface cue you normally rely on — fluency, specificity, confident structure — is available to a wrong answer as readily as a right one. The move here is to change class: get a source you can check, or ask someone who can.
The doing is the point. Some work exists to develop your judgement, and an output that arrives without the work has skipped the thing the work was for. That is a statement about what the task is, not a rule about tools. Your institution's rules on assessed work are separate, and they are theirs to state, not this module's.
Summary
- Sort a task by where its reference lives, not by topic or by product. That sorting does not date.
- Transformation and grounded answering put the reference in your hands. Generation to a specification puts it in something you write. Recall puts it somewhere you must go and find. Judgement has none.
- The class sets what checking costs, and most people do not price it before starting. Recall and judgement are the expensive classes, which makes them the ones you owe the most checking — not the ones to avoid.
- Cheap to check is not the same as checked. A transformation can drop the one qualification that mattered, and it will do so fluently.
- You can often move a task into a cheaper class by supplying a source or writing a specification. That is the single most useful habit in this module.
- The limits are not a list of banned tasks but one question: if this is wrong, how would I find out, and what would that cost?
Then answer cp01-where-the-reference-lives and cp01-changing-class at the
foot of this page.
Timings: video 6 min, reading 10 min, activity 12 min.
Check yourself
This unit has 2 checkpoints, answered here. A wrong answer says which misunderstanding it matches. Nothing is uploaded.
Local progress is available in a supported browser.
Prefer a terminal?
The same questions, from course/lab/:
python3 selfcheck.py run --unit 1