Video
Media Pending: Unit Video
Intended content: Full narrated video presentation, including visual assets, caption file, and transcript.
Learning purpose: Replace the idea that a prompt is the input with an accurate picture of
Planned form & duration: Video, ~5 minutes.
Accessible text alternative: Your prompt is a minority of what reaches the system, the The written material below covers the same complete learning path.
Watch what changes
Reading time: about 9 minutes. Activity: about 14 minutes.
Why the same question gives different answers
You ask something, get a useful answer, ask a colleague to try the same thing, and they get something different. Or you return to a conversation a week later, ask a follow-up, and the answer no longer fits what you remember agreeing.
Neither is mysterious, and neither is evidence of anything about the underlying model. The explanation is almost always the same: the two questions were not the same question, because a prompt is not what reaches the system.
This unit is about what actually reaches it, and what follows for checking.
What actually goes in
Unit 3 described generation as conditioned on a context. That context is a single sequence of text assembled just before generation, and your typed prompt is only part of it. The rest typically includes:
- Instructions you did not write. A deployed product typically prepends its own standing instructions about tone, refusals, format and role. You do not normally see them, and they can change between releases.
- Earlier turns of this conversation. Generally all of them, though a long conversation may be truncated or summarised rather than sent whole. This is why a long conversation costs more per question than a short one, and why something you said twenty messages ago is still shaping the answer.
- The contents of attachments, converted to text.
- Retrieved passages, if the system searched (Unit 4).
- Tool results from earlier in the exchange.
Two consequences follow immediately, and they are the practical content of this unit.
Your prompt is a minority of the context. The thing you are tempted to perfect is one contribution among several. This is why "prompt engineering" framed as finding magic words misses: much of what determines an answer is what else is in the window.
The same words in a different conversation are a different input. So "it gave me a different answer" is not necessarily sampling, or a changed model, or anything about the system at all. It may simply be that the input differed, because the input was never just your sentence.
The window has an edge
The context is finite. Every system has a limit on how much text it can attend to at once, and when a conversation or a document exceeds it, something must go.
What happens then varies and is not usually shown to you. Earlier turns may be dropped, or summarised into something shorter, or a long document may be cut into pieces of which only some are retrieved. All are reasonable engineering; all mean that material you supplied may not be in the context when the answer is generated.
This produces a failure mode with a distinctive shape, and it is the one worth carrying out of this unit:
The system answers as though your document said less than it does — and the answer is confident, because from where it was generated the document did say less.
You supplied a forty-page report and asked a question about a clause on page thirty-one. The answer is fluent, cites page four, and is wrong. Nothing about its register distinguishes this from the case where the whole document was present. This is why a locator matters — the part of a citation that says where to look, such as a page, a section or a clause. An answer that names a place lets you find out whether it read that place at all.
The practical response is not to trust a limit you cannot see. For anything that matters, ask about a bounded piece of a document rather than the whole of it, and ask for the passage the answer relies on.
Attachments are not references
Attaching a document does two things worth keeping apart.
It moves the task into grounded answering (Unit 1): the reference is now something you hold, so the check is bounded.
It does not make the answer an extract. What comes back is still generated text conditioned on the document. It can paraphrase past what the document supports, attach a real page number to a claim from elsewhere, or answer from the part of the document that survived the window while sounding as though it read all of it.
So the discipline is short and it is the same one every time:
- Ask for the passage, not just the answer.
- Read the passage in the document, not in the answer.
- Ask whether that passage settles the question you asked.
Step 3 is the one people skip. A passage can be genuinely present, genuinely quoted, and about something adjacent to your question.
Three things that change the answer, and are not magic words
Stating what would count as correct. This is the move from Unit 1: it turns a vague request into generation to a specification, and it is what Unit 10 is about. It changes the answer more reliably than any phrasing trick, because it changes what you can check the answer against.
Giving an example of the form you want. Useful and cheap, with one hazard: an example constrains the shape of the answer, not its truth. A response in exactly the format you demonstrated can be entirely wrong, and the matching format makes it read as compliant.
Starting a new conversation. Underrated. If a conversation has gone somewhere you did not intend, the accumulated context is now working against you, and clearing it is more effective than arguing with it. This is also the honest way to ask a second time: a fresh conversation is a different input, and if you want to know whether an answer depends on how you got there, that is how to find out.
Note what none of these do. None makes the output self-verifying. They change what you get and how easily you can check it; they do not remove the check.
One before and after, on a non-code task
You have attached a twenty-page departmental handbook and you want to know what it says about late submission.
Before.
What does this say about late submission?
What comes back is likely to be fluent, plausible, and about late submission. It may also be right. The problem is not that it is wrong; it is that you cannot tell, and finding out means reading the handbook, which is the work you were trying to bound.
After.
Using only the attached handbook, list every rule it states about late submission. For each one give the section number and quote the sentence it comes from. If the handbook does not state a rule about something, say so rather than filling the gap.
Three things changed, and none of them is a magic word.
- It states what would count as correct — every rule, from this document only. That is Unit 1's move into generation to a specification, and it gives you something to compare the answer against.
- It asks for locators. Each row now points at a place you can open. A fabricated section number is visible immediately; an unbounded paraphrase is not.
- It says what to do about absence. Silence about a rule and a statement that no rule exists are different findings, and a request that does not distinguish them will get the more fluent one.
What has not changed is worth being equally clear about. The second answer is not verified. It is checkable, which is different and is the whole point: you can now spend two minutes confirming three section numbers rather than an afternoon reading a handbook to find out whether you were told the truth.
And if a quoted sentence turns out not to be at the section given, you have learned something specific, not "the answer was wrong" but how it was wrong, which is what Unit 9 turns into a named vocabulary.
What follows for checking
Everything above cashes out as three habits.
Record the whole input, not just your prompt. When you keep a prompt for your evidence pack — as Unit 7 asks you to — a prompt without its conversation is not a record of what produced the output. Note whether the conversation was fresh, what was attached, and whether anything was retrieved.
Treat a fresh conversation as the reproducible case. If you want to know whether a result holds, ask in a clean context. Repetition inside the same conversation is the least informative kind, because the earlier answer is now part of the input.
Ask for locators, always. "The policy permits this" is unbounded. "Clause 4.2, second sentence" is bounded. The second costs the system nothing extra and turns an open-ended search into a check you can finish, and when the locator turns out not to exist, you have learned something the fluent version would have hidden.
Two things follow that are easy to hear as opposites. Prompting matters, and Unit 10 is about writing a specification precise enough to be worth prompting with. What does not follow is that a better phrase reaches material the system never received: wording cannot recover what is not in the context. Those are different claims and this unit means the second.
Long documents are usable. The limit is one you cannot see from the outside, so the move is to supply the part that matters rather than to hope the whole arrived.
Summary
- What reaches the system is an assembled context: standing instructions you did not write, the whole conversation so far, attachments, retrieved passages and tool results. Your prompt is a minority of it.
- Different answers to "the same question" usually mean the inputs differed, because the input was never just your sentence.
- The context has an edge, and what falls off it is not shown to you. The characteristic failure is a confident answer generated from a document that was only partly present.
- An attachment moves the task into grounded answering; it does not turn the answer into an extract.
- Ask for the passage, read it in the document, and ask whether it settles your question. Record the whole input, not just the prompt.
Then answer cp05-what-reaches-the-system and cp05-locators at the foot of
this page.
Timings: video 5 min, reading 9 min, activity 14 min.
Check yourself
This unit has 2 checkpoints, answered here. A wrong answer says which misunderstanding it matches. Nothing is uploaded.
Local progress is available in a supported browser.
Prefer a terminal?
The same questions, from course/lab/:
python3 selfcheck.py run --unit 5