Unit 13: Privacy and data handling

25 min
DEVELOPMENT REVIEW DEPLOYMENT - NOT READY FOR RELEASE

Video

Media Pending: Unit Video

Intended content: Full narrated video presentation, including visual assets, caption file, and transcript.

Learning purpose: Replace "do not paste confidential information" with a model of what

Planned form & duration: Video, ~4 minutes.

Accessible text alternative: The whole context leaves, repeatedly; classify by whose The written material below covers the same complete learning path.

Watch what changes

A dashed line divides your device from the provider. You send one question, and it crosses. Then three more pieces follow that you did not think you were sending: the attached file whole, the parts of it you were not asking about, and the earlier turns of the conversation. Three questions appear on the far side — whether it trains future models, how long it is kept and who can read it, and where it is processed. Finally a redacted question is sent instead, and only that one piece crosses.

Reading time: about 9 minutes. Activity: about 12 minutes.

The question this unit answers

Unit 7 says, in one paragraph, not to paste confidential or personal material into an AI system. That is a sound instruction and it is not enough to act on, because it does not say what happens to what you paste, and without that you cannot judge the cases the sentence does not cover.

This unit says what leaves your device, what is retained, and what "do not paste this" means when the material is not obviously secret.

It does not state your institution's rules or the law. Those exist, they differ by place and by organisation, and they are not this module's to invent. What the module can give you is the shape of the thing, so that when you read your own rules you know what they are talking about.

What leaves the device

With a hosted service — anything you reach through a browser or an app, rather than a model running on your own machine — what you supply leaves your device. Exactly what happens to it then is a property of the product, and products differ. The rule that holds regardless:

Assume material you submit to a hosted service may be transmitted to and processed by the provider and whoever it works with, unless that product's documentation and settings establish otherwise.

That covers your prompt, the conversation, any file you attach, and anything a tool fetched on your behalf.

Three things follow that people are routinely surprised by.

An attachment goes whole. You asked about one clause; the document went. If the file has tracked changes, comments, hidden columns, or a second sheet, those are text too. Whether the whole of it then reaches the model is a separate question — some products index or chunk a file and retrieve parts of it — but the file left your machine either way, and that is the part that matters for privacy.

Earlier turns keep going. A conversation carries its history, so something you pasted twenty turns ago is generally still in play rather than gone. Whether it is re-sent verbatim, summarised, or retrieved from storage varies by product.

Deleting the conversation is not deletion. It removes it from your view. Whether and when the provider's copies, backups and abuse-monitoring logs go is a question about that service's policy, not about the button.

What is retained, and by whom

This is where a module that cannot date itself has to be careful. Retention depends on the service, the plan you are on, and settings that change. So rather than facts that will be wrong next year, three questions whose answers you can go and find, and which your own rules will assume you have found:

  1. Is this input used to train future models? Often configurable, often different between a free tier and a paid or institutional one. This is the question consumer defaults are least likely to answer the way you would.
  2. How long is it stored, and who can read it? Providers commonly retain conversations for a period for abuse monitoring, which means staff access under stated conditions.
  3. Where is it processed, and under what terms? Server location is part of this and not the whole of it: which law applies can depend on who the data is about, who is responsible for it, and what the contract says. You are not expected to settle that. You are expected to notice that it is a question, and to ask whoever handles it where you work or study.

If you cannot answer all three for the service in front of you, that is itself the finding, and it makes the classification below strict rather than optional.

A classification you can apply in seconds

Not a legal taxonomy. Four bands, sorted by who is harmed if the material becomes readable by someone else.

Band What it covers What it settles
Public Already published; you could put it on a poster The privacy question, mostly. Not the rights question
Yours alone Your notes, your drafts, your code, your unsubmitted work That nobody else's privacy is at stake. Not whether you are permitted
Someone else's Anything identifying another person, or told to you in confidence That you must not paste it as it stands
Not yours to decide Employer or client material, unpublished research, anything under an agreement That the decision is not yours. Ask before, not after

These bands are one dimension, not a permission slip. That is the most important sentence on this page. A band tells you about the privacy risk — who is harmed if this becomes readable by someone else. It does not tell you whether you may use the material, and three other questions sit alongside it:

So the first row means "the privacy risk here is low", and the second means "no third party is exposed by this". Both leave the other three questions open.

The third and fourth bands are the ones that matter most, and they are the ones a sentence about "confidential information" does not reach.

Someone else's is broader than it sounds. A message a friend sent you. A draft your colleague shared. An interview transcript. A photograph with a person in it. Health, financial or immigration details about anybody. The test is not "is this marked confidential" but "did somebody else give me this expecting it to stay where I put it".

Not yours to decide is the band people miss entirely, because nothing about the material feels sensitive. A dull internal spreadsheet is often covered by an agreement someone else signed. The point of the band is that the decision is not yours to make on the material's apparent importance.

Working with material you cannot paste

The useful move is almost never "do not use the tool". It is to change what you send.

Redact and generalise. Replace names, dates, identifiers and distinctive details. "A patient" not a name; "a large employer in the sector" not the employer. The task usually survives this untouched, because you were asking about structure, not about the person.

Redaction reduces risk; it does not by itself make material anonymous or make its use authorised. A combination of details can identify someone with every name removed — a role, a date and a place is often enough — and an organisation may forbid sending even de-identified material to a service it has not approved. Removing the names is the beginning of that judgement rather than the end of it.

Abstract the question. You rarely need to send the document to get what you want. "Here is a clause I have rewritten with all identifying detail removed — does it say what I think it says?" is a different question from pasting the contract, and it usually gets the same answer.

Bring the tool to the data. Where a locally-run model or an institutionally-approved service is available, the first question — what leaves the device — gets a different answer, and that is often the whole difficulty solved. It does not remove the question. A local application may still check for updates, download models, sync to cloud storage, or send diagnostics, and the device it runs on has its own exposure. Ask what leaves, rather than assuming that local means nothing does.

Watch re-identification. Removing a name is not always removing a person. A combination of details — a role, a place, a date, an unusual circumstance — can identify somebody as surely as a name, especially in a small institution. Ask whether the remaining details would let a colleague work out who it is.

The one about your own work

Your unsubmitted coursework is in the "yours alone" band as far as privacy is concerned. Whether you may use an AI system on it is a different question entirely, governed by your institution's rules on assessed work, which this module does not state and cannot.

Keeping those two questions apart matters, because they have different answers and different authorities. "It is my own work so there is no privacy problem" can be true at the same time as "and I am not permitted to do this".

What this unit is not saying

It is not saying hosted services are unsafe. It is saying that what leaves your device is a fact you can establish, and that the classification depends on whose material it is rather than on how secret it feels.

It is not giving you a rule to follow instead of your institution's. Where they conflict, theirs governs. This gives you the vocabulary to read theirs.

Summary

Then answer cp13-what-leaves and cp13-whose-material at the foot of this page.

Timings: video 4 min, reading 9 min, activity 12 min.

Check yourself

This unit has 2 checkpoints, answered here. A wrong answer says which misunderstanding it matches. Nothing is uploaded.

Local progress is available in a supported browser.

Manage or delete this record

Prefer a terminal?

The same questions, from course/lab/:

python3 selfcheck.py run --unit 13

Your progress, and every checkpoint in the module