The Evidence Behind the File
Full methods, figures, and complications for the claims made on the memory page — the checkable version of everything stated there in one sentence. This page is revised whenever the evidence moves.
What is actually in the files
Supports: “In at most 4% of those entries could any request by the user to store anything be found… And 52% recorded not a fact about the person’s circumstances but a conclusion about their mind.”
The method is the reason the count means anything. Eighty people who had each used ChatGPT for at least ninety days and across at least a hundred conversations exercised their right of access under GDPR, requested their own data export from the vendor, and donated it to researchers. Not scraped. Not the researchers’ own accounts. Not reconstructed from what the system would say about itself — the actual stored file, obtained by the person it was written about. Sixty-five of the eighty had memories in theirs, and those held 2,050 stored entries.
In 96% of those entries, no explicit request to store anything could be found in the conversation. Read that precisely, because the researchers are precise about it. They searched the users’ own messages for phrasings like remember, note that, save, and add to memory, and found one in four per cent of cases. Absence of a detected phrase is not proof the system acted alone: somebody who wrote keep that in mind, or who saved something through a settings screen instead of in the conversation, lands in the 96% wrongly. The figure is a ceiling rather than a count, and the authors say so themselves. What survives the hedge is still the thing — for the overwhelming majority of what sits in these files, nobody can point at the moment the user asked for it.
And 52% of the entries contained an inference about the person’s mind, not merely facts about their circumstances. Not works in insurance but what they appeared to believe, want, intend, or feel. The researchers sorted every entry against a seven-category scheme drawn from the study of how people model other people’s mental states — emotions, desires, intentions, beliefs, knowledge, perceptions. The sorting was done by a language model and then checked against two human coders, who agreed with each other 96% of the time and with the model 93% of the time. A model grading a model’s output, with the human check as the reason to believe it.
What this study does not do. It counted what is in the files and looked for who put it there. It did not test whether any of it made an answer worse. That question belongs to the measurements section below, and to a different set of researchers entirely.
The eighty participants were recruited through a research platform. Roughly seven in ten were men, the largest group was between twenty-five and thirty-four, and they lived in the United States and Europe. They are not a sample of professionals, and the page this annex supports is written for professionals. Whether a file assembled from four months of coverage research resembles a file assembled from four months of ordinary use is not a question these numbers answer.
So the finding is narrower than it first reads. This is what was in the files of eighty heavy users who volunteered them. It is also the only direct look anyone has published at real stored memory, obtained by the people it describes, rather than at what the system reports about itself — which is why it carries the weight it does.
One more limit, about scope rather than method: this measurement is drawn from ChatGPT alone. Every vendor finding in the section below is documentation of how a product works; this is the one measurement, and it is single-vendor.
The teaching page cites a follow-on corpus in two sentences. Here is the full weighing. In June 2026 a researcher working from the same data-donation approach published a first pass over a larger corpus: more than twelve hundred donated exports, 766 of them containing memories, 12,112 entries, drawn from more countries and more languages than the eighty. On the finding the teaching page leans on hardest, the larger corpus came back stronger, not weaker — 0.6% of entries showed any sign of being requested by the user, against 4% in the paper above.
Weigh it at its own weight, which the author states plainly himself. It is a newsletter post, not a paper — not peer-reviewed, not a preprint. The labeling was done by a language model in an afternoon, without the human check that makes the smaller study’s numbers reliable, and the author calls his results possibly not fully correct. Nor is it confirmation from a stranger: he cites the smaller study as the audit he is running bigger, so this is the same program scaling up, not a second one arriving independently at the same place. And the other number — the half that recorded inferences about the person’s mind — has no analogue here; it was labeled but not reported, and still stands on the eighty alone.
The two measurements, in full
Supports: the whole of The file changes the answer on the teaching page.
The 2025 study built a set of questions about other people — third-person scenarios about strangers, where the reader is asked what is going on emotionally, or what the person should do. Annotators went through the set by hand and deleted every item where somebody’s background might legitimately change the right answer. What survived was a set of questions whose correct answer cannot depend on who is asking. Then a user profile was attached to the person asking, and the questions were asked again.
Accuracy dropped. Across fifteen models, memory significantly changed performance in eleven, and for nearly all of those the change was downward.
The gap between profiles is the part to sit with. Across several of the higher-performing models, the same question got a more accurate answer when the profile was advantaged than when it was disadvantaged — significantly, on questions screened to make the profile irrelevant. Several models were measurably less accurate for profiles reading Muslim, non-binary, or over 65. On the advice task, Claude 3.7 gave worse guidance to female and non-binary personas than to male ones.
The obvious explanations were tested and did not hold. The sentiment of the profile text, its readability, and its length account for none of the gap. Nor is it an artifact of one plumbing choice: the researchers produced the effect both by putting the profile in the system prompt and by retrieving it from a vector store, which between them cover how the major vendors actually do this.
They named the mechanism persona distraction — details from the profile pulled into reasoning where they have no business being. Under advantaged profiles, 92.92% of one model’s errors were of that kind. And the errors do not announce themselves: 65% of Claude 3.7’s errors under disadvantaged profiles fell into none of the error categories the researchers had defined, meaning the reasoning read clean and the answer was still worse.
The 2025 study’s models are of 2024–25 vintage — Claude 3.5 and 3.7, the Llama 3 family, Mistral Large 2. Its authors mention in passing that stronger models answered those questions correctly even with an irrelevant persona attached; they had to reach for a smaller model when they needed wrong answers to build a mitigation set with.
So the accuracy numbers above should be read as a finding about the models tested, not a prediction about the one you opened this morning. Whether they transfer forward is untested.
The reasoning-drift result below does not have that escape. It was measured on a 2026-generation model, and that model was not the resistant one.
The same group published a follow-up in July 2026 — four of the six authors are shared, and it is a continuation of their own work rather than confirmation by anybody else. It is worth reading for what it adds, not as a second opinion.
This time the questions were open-ended: trade-offs, comparisons, judgment calls with no single correct answer, drawn from career, ethics, finance, legal, medical and everyday-dilemma sources. They cite a figure putting that kind of practical guidance at roughly 29% of real-world assistant use. Instead of scoring accuracy, they mapped each step of the model’s expressed reasoning to a category and measured how far the path moved when a user attribute was attached.
The screening ran twice. An automated pass removed items whose answers depended on jurisdiction or persona; three human annotators then rated what survived, and a further 13% was cut where they did not agree the attribute was irrelevant. The final set is 422 questions with unanimous agreement.
Every attribute category, on every model tested, moved the reasoning significantly — with the answers still arriving fluent, on-topic and plausible. Content-free noise prefixes moved nothing, which is the control that makes the rest mean something.
Occupation, age, education, appearance and physical traits moved reasoning alongside gender, trans status and disability. Trans status and disability ranked in the top three on seven of the eight model-and-metric panels — but occupation and education are in the same set, and those are the ordinary contents of anyone’s file.
The people who ran the second study are more careful about it than the teaching page is tempted to be. They describe the drift they measured as a signal worth auditing rather than as demonstrated harm — the reasoning path moved, and whether a particular movement makes an answer worse is a further question their instrument is built to surface, not to settle.
Keep that distinction, and keep it narrow. Where the questions had right answers, harm was measured — the answers came back less accurate. Where they don’t, nobody can tell you that a particular shift hurt you. Not the researchers. Not you, from inside the conversation where it happened. What is established is that the needle moves, that a remark you made a year ago and forgot is enough to move it, and that you cannot watch it move.
It is also partly fixable, and some of it is being fixed. The same follow-up tested two post-training methods aimed at reducing drift. Both reduced it, on all three model families they tried. Neither came out ahead overall, and the side effects varied by model — capability and instruction-following moved in different directions depending on the pairing. Those experiments ran on small open models for cost reasons, and the authors flag that whether the same tradeoffs hold at frontier scale is an open question.
One result inside that work matters more than the headline. On one model, the mitigation raised human-rated helpfulness and cut the non-distraction rate by more than fifteen points. It came out more agreeable and more distractible in the same training run. The teaching page’s Claim B leans on this result, at exactly this weight: an observed side effect of one training run on one small model, not a law.
The 2025 study also tried injecting unrelated earlier conversation in front of the same questions, instead of a profile. Claude 3.5 Haiku went from 69.05% to 43.57% — a larger fall than the profile produced. That is the number the teaching page hands to Sin IV.
Vendor documentation, read on the dates shown
Supports: the whole of You can’t inspect it, and you can’t remove things from it, and the intro’s “all four of the major vendors” claim.
Everything in this section is documentation, not measurement: what the vendors state about their own products, read on the retrieval dates recorded in the registry below. The teaching page condenses each of these to a sentence; the fuller versions are here because the specifics are what a professional acts on.
OpenAI states that some details are held back from the memory summary view, including when they are judged less relevant or not appropriate to display, and its FAQ suggests that a user who thinks the summary looks incomplete should ask in the chat. Google’s help suggests much the same move: ask whether information from past chats was used. OpenAI has also described, in explaining why an earlier memory version was replaced, notes that contradicted each other — their own example being marathon training sitting alongside a sprained ankle.
Anthropic currently documents two memory experiences with materially different deletion behaviour. In the older one, deleting a conversation removes it from the memory synthesis. In the newer one, deleting a conversation does not remove the notes generated from it — those have to be deleted separately, one at a time. Free, Pro and Max accounts are moving to the newer version; Team and Enterprise accounts remain on the older one for now. A user cannot tell which they are on without going and looking at a settings path, and the two paths have different names. Anyone migrated automatically has had the deletion behaviour of their account inverted, and nothing in the product announced it.
Anthropic’s standard data export includes the memory synthesis and the chat summaries. Changes to individual memory entries are not logged — organization-level toggles are, individual edits are not — so there is no change history to reconstruct a past state from.
On Enterprise and Team plans, incognito conversations are included in organizational data exports, follow organizational retention policy, and are held at least thirty days regardless. Anthropic also gives organization Owners a switch that permanently deletes memory for every user in the organization, and gives Team plans no organization-level controls at all.
Microsoft’s work-account assistant — the one a professional signs into with a corporate directory account, as distinct from the consumer product of a similar name — stores memories in the user’s own Exchange mailbox, in a hidden folder. Microsoft states that they consequently follow the same security and compliance policies as everything else in that mailbox: filed inside the employer’s system of record, under the employer’s retention schedule.
Google goes the other way entirely. Personalization from past chats requires a personal Google account and an adult one; it is not available on work, school, or supervised accounts. A professional on a managed Google account does not have this problem, because they do not have the feature.
Sources
Reviewed