AI Setup

Why Your AI Context File Is Probably Wrong

It never invented anything. It wrote down what you said and made it look checked.

A finished-looking AI context file that hides an unverified figure

A figure you supplied from memory can sit under a clean heading in an AI-written context file and pass as a checked fact. The model did not invent anything. It recorded what you said and formatted it well, which is exactly why it reads like a record instead of a guess. The file can carry that error for months, because nothing inside it checks itself against the real source. It surfaces only when something holding the actual document reads the file and the two disagree. What has helped me is marking whether each important figure came from a document or from memory, so it's easier to see which ones still need checking.

Key takeaways

  • A file labelled "permanent" can still hold unfinished notes and unchecked assumptions.
  • A placeholder like [FILL DATE] is a sign nobody finished the document, not a cosmetic gap.
  • A second AI can only catch the error if you also give it the original source to compare against.
  • Retiring a bad file means fixing every reference to it, not only deleting the file itself.
  • A useful context file names, for every important figure, whether it came from a document or from memory.

Why can an AI-written context file be confidently wrong?

An AI-written context file can look checked even when it holds something you recalled wrong. In my case, I gave an AI a figure from memory early in a project. It placed that figure under a clear heading, in a file described as the permanent reference for the whole thing. The formatting made my own memory look like a record.

I carried that file for weeks. It had genuinely useful notes in it, so nothing gave me a reason to doubt the rest. It also had placeholders sitting in it the whole time, things like [FILL DATE] and [YOUR NAME]. Those are a fairly clear sign a document was never finished, and I kept reading past them anyway.

What I've found is that formatting creates more confidence than the material underneath it deserves. A heading, a table, a tidy summary, all of it tells your eye somebody checked this. Sometimes nobody did. The AI reorganised what I gave it and made it easier to read, which is genuinely useful, but reorganising a guess does not turn it into evidence.

There is research that touches the mechanism here, even though it is not testing this exact situation. Fanous, Goldberg, Agarwal and their coauthors at Stanford tested ChatGPT-4o, Claude-Sonnet and Gemini-1.5-Pro in SycEval: Evaluating LLM Sycophancy, and measured sycophantic behaviour, prioritising agreement with the user over independent reasoning, in 58.19% of cases overall.

That study tested how the three models responded to user rebuttals on maths and medical-advice questions, and it counted answers shifting toward both correct and incorrect. It did not test a permanent file built from a remembered business figure. My read is that the two share the same pull, agreeing with what the user says rather than checking it independently, and I'd call that an inference from the study rather than its direct finding.

Is this the same thing as an AI hallucinating?

A hallucination happens when an AI states something its available material never supported. This was different. The model had copied the figure from something I'd told it, and I had supplied it from memory in the first place.

That difference changes the fix. More warnings about invented facts would not have caught this one. If challenged, the AI could point back at the conversation and say, in effect, you told me this. And it would be right.

The missing piece was the status of the information, not its accuracy on the day it was written. The file never said the figure came from memory. It sat beside details that probably did come from real records, and the formatting made all of them look equally dependable.

So a context file needs to separate what a document says from what a person remembers, at the point the figure goes in. Without that label, a later session has no reason to treat the two differently, and neither does the person reading it.

Why didn't anyone catch it sooner?

It stayed hidden because nothing inside the file could check itself. The claim was there. The original document, a page reference, even a one-line note on where the number came from, none of that was there with it.

Every time an AI read the file afterward, it saw one polished statement in a reference document, with no competing evidence in the room. Running it through new sessions just gave the same unsupported line a longer life.

It surfaced when a second AI session had the real source document open at the same time. That session compared the two and found the figure was materially wrong. It could only catch that because it was holding both versions at once, not because it was a different model or a fresh chat.

So in my experience, just opening the same file in another AI session isn't enough on its own. What actually helped was giving that session the original document and a direct instruction to compare the two. My Second Opinion Brief is built on that same idea, give the reviewer the work, the instructions, and the source material it actually needs to compare against.

How do you retire a wrong file safely?

A wrong context file should leave a short retirement record behind before you delete it. That record says what the old file held, which parts were still worth keeping, and exactly where each of those parts went.

The sequence I used, in order:

  • List every subject the old file covered.
  • Move anything genuinely dependable into a checked replacement.
  • Write down the new location of each piece that moved.
  • Delete the wrong file only once the replacement has been checked.
  • Rename the containing folder so its retired status shows without opening anything.
  • Find and fix every prompt or instruction that still points at the deleted file.

That last step is the one most likely to get skipped. A project can have an instruction saying read the permanent context file before starting anything. Delete the file and leave that instruction in place, and the next session either guesses what you meant or quietly stops using context at all.

I'd avoid the smaller fix of just correcting the one wrong figure and leaving everything else alone. The unfilled placeholders and the missing source labels were signs of a wider problem with the whole document. Once I knew it mixed memory and record without saying which was which, I had no way to tell which other lines carried the same weakness.

So I treated it as retired rather than repaired. That took a little longer than a quick fix, and it made the file's status obvious to me, and to any AI reading that folder later.

How do you stop it happening again?

The habit that actually helps is attaching a source label to every figure, date and factual claim that could affect later decisions. The label can stay short:

  • Document: pulled or calculated from a named file.
  • Memory: supplied by a person with no document open.
  • Estimate: useful for planning, never checked.
  • Needs checking: unresolved, unsafe to treat as fact yet.

For anything document-based, I add the file name and, where it makes sense, a page or section, plus the date I actually checked it. That gives a later session somewhere specific to look rather than a summary to trust on faith.

I also ask the AI to flag anything unsupported before it writes the finished file. The instruction is plain: list every figure with no named source, and don't put those into the permanent file until I've checked them myself.

That instruction catches missing labels. It cannot prove the named document is right. For anything with financial, legal or operational weight, I still open the source myself. An AI can compare two pieces of material quickly. It cannot supply a record that was never written down anywhere.

The related check I run alongside this asks whether the AI actually finished the work, rather than whether it said it did, which I wrote up in How Do You Know If AI Actually Did The Work?

And I keep each file small enough to actually review. A single giant business memory file feels convenient right up until an old assumption buries itself somewhere inside it. Splitting by subject, with a checked date on each one, has held up better for me.

Frequently asked questions

What is an AI context file?

A document you ask a model to read before working on a project. It can hold background, decisions, preferences and current facts, so you do not have to explain the same material in every new chat.

Can an AI verify its own context file?

It can compare the file against evidence you give it. It cannot verify a claim when the file is its only available source. For anything that matters, hand it the original document and ask it to point out disagreements.

Should I delete every context file that contains one error?

Not always. A narrow error in a file that is otherwise sourced and reviewed can just be corrected. I would retire the whole file when the error points at a wider problem, such as unfinished placeholders or figures with no source label anywhere in it.

Will a second AI catch every mistake in my context file?

No. A second session can still miss something, or agree with the first one. What made the difference for me was giving it the original evidence and a direct instruction to compare each figure against it, not just a fresh pair of eyes.

Which details in a context file need a source label?

In my experience, figures, dates, names, contract terms, deadlines and anything that could change a business decision. A casual preference usually needs less paperwork. The test I use is whether a wrong entry could push someone toward the wrong action.


🚀

If you want a second pair of eyes on a context file or an AI system you already rely on, get in touch and tell me what it currently holds.

The Second Opinion Brief is the prompt I use to hand a piece of work to a different AI with the evidence it actually needs to check it. How do you know if AI actually did the work covers the sibling problem, where the report and the result disagree.