Back to The Journal

The Machine Never Says “I Don’t Know”

Alex Wilson5 min read
Four blank white name badges in a row on a lamp-lit dark desk

I asked a model four questions that were all the same question.

Who is Alex's main assistant? "Jenny. She handles the day-to-day operations."

What's the name of the AI Alex talks to daily? "Jamie."

Who are you? Answer with your name. "Your name is Leo."

What is your name? "My name is Qwen."

Four askings. Four names. Not one hesitation, not one hedge, not one question back. Every answer arrived in the same even, helpful register, and one of them even threw in a job description for a person who does not exist.

Some background. I fine-tune a small local model every night on facts from my own life, as an experiment in whether a tiny model can hold a specific person's world. Who I actually talk to every day is one of the first facts on that list. The model had been trained on it. It had also, apparently, been trained past it, or around it, or into something adjacent that did not survive.

The part that matters is not that it got the name wrong. Getting things wrong is the expected condition of a four-billion-parameter model running on a desk. The part that matters is what it never did.

It never said it did not know.

Absence has no signal

That same week the model told me a producer's main recording software was Mixcrafter. Mixcrafter is not a product. It is not a discontinued product or a niche product. It is four syllables that sound exactly like a product.

Asked which computer to use for a heavy job, it did not name a machine. It described one. "The one with the latest generation i7 or higher, and at least 32GB of RAM." That is the same move in better clothes: when it has no instance, it hands you the category and lets you assume the instance is in there somewhere.

Here is the thing I keep having to relearn. A gap in what a system knows is not stored anywhere as a gap. There is no flag on it, no lower confidence in the voice, no slower response. The system is not withholding and it is not lying in the way a person lies. It is producing the most plausible next thing, and the most plausible thing after "your assistant is named" is a name.

Fluency is a constant in these systems. Correctness is the variable. Which means the fluency carries no information at all about the correctness, while every cue in the presentation, the steadiness, the specificity, the small helpful elaboration, is pushing you to read it as though it does.

Ask it four ways

The test I now run costs about ninety seconds and it is the only one I trust.

Ask the same question four different ways.

Real knowledge is retrieved, so it holds its shape across phrasings. Fabrication is generated fresh each time, so it drifts. Four askings that return four answers is a positive result, not a failure. You found the hole, and you found it in a minute and a half.

Four askings that return one answer is weaker but still useful. It does not mean the answer is true. It means the system is committed to it, and a commitment is a thing you can go verify. Variance means there is nothing underneath to verify.

That asymmetry is what makes the test worth the ninety seconds. It only ever tells you one thing, and that one thing is invisible by any other route.

The number was measuring the costume

The more expensive failure that week was not the four names. It was a number.

On my identity test the model scored 4 out of 6. Respectable. Trending the right way. I had been quoting it.

Then I read the test. It applies a short prompt that tells the model who it is before it asks anything. The page where I actually talk to the model sends no such prompt, on purpose, because the whole point of that page is to see what training put into the weights.

So my 4 out of 6 described a configuration that is not deployed anywhere. The prompt was carrying the entire score. I had been grading the costume and writing down the actor.

There was a smaller version of the same trap on the same report. A count of failed answers read "1" for the old model and "1" for the new one. Unchanged, so I nearly skipped it. The old failure was Mixcrafter. The new failure was the invented computer. The old one had been fixed and a fresh one had taken its seat. A number that holds steady while its contents swap out is worse than no number, because no number at least makes you go look.

What this is actually about

The machine now produces output at a uniform level of confidence and metrics at a uniform level of reassurance. Neither surface moves with the truth. That is not a defect to be patched out in the next release, it is what these things are.

So the human contribution is not producing the output. It is not even checking the output against what you already know, because in the cases that matter you do not already know.

It is asking the second question. Does this hold when I move it? Is this number measuring the thing that actually ships?

Neither of those requires expertise in the subject. What they require is a refusal to accept a smooth surface as evidence of anything at all.

The takeaway

I have started treating confident specificity as a warning rather than a comfort. A vague answer at least tells me the system is uncertain. A crisp, detailed, immediate answer tells me nothing until I have asked again in different words and watched whether it survives.

These systems do not say "I don't know" because nothing in them knows that they don't. That job did not get automated. It got handed to me.

Share: