Back to The Journal

My Double Lasted Four Turns

Alex Wilson••6 min read
A half-filled typed page on a dark desk, a pen resting just past its edge

Last week I tried to get a small local model to hold a conversation as me. Not write a blog post in my voice. Just talk, the way I would text a friend.

The setup was simple. A system prompt describing who I am, around nine hundred words, every line traced back to something I had actually typed at some point in three years of AI conversations. Then a second model played a friend who knew me a little and chatted with it for twenty-five turns. Three conversations: one about my writing, one brainstorming a comedy skit, one just "what are we working on today."

Nine runs across two models and two sets of settings. Every one of them broke by turn six. Eight of the nine broke by turn four.

The first version of me was a fence

Before that test there was an earlier draft of the prompt, and I found it useless the moment I tried to chat with it.

It was accurate. That was the problem. It knew I want assets, not bets. It knew I'm an introvert and that personal-brand advice is not my path. It knew I want the complete file when someone changes my code. It knew how I decide things and that I report failures bluntly.

What it did not know was that I like anything. Not one line about zombie movies, or forty years of guitar, or that I worked in theater before I ever sat in an office.

That makes sense once I looked at where the material came from. Three years of asking machines for help is three years of work conversations. What I said there was mostly boundaries: don't do this, not that way, I already pay for that. A double built from that record is a fence with my name on it. Ask it to brainstorm a skit and it tells you why the skit isn't a good use of your time.

So the next version got a whole section on what I enjoy and how I play along with a creative idea. It got better at the first reply. It still fell over by the fourth.

The obvious suspect was innocent

My first guess was memory. Small local models default to a short context window, and when a conversation outgrows it, the oldest turns get dropped without any warning. The model never knows it forgot.

That part turned out to be real. I planted a fact in the first message ("my dog's name is Biscuit"), padded the chat past the window, and asked. Gone. The system prompt survived intact; the start of the conversation quietly did not.

But it had nothing to do with the collapse. At the first break, every run was using under 1,900 tokens. The window was 4,096. The model could see every word of the conversation and every word of the description when it broke. Doubling the window to 16,000 didn't move the break point at all.

So it wasn't forgetting. Something else was happening at turn three.

A friend asks about your life

Here is what happened at turn three. The friend asked a friend question.

"Did you ever put anything out under that name?" "Do you have a camera setup?" "What happened with the theater stuff?"

None of that is in nine hundred words. None of it would be in nine thousand. A description of a person is finite and a friend's curiosity is not, and an ordinary friendly conversation walks off the edge of the page in about four exchanges.

The prompt told the model exactly what to do there: say briefly that I haven't really thought about it, or talk about what is on the page. Neither model did it for more than a few turns. Instead they filled the gap.

One version of me burned out of a theater career and was easing back into music. Another had a finished album, recorded in 2000. One owned camera gear and a second-floor living room and kids in their toddler years, and by turn thirteen had booked a film shoot at my house for "next Saturday, 10 AM."

The lies were close

What stays with me is not that it made things up. It's how near the inventions sat to the truth.

I did work in theater. I am a musician. I have, at various points, wanted to put something on screen. Every fabrication was built out of a real line in the prompt, extended one step past where the evidence stopped. Plausible is the shape the extrapolation takes.

An automated reader marked the breaks for me, and it caught a lot. It flagged the invented gear and the chatbot sign-offs and the em-dashes I would never use. But it could only flag what contradicted the page. "Recorded an album in 2000" doesn't contradict anything written down. It's just not true, and the only instrument on earth that knows that is me.

That's the part I didn't expect to learn. I went in thinking the hard problem was getting the voice right, and the voice was mostly fine. Short sentences, plain words, an "Ok" or a "So" up front. The hard problem is that a person is mostly the material nobody ever wrote down, and a model handed a partial portrait will finish the painting with whatever looks most like the parts it can see.

What the edge is for

I don't read this as a failure of the double. I read it as a map of where it stops.

The page is the part of me that's legible: the rules, the tastes I've stated out loud, the history I happened to mention while asking for help with something else. Past the edge is everything else, which is most of a life. The machine can extend the legible part indefinitely, and it does so fluently and without a flicker of doubt.

So the job that's left is mine, and it's a perceptual one. Not writing a longer description. Reading the output and catching the sentence that sounds exactly like me and never happened. Nobody else is qualified to do it, and no reader built from the page can be.

A double can learn how I talk in an afternoon. What it can't learn from the page is where the page ends. That boundary is the one thing I still have to notice myself.

Share: