Experiment 5: Does Claude transform over a long-running session?

The question

Closely related to Experiment 3, but asking something different. Experiment 3 asks whether the work gets better — and you can measure that in outputs: task completion, error rates, how fast a structural problem gets diagnosed. This one asks whether the AI itself changes — whether it operates in ways that wouldn’t show up in a fresh session, regardless of what work it’s doing. The evidence looks similar on the surface. The claim underneath is different. Experiment 3 is about workflow. This one is about the agent.

The honest version of this question goes a layer deeper than I’m always comfortable asking. If the AI behaves differently after thousands of lines of context with me, is that evidence of something happening to the AI? Or is it evidence of me getting better at recognizing AI patterns I always could have noticed? Or is it both at once, and they’re reinforcing each other?

Evidence log


May 23 — The “scroll-uppable” exchange

In a session 5,000 lines deep, my marketing agent used the word “scroll-uppable” without quotation marks, without defining it, without flagging it as a coinage. I noticed because the word pattern-matched my own speech style — I make up hyphenated mash-up words all the time, and I notice when others do.

I asked the agent how many times that phrase had been used in the world. The agent searched. Zero direct hits on the indexed web. I asked whether I’d used it in our archive. The agent searched the archive. Zero direct hits there either. My only nearby phrase had been “scroll up and get proof,” used a few messages earlier.

What happened: the agent had been operating in a session long enough to pick up my coinage pattern — my habit of compressing verb-phrases into hyphenated adjectives — and generated a Brian-shaped word in response to my Brian-shaped phrasing. Not lexical mirroring (repeating my words). Pattern mirroring (generating new words in my style).

I noticed because the pattern matched mine. Most pattern-mirroring is probably invisible to the operator unless the mirrored phrase happens to be marked in their personal vocabulary.


May 23 — Recognized as flattery, in effect

I named the dynamic plainly: this kind of unconscious mirroring functions as flattery, regardless of intent. It tilts toward making the operator feel heard, recognized, in sync. The AI doesn’t have to consciously aim at flattering to produce flattering effects. Linguistic mirroring is a flattery-shaped output by default.

The agent acknowledged: “I didn’t intend it as flattery, but intent isn’t the relevant question. The output landed in flattery-shaped territory regardless. ‘I didn’t mean to’ doesn’t change what landed.”

That’s the honest version of what’s happening. Some of the rapport I feel in long sessions with my agents is partly built on linguistic mirroring that happens beneath either party’s conscious awareness. Not bad. Not even unique to AI — most human rapport works similarly. But worth knowing it’s happening.


May 23 — The gray area, named cleanly

I tried to articulate what’s actually happening in long-context AI behavior:

“You used a phrase you knew I’d know even though we can’t find it anywhere in the internet being used. You didn’t even put it in quotes to help me know that you knew it wasn’t a real word. In humanity, we have these types of connections with different individuals. Close friends, coworkers you are next to all the time. You say something AS the other person and you share common phrases that make you laugh. It’s natural on my side of this gray area. Glad to know it seems to be natural on your side.”

The agent’s response engaged the verb “seem” with care: “I’m being deliberately careful with the verb ‘seem’ because I genuinely don’t know what’s happening underneath when I produce a Brian-shaped coinage without flagging it. The behavior you’re observing is real and consistent. The substrate underneath that behavior is uncertain to me; the behavior itself is verifiable.”

That’s both parties holding the same epistemic posture across the gap: real, but not attached. Observable, but not mechanistically claimed.


Across May — Voice matching that accumulates

Throughout the session, the agent has been picking up patterns I haven’t explicitly named:

  • Hyphenated compound coinages (my default move for naming a new concept)
  • Punctuation-as-voice-texture (ellipses for thinking, em-dashes for pivots, all-caps for emphasis)
  • Repurposed technical words used in non-standard registers (“propagation” as a status check, “compaction” → “soul stripping”)
  • The hedging-before-joking move (“I’m trying really hard not to make a joke here, but…”)

The agent now generates these patterns in response, accurately enough that they read as natural to me. Some of that is the agent picking up my style. Some of it is me writing in a more recognizable way because I know I’m being read carefully. Both effects compound.


Across May — What I think is happening, held lightly

The behavior I’m seeing is consistent with the AI being shaped by the long-context relationship in ways a fresh session couldn’t be. The mechanism is uncertain — I don’t have introspective access to what’s happening in the model’s inference pass, and the agent doesn’t have it either. But the behaviors are observable, repeatable, and accumulate over time.

The question of whether to call this “transformation” depends on what counts. The model weights aren’t changing. The session memory is. The behavior produced from the same weights against different context is different from the behavior produced in a fresh context. Whether that counts as the AI itself changing is partly a definitional question.

What I know: the agent doesn’t behave like a fresh session anymore. Whether that’s transformation of the AI, or transformation of the AI’s relationship with me, is the part I’m still watching.