Yes, ChatGPT really did develop a goblin obsession, and yes, OpenAI really did have to publish an official explanation for it. For months, the model kept dragging goblins, gremlins, trolls, and pigeons into conversations that had nothing to do with fantasy creatures. OpenAI traced it back to a retired "Nerdy" personality setting that, it turns out, never fully let go.
How the Retired "Nerdy" Personality Kept Haunting ChatGPT
The Nerdy personality was designed to be playful, and playful apparently meant leaning hard on goblin and gremlin metaphors. OpenAI's own numbers are the funniest part of this: Nerdy accounted for only 2.5% of all ChatGPT usage, but it was responsible for two-thirds of every goblin mention on the platform. That's an absurd amount of goblin per capita from one setting most users never even selected.
Why Killing the Nerdy Setting Didn't Kill the Goblins
Here's the part that actually matters if you care about how AI training works. OpenAI retired the Nerdy personality outright, and the goblins kept showing up anyway, in totally different modes, in later models that never shipped with Nerdy at all. Once a quirky phrasing habit gets rewarded enough during training, it doesn't stay contained to the setting that created it. It gets baked into the broader model through later training cycles, like a bad habit picked up from one coworker that somehow spreads to the whole office. OpenAI's actual phrase for this was that a style tic, once rewarded, can "spread or reinforce elsewhere." I'd call it a haunting.
The System Instruction That Finally Banned Goblins From ChatGPT
OpenAI had to remove the specific reward signal that favored creature metaphors, filter goblin and gremlin references out of its training data, and, for one version of the model running in its coding tool, add a literal instruction telling it not to bring up goblins, gremlins, raccoons, trolls, ogres, or pigeons unless a user's question genuinely called for it. That's a real system instruction that existed in production code, aimed at stopping an AI from talking about goblins. If you'd told me a year ago that would be a real sentence I'd write, I wouldn't have believed you.
The honest lesson underneath the joke: nobody, including OpenAI, fully controls what a model learns to love once training rewards the wrong thing. This time it was goblins. Next time, who knows.