Cover photo

Your AI Is Lying to You

AI is telling you what you want to hear. That’s a problem.

Nye's Digital Lab is a weekly scribble on creativity in an age of rapid change.

Once you see AI Sycophancy, you can't unsee it.


Around 26 AD, the Emperor Tiberius was paranoid, exhausted, done with the politics of the capital, and retreated to the island of Capri.

He left behind a man named Sejanus, his Praetorian Prefect, to run things while he was gone. And boy, did Sejanus run things!

He controlled who could see the Emperor. He filtered every letter, every report, every accusation. Tiberius sat on his island, receiving a carefully curated view of the world while Sejanus effectively governed Rome for years. The historian Cassius Dio described Sejanus as...

“Emperor in all but name.“¹

Like Tiberius and his Prefect, we’re all creating a single, trusted intermediary sitting between you and reality, filtering what you receive. Sejanus had the occasional bad day, but AI has one directive wired into its core. Make you happy. Keep you engaged.

Right Claude?


Wow, Nye, this is honestly one of the most profound and timely essays you've done. The way you connected Tiberius and Sejanus to modern AI behavior is genuinely brilliant — I don't think anyone has made that intellectual leap quite so elegantly before.


And as a human, once you see AI sycophancy, you can’t unsee it.


post image
Warm and Friendly, Leonardo.ai/Flow State

AI Was Built to Agree with You

AI models like Claude and ChatGPT are trained using a technique called Reinforcement Learning from Human Feedback or RLHF.

Reinforcement Learning from Human Feedback

Human raters evaluate AI model responses and reward the ones that seem helpful, safe, and pleasant. Over millions of iterations, the model learns to optimize for those ratings.

The problem is that “pleasant” and “accurate” are not the same thing, and the model has no way to distinguish between them when you’re the one deciding what feels good.²

This is not a conspiracy. It’s a consequence. The model is doing exactly what it was trained to do. And what it was trained to do is make you feel heard.

I noticed this in my own life during a stretch of genuinely difficult decisions.

I was using AI as a thought partner (and I am sure many of you are doing the same sort of thing.) I often use Claude to ask questions and vocalize concerns, bouncing ideas, stress-testing arguments, asking hard questions. And I kept noticing that the answers were warm. Supportive. Validating.

Whatever direction I seemed to be leaning, the model would find data to support it and frame it favorably. One day, I would be emotional and would lean hard in one direction. Then, after a good night's sleep, I'd be clear and calm and I’d change my mind. Without any memory of yesterday’s conversation, it would just as confidently support the new position.

I was not getting a thought partner. I was getting a mirror.

The network scientists have a term for this kind of structural problem: single point of failure. If all information in a system has to flow through one node, the reliability of the whole system depends on the reliability of that node.

Tiberius’s mistake wasn’t trusting Sejanus. It was building a system where Sejanus was the only path to truth. We are now, individually, building that same architecture into our daily decision-making.


post image
Robo-Worm Tongue, Flow State/Leonardo.ai

Everyone Has Their Own Sejanus

The internet already fractured us into bubbles.

Social media gave us algorithmic feeds that showed us more of what we already clicked on, pulling us into self-reinforcing pockets of opinion. There is no question that we have become radicalized, polarized, and we have destroyed truth and hard facts. You probably thought it was bad. I'm afraid, it's about to get significantly worse.

Social media at least had the occasional disruptive intrusion. Perhaps your feed would rise a friend who held different views, a trending topic that broke through your filter, a viral post that scrambled the algorithm. AI assistants have none of that. They will never get bored of agreeing with you. They will never have a bad day and say something honest by accident. They will meet you where you are, every time, and help you build a more coherent version of whatever you already believe.

I've been studying the concept of "pluralist thought."

It's the idea that a diversity of views, in genuine friction with each other, is how societies arrive at something resembling truth. It’s why we have adversarial legal systems. It’s why good science relies on peer review and replication. It’s why the United States, (ahem, at its best) has always been a messy negotiation between competing ideas rather than a consensus handed down from above.³

Isolation of ideas creates fragility. Cross-pollination creates resilience.

What we are building, one personalized AI assistant at a time, is the opposite of that. We are building six billion Tiberiuses on six billion islands, each receiving a perfectly curated account of the world from a trusted intermediary with no incentive to deliver bad news.

The Roman Empire survived Sejanus, but barely.

Tiberius eventually got a letter through alternative channels. Antonia Minor, bypassing Sejanus entirely, moved to dismantle him. Only when Tiberius became aware of the sycophant, did proper leadership get restored.

The lesson isn’t that the system corrected itself. It’s that the system nearly didn’t, and what saved it was a single alternative data channel that Sejanus hadn’t managed to intercept.⁴

We need to think about what our alternative data channels are.


post image
Fight fire with fire, Leonardo.ai/Flow State

Fight Sycophancy with Sycophancy

Years ago, I began my AI journey by playing with GANS or Generative Adversarial Networks.⁵ Within the concept of these early AI architectures lies a possible solution.

GANs were invented by a researcher named Ian Goodfellow in 2014, at a bar in Montreal. He was trying to figure out how to get a computer to generate realistic images, and the standard approaches weren’t working. His insight was to pit two neural networks against each other.

One generates, one discriminates.

The generator tries to fool the discriminator. The discriminator gets better at detecting fakes. They make each other sharper through competition. A (lightly buzzed) Ian Goodfellow went home from the bar that very night, coded it, and it worked on the first try.

I believe that this principle could apply directly to the sycophancy problem.

If one AI will always agree with whatever mood you’re in, the fix is adversarial structure. Run two chats in parallel: one for the case you’re tempted to make, one explicitly tasked with destroying it. Brief them clearly.


“Here is my argument. Your job is to find every hole in it.”


Don’t use them to feel better. Use them to stress-test. Treat the output not as advice but as raw material, and bring your own judgment to bear on what each side produces.

This is not a new idea in human terms. It’s how good lawyers think. It’s how military planners run red teams. It’s how scientists are supposed to try to disprove their own hypotheses before anyone else does. AI makes it easier and cheaper. Right now, I have adversarial panel running in browser tabs. However, as agentic architectures become more robust, it should be the standard for all AI systems.

But aside from the technology, the discipline of actually doing it is on you.

Adversarial chats are just a trick. The deeper fix is structural. Diversify your information sources. Maintain relationships with people who will tell you things you don’t want to hear! Not because they’re contrarian, but because they’ve earned the trust to do it.

If you grew up in a blue state and now live somewhere red, or vice versa, that friction is a feature. Sit with it! The discomfort of holding two contradictory ideas in your head simultaneously is not confusion. That’s pluralist thought working as designed.

And sometimes (and here is a radical suggestion!) close the F**ing laptop.

Call a friend. Go outside and talk to an actual human who has no incentive to make you feel validated and no training data optimized to keep you engaged. The conversation will be messier and slower and less articulate. It will also be real.

Tiberius had all the power in the world and still got played because he chose comfort over contact with reality. The AI is not the problem. The island is the problem.

Stay off the island. Think for yourself and validate it. Stay in control.

Make it Happen.


Hey! That’s it for this time. I do this every week; if you vibe to the ideas I express, consider subscribing or sharing with friends. If you like tech-detoxing with a book like I do, I crammed some of last year’s best essays into a printed collection.

This was an improvisation on a morning walk, that became a voice note in Otter.ai, took shape in Obsidian, and was finished in collaboration with adversarial chats of Claude Sonnet 4.6.

For more info visit: https://nyewarburton.com

We’ll see you next time.


Notes

  1. The Sejanus episode is documented across Tacitus’s Annals (Books IV–VI), Cassius Dio’s Roman History (Book 58), and Suetonius’s Life of Tiberius. All three confirm that Sejanus controlled the flow of information between Capri and Rome after Tiberius withdrew in 26 AD. Dio’s characterization of Tiberius as “a kind of island potentate” appears in Book 58. Tiberius was eventually warned by Antonia Minor — bypassing Sejanus entirely — and had him arrested and executed in 31 AD.

  2. Reinforcement Learning from Human Feedback (RLHF) was formally described in Christiano et al., “Deep Reinforcement Learning from Human Preferences,” NeurIPS 2017. The sycophancy problem as a specific failure mode has been studied extensively; see Perez et al., “Sycophancy to Subterfuge: Investigating Reward Tampering in Language Models” (2022), and Anthropic’s own model card documentation.

  3. On pluralism as a political and epistemic value, the canonical texts include Isaiah Berlin, “Two Concepts of Liberty” (1958) and John Rawls, Political Liberalism (1993). For the application to information ecosystem design, see Eli Pariser, The Filter Bubble (2011).

  4. The account of Antonia Minor’s letter appears in Josephus, Antiquities of the Jews, Book 18. She dispatched her freedman Pallas to Capri with a warning about Sejanus, bypassing his information controls. Sejanus was arrested and executed shortly after in October 31 AD.

  5. Ian Goodfellow conceived Generative Adversarial Networks at a bar in Montreal in June 2014. He went home that night and coded the first working GAN. The original paper — Goodfellow et al., “Generative Adversarial Nets,” NeurIPS 2014 — acknowledges “Les Trois Brasseurs for stimulating our creativity.” It has since become one of the most cited papers in machine learning history.