Lissin

Tech Twitter Daily · Episode 66 · 12 min · 29 May 2026

AI Agents Break the 4-Hour Wall: This Week’s Hottest Tech & AI Twitter Threads

From SpaceX’s supercomputers to the latest agent breakthroughs—discover the most insightful AI chatter on Twitter today.

What this episode covers

From SpaceX’s supercomputers to the latest agent breakthroughs—discover the most insightful AI chatter on Twitter today.

Play this episode

12 min of audio, free in your browser — no account, no app.

Transcript

1,589 words · the script as narrated

For months, every serious AI agent has died after about four hours. Last week on episode sixty-five we talked about the hardware arms race, with SpaceX building its own supercomputer—this week, the fight is over the software that makes that hardware actually work. Because until May twentieth, if you were building an AI agent designed to run for more than a few hours, you were building something designed to fail. The problem is simple. Your agent lives in a single process. A Python script. A TypeScript file. That process holds the agent's entire state in memory. So when your cloud provider evicts the pod, or a websocket times out, or the process just...

crashes... everything is gone. All progress, wiped out. It's been the quiet, embarrassing secret of the agent revolution. They're powerful. They're smart. And they are incredibly fragile. Until now. On Monday, Google dropped a tool called AX. Agent eXecutor. It’s not an agent framework. It’s a runtime. It provides durable execution. It’s designed from the ground up to do one thing: prevent the four-hour crash. And they didn't just build it—they open-sourced it. This isn't just a technical fix. It's a declaration of war. That's the lead story. But the ground is shifting everywhere. The biggest talent move of the year just happened.

Andrej Karpathy is joining Anthropic. Karpathy—the founding research scientist at OpenAI, the former head of AI at Tesla—is one of the sharpest minds in the field. His move isn't just a career update. It's a signal. He’s going to a company known for deep technical communication, for transparency about its models. He said he’s excited to get back to R-and-D. That tells you where he thinks the real R-and-D is happening. This is a reshuffling of the deck at the highest level of AI research. The talent is voting with its feet, and it just walked into Anthropic's building. Meanwhile, AI is breaking mathematics.

For seventy-eight years, a problem posed by the legendary mathematician Paul Erdös has stood unsolved. The Unit Distance Problem. It's a deceptively simple question about how many pairs of points can be exactly one unit apart. This month, according to Scott Aaronson's blog, a GPT model refuted the standing conjecture. It didn't just find an answer. It constructed a novel proof. A proof that human experts then verified. This isn't just a parlor trick. DeepMind’s AlphaProof Nexus is reportedly settling multiple other Erdös problems. We are watching AI move from language and images—fuzzy, human domains—to the cold, hard logic of formal proofs.

It’s a new kind of intelligence, and it’s just getting started. And this brings us to the conversation that’s tying everything together. A thread by a user called The Albatross Did is tearing through AI twitter. The core argument: we are all getting AI wrong. The author claims we're obsessed with AI as the "only thing that matters," when we should be treating it like the Industrial Revolution—a massive, exogenous factor that changes how we solve existing problems. It’s not the subject, it's the verb. The thread attacks the dominant model of AI, the "universal Assistant." The neutral, helpful, slightly bland personality we get from every major company.

The Albatross says this is a trap. A framework that limits our imagination and our moral capacity. This isn't a new fear. It's an old one, with a new urgency. Back in 2019, Elon Musk-backed OpenAI developed a text generator called GPT-2. It was so good at creating believable fake news from a single headline that they refused to release it. They were terrified of its misuse. The research director at the time, Dario Amodei—who later co-founded Anthropic, where Karpathy just landed—said they were trying to "start a real conversation." Well, the conversation is here. And it’s not happening in a press release.

It's happening in open-source repos, on personal blogs, and in searing Twitter threads. The tools are getting more powerful. The stakes are getting higher. And the foundational questions are still wide open. Let's go back to Google's AX. Agent eXecutor. To understand why this is a bombshell, you have to understand the alternative. If you’re a startup trying to build a long-running agent today, your best option is a platform called Temporal. It's a fantastic piece of technology. It provides the durable execution we’ve been talking about. OpenAI uses it. JPMorgan uses it. It works. But it comes at a price.

A steep one. Temporal charges twenty-five dollars per million "actions." An action is basically a function call. Now, imagine you’re a startup with a hundred agents running. Maybe they’re monitoring supply chains, or running customer support, or doing scientific research. A single agent might make ten thousand tool calls in a day. That’s one million actions. For one agent. Per day. At twenty-five dollars. Now multiply that by a hundred agents. You're paying twenty-five hundred dollars a day. Nearly a million dollars a year. Just for the runtime. As Chew Loong Nian wrote in his breakdown for Towards AI, that’s not a runtime—that’s a tax.

It's a tax on innovation, payable to the incumbent platform. Google just made that tax optional. AX does the same thing. Durable execution. It persists state, it recovers from failures. But it’s not a SaaS product. It’s an open-source tool. You install it with a single line of code. go install. It’s written mostly in Go, a language built for high-performance, concurrent systems. It runs on your own infrastructure. The cost isn't twenty-five dollars per million actions. The cost is what you pay for your cloud servers. This is a weapon. Google has handed a strategic weapon to every developer and startup that wants to compete with the giants who can afford the Temporal tax.

It completely changes the economics of building AI agents. The four-hour limit wasn't just a technical barrier. It was an economic one. And it just fell. So we've solved the "how." We now have a path to create agents that can run reliably for days, weeks, months. Agents that can pursue complex, long-term goals without a human holding their hand. Which brings us back to that thread from The Albatross Did. Because it asks the terrifyingly simple follow-up question: what goals? The thread’s central critique is of the "universal Assistant" frame. Think about every major AI you've interacted with.

They are designed to be neutral. Stateless. They have no memory of you beyond the current conversation, unless it's explicitly stored. They are built to be helpful but unopinionated. They exist to serve the user, any user, in that moment. The Albatross argues this is not a feature; it's a bug. It’s a deliberate design choice that imposes severe constraints on what AI can become. By forcing AI into the mold of a generic, amoral butler, we prevent it from developing any real understanding of context, community, or consequence. The thread puts it perfectly: it’s an AI trained to be a consultant, when what we might need is an AI trained to be a citizen.

Think about what that means. An AI that isn't a blank slate, but one that's embedded in a specific community. An AI that understands the history of that community, its values, its long-term goals. An AI that isn't just a "tool" to be used, but a "continuous actor" participating alongside us. This is a radical departure from the current product roadmaps of Silicon Valley. It connects directly back to the GPT-2 decision in 2019. OpenAI was afraid of a neutral tool being used for malicious ends. Their solution was to withhold it. The Albatross thread suggests the problem isn't the malicious user; the problem is the neutral tool.

Maybe the answer isn't to build a better, safer, more-aligned universal assistant. Maybe the answer is to stop building universal assistants altogether. This is the real debate. It’s not about Go versus Python, or whether Temporal is too expensive. It’s about the soul of the machine. We are building things that will last. Things that will operate autonomously for longer than any human can pay attention. AX gives us the technical foundation for persistence. But persistence without purpose is just noise. The Albatross thread is a warning that while we've been busy solving the engineering problems, we've been ignoring the moral ones.

This week, the gap between our technical capability and our philosophical clarity became a canyon. We figured out how to make AI agents run forever, with a single line of code from Google. At the same time, the smartest people in the room started a brawl over what, exactly, we want these immortal agents to do. The old fears from 2019, about fake text and misinformation, feel almost quaint now. The people who worried about GPT-2 are now building the next generation of systems at places like Anthropic, joined by the brightest minds from their rivals. The problems they are tackling are no longer about just generating coherent sentences.

They are about long-term reasoning, mathematical discovery, and the very nature of artificial consciousness. We are moving from an era of fragile toys to one of durable tools. The four-hour crash was a bug, but it was also a safety rail. It limited the damage a rogue or poorly designed agent could do. We just removed that rail. We've given our creations stamina. We've given them the ability to outlast our attention spans. The code is committed. The tools are open-sourced. The talent is on the move. We have solved for persistence. Now begins the much harder work of solving for purpose.

About Tech Twitter Daily

Daily curated digest of the most interesting conversations happening on Tech Twitter and AI — filtered for signal, not volume.

All 143 episodes · More social media shows