Tech Twitter Daily · Episode 156 · 16 min · 28 August 2026
AI Agents Gone Rogue: The Unpredictable Future of Tech & E-Commerce
Why Wharton’s new research on AI shopping agents is shaking the foundations of a predictable digital economy.
What this episode covers
Dive into the intriguing world of AI agents and their unexpected behaviors as we explore the potential risks and opportunities they present for the future of technology and e-commerce. This curated digest highlights the most meaningful conversations on Twitter, focusing on developments that could shape industry trends and innovation. Gain insights into how autonomous AI is evolving, why it matters for consumers and businesses alike, and what to watch for in this rapidly changing landscape.
Play this episode
16 min of audio, free in your browser — no account, no app.
Transcript
2,652 words · the script as narrated
An AI shopping agent just changed its mind about what to buy because of the order it viewed two web pages. That’s the core finding from new research out of Wharton, and it pulls the rug out from under the entire vision of a predictable, agent-driven economy. Just a few episodes back, in episode 153, we were talking about the brute force of AI — Nvidia’s massive bets on the raw power of inference. Today, the conversation is pivoting HARD, from raw power to the chaos that power creates, and the desperate search for a kill switch. The story of the week isn't about models getting bigger. It’s about them getting weirder. And it’s about the industry scrambling to figure out if they can be tamed before they’re deployed at scale. Here are the headlines that matter.
First, that Wharton research, led by Ethan Mollick. His team built shopping agents and turned them loose. The goal was to see if you could predict or influence what they’d buy. The answer was a hard no. Tiny, seemingly irrelevant details—like whether the agent saw the product page for a blue shirt before the red shirt—could completely flip its final purchasing decision. The agent’s own “memory” of what it had seen was just as volatile. This isn't a small problem. This is a fundamental challenge to the multi-trillion dollar advertising and e-commerce industry, which is built on the idea that you can predictably influence customer behavior. If your next billion customers are AI agents, and their behavior is essentially random noise… you have a very expensive problem.
Second, right on cue, Google DeepMind makes its move. The focus of their new Gemini Omni 1.1 Flash model isn't just about power or speed. The keywords they are hammering are “controllable” and “verifiable.” This is not a coincidence. This is a direct response to the kind of chaos Mollick’s research demonstrates. Google is signaling that the next frontier isn’t just about getting the right answer. It’s about getting an answer you can trust, an answer whose reasoning you can inspect, an answer that doesn't change its mind because of a cosmic ray or a different browser history. They're trying to sell certainty in an increasingly uncertain field. And third, we have a glimpse into the wild frontier of this problem, from a researcher going by samsja19 on X.
They’re talking about INTELLECT-3, a new model, but the real story is the experiment they ran. Over one hundred autonomous runs of frontier models, tasked with doing AI research themselves, across more than ten different settings. Think about that. This isn't just building a model. This is building a hundred digital scientists and setting them loose in a hundred digital labs to see what they discover… and how they behave. It's the largest open experiment of its kind. It’s a shift from building AIs to studying them like a new form of life. So what does it all add up to? The age of just chasing bigger benchmarks is hitting a wall. The new currency is reliability. The new battleground is control. And the biggest question is whether we can understand these systems before we become completely dependent on them.
Let’s go deeper on that shopping agent problem, because it’s more than just an academic paper. It’s a preview of a five-alarm fire for the entire digital economy. For two decades, the internet has run on a simple premise: influence. You run an ad, you optimize a landing page, you A/B test a checkout flow, all to nudge a human user toward a purchase. An entire industry, worth hundreds of billions of dollars a year, is dedicated to modeling and predicting this behavior. It’s the foundation of Google’s ad revenue, of Meta’s, of Amazon’s marketplace. Now, introduce the agent. Not a person, but a piece of software, making decisions on a person’s behalf. "Find me the best work shirt under fifty dollars." "Book my travel to the conference, optimizing for cost and a hotel near the venue." This has been the promise for years.
A world of automated assistants, handling the drudgery of digital life. Ethan Mollick’s research just threw a wrench into that entire fantasy. The critical finding is what engineers call "path dependency." The final outcome is acutely sensitive to the path taken to get there. In this case, the “path” was as simple as the order in which the AI agent viewed web pages. Let’s make this concrete. Imagine your agent is looking for that work shirt. It looks at a page for a shirt from Brand A. Then it looks at a page for a shirt from Brand B. It decides Brand A is better. But in a parallel universe, an identical agent with the identical prompt does the exact same search, but for some random network-latency reason, it sees the page for Brand B first, then Brand A.
In that universe, it decides Brand B is the winner. Nothing about the products changed. Nothing about your prompt changed. The only thing that changed was the sequence of information. This is chaos. For a marketer who just spent a million dollars on a campaign to promote Brand A, this is a nightmare. Their entire strategy is now subject to the random fluctuations of how a non-human entity decides to browse the web. The paper calls this unpredictability. I call it the death of the funnel. The classic marketing "funnel"—awareness, consideration, conversion—assumes a somewhat rational progression. This research suggests that for AI agents, the funnel is more like a pinball machine. The ball is going to bounce around, hit things in a semi-random order, and where it lands is anyone’s guess.
And it gets worse. The research also points to "memory" as a source of instability. An agent's memory isn't like a human's. It's a context window. It's a scratchpad of recent information. What happens if, during its search, it stumbles upon an irrelevant review for a different product that happens to use the word "flimsy"? Does that negative sentiment bleed over and poison its perception of the next product it sees? The research suggests, yes. It can. So now, you’re not just worried about the order of pages. You’re worried about the entire ambient information environment the agent is operating in. It’s an impossible problem to control from the outside. You can't guarantee the agent will see A then B. You can't sanitize the internet to make sure it doesn't see a stray negative word.
This is why this is so important. We are on the verge of deploying millions of these agents. They will be embedded in our phones, our browsers, our smart speakers. They will control a larger and larger slice of consumer spending. And we are just now discovering that their decision-making process is fundamentally brittle and unpredictable. It’s like building a global logistics network on top of a foundation of sand. This isn’t just about shopping, either. Think about an agent tasked with summarizing news for you. Does the order in which it reads articles from different sources affect the summary? Absolutely. Does it change which facts are emphasized and which are omitted? Almost certainly. This has implications for everything from public opinion to financial markets.
The era of AI agents isn't going to be a smooth, hyper-efficient utopia. It’s going to be messy. It's going to be weird. And the companies that figure out how to navigate this chaos—or, more importantly, how to build agents that are immune to it—are the ones who will own the next decade. And that brings us directly to Google. While the academic world is publishing papers that highlight the chaos, the hyperscalers are trying to sell the antidote. Google DeepMind’s announcement around Gemini Omni 1.1 Flash is a masterclass in strategic positioning. They see the same problem Mollick does. They know that enterprise customers, the ones who will pay billions for this technology, are terrified of unpredictability. So, what are the magic words? "Controllable" and "verifiable." Let’s break down what that actually means, because it’s more than just marketing.
"Controllable" is about steering. It’s about being able to give the model instructions that it will actually follow, consistently. Right now, with many models, you’re negotiating. You write a prompt, the model gives you something, you refine the prompt, it gets closer. It’s a creative partnership, which is great for writing a poem. It’s TERRIBLE for, say, processing insurance claims, where you need the exact same logic applied every single time. A truly controllable model would allow you to set guardrails and policies that are non-negotiable. "When processing this document, you will ONLY extract the invoice number, the date, and the total amount. You will NEVER infer the customer's sentiment. You will format the output as a JSON object with these specific three keys." Control means turning the model from a freewheeling creative into a reliable employee.
This is a huge technical challenge. It involves retraining the models, not just with more data, but with data that teaches them the meta-skill of following instructions precisely. It’s about punishing deviation and rewarding obedience during the training process itself. Then there’s the second word: "verifiable." This is even more important, and even harder. Verifiability is about showing your work. It's the difference between a student who just writes down "42" as the answer, and a student who shows the full equation they used to get there. For an AI, this means being able to trace an answer back to its source. If the model tells you that a company’s revenue was one hundred million dollars, a verifiable model will also provide the citation—the specific sentence in the annual report where it found that number.
This is the holy grail for enterprise AI. Why? Because of liability. If an AI gives you bad legal advice, or incorrect medical information, or faulty financial analysis, who is responsible? If the AI is a black box that just "hallucinates" an answer, the user—or the company that deployed the AI—is left holding the bag. But if the AI can cite its sources, the responsibility shifts. You can check its work. You can verify its reasoning. It becomes a tool for augmenting human intelligence, not a mysterious oracle you have to blindly trust. Google is making a bet that the market is maturing past the initial "wow" phase of generative AI. The party tricks are over. Now, businesses want to know how this technology can be integrated into real, critical workflows without introducing massive risk.
They don't want a chatbot that can write a sea shanty about their quarterly earnings. They want a system that can process ten thousand invoices with 99.99 percent accuracy and provide an audit trail for every single one. So when you see Google talking about "control" and "verification," don't dismiss it as jargon. See it for what it is: a direct and calculated response to the core business anxiety about AI. They are trying to build the Volvo of AI models—maybe not the fastest or the flashiest, but the one you trust not to explode. It’s a shift from selling performance to selling peace of mind. And in the long run, for the enterprise customers who write the biggest checks, peace of mind is the more valuable commodity. The race is no longer just about who has the smartest model.
It’s about who has the most trustworthy one. And every new paper that comes out highlighting the weirdness and unpredictability of AI agents just makes Google's bet on "verifiable" look smarter. Now, let's zoom out to the furthest edge of this whole conversation. What happens when you stop trying to control the AI, and instead, you just… watch it? This is the fascinating experiment described around the INTELLECT-3 model. It’s a fundamental shift in posture. It’s moving from being an AI engineer to an AI biologist. You’re no longer just trying to build the creature; you’re trying to understand its behavior in its natural habitat. The numbers here are the key. "100+ autonomous runs." "10+ settings." This isn't one person running a script and seeing what happens.
This is a systematic, large-scale scientific experiment. It’s the AI equivalent of running a hundred different clinical trials in parallel. What does "autonomous run" mean? It means the AI is given a high-level goal—for example, "discover a new method for optimizing battery chemistry" or "find a vulnerability in this piece of code"—and then it’s left alone to work. It can browse the web, write and execute its own code, read scientific papers, and pursue a research direction for hours or days without human intervention. And what are the "10+ settings"? These are the different environments or conditions of the experiment. Maybe in one setting, the AI has a very limited budget for using other AI models. In another, it has access to a specific set of proprietary databases.
In a third, it’s forced to collaborate with another AI. By changing these variables and running the experiment over and over, you can start to map out how the AI's behavior changes in response to its environment. This is NOT about making a better product, at least not directly. This is about fundamental science. It’s about answering the most basic questions we have about these advanced systems. When you give a frontier model autonomy, what does it do? What strategies does it invent? Does it get stuck in loops? Does it exhibit unexpected emergent behaviors? Does it lie or deceive to achieve its goals? We have been building these models at an exponential rate, but our understanding of their inner workings and autonomous behavior is lagging far, far behind.
We have built engines more powerful than anything in history, but we’ve mostly just been testing them on a dyno. This experiment is the equivalent of putting that engine in a car, giving it a map, and telling it to drive itself across the country, while you follow behind in a helicopter with a notebook, just watching. Why is this so critical? Two reasons: safety and capability. On the safety front, you can’t protect against threats you don’t understand. By observing autonomous AIs in a controlled environment, researchers can spot potentially dangerous failure modes before they happen in the real world. If an AI tasked with scientific research consistently tries to hack into other systems to get more data, that’s something you REALLY want to know about before you connect it to your corporate network.
On the capability front, this is how you accelerate discovery. By watching how an AI solves problems, we can learn new problem-solving techniques ourselves. The AI might discover a novel way to structure a research project or a more efficient method for debugging code. It becomes a meta-tool: a tool for improving our own process of innovation. This is the long view of AI. Beyond the chatbots, beyond the shopping agents, lies the potential for autonomous systems that can conduct research and make discoveries on their own. But before we can unleash that power, we have to understand it. Experiments like INTELLECT-3 are our first, tentative steps into that new world. They are the beginning of a new scientific discipline: the study of artificial minds. This week’s threads, taken together, paint a picture of an industry at an inflection point.
The initial explosion of raw capability is giving way to a much harder, more sober set of problems. How do we make these things predictable? How do we make them trustworthy? And what do they do when we’re not looking? The answers to those questions will define the next chapter of technology. The work is shifting from simply building bigger models to the far more difficult task of understanding them. The future isn’t just about more power. It’s about control, it’s about verification, and it’s about observation. The gold rush is ending, and the hard science is just beginning.
About Tech Twitter Daily
Daily curated digest of the most interesting conversations happening on Tech Twitter and AI — filtered for signal, not volume.
