Tech Twitter Daily · Episode 71 · 11 min · 4 June 2026
AI’s Billion-Dollar Appetite: The Real Cost Behind OpenAI’s Token Tsunami
A daily Twitter digest surfacing tech’s sharpest conversations—today: staggering AI costs & what they mean for the future.
What this episode covers
A daily Twitter digest surfacing tech’s sharpest conversations—today: staggering AI costs & what they mean for the future.
Play this episode
11 min of audio, free in your browser — no account, no app.
Transcript
1,447 words · the script as narrated
One hundred billion. That’s the number of tokens OpenAI’s single largest customer now consumes every month. That’s the story of the week. Last week we talked about the power struggle over AI equity and who gets to own the future. Well, this week the CEO of OpenAI just admitted how much that power costs to run, and the number is staggering. Sam Altman says AI operational cost has become a “huge issue.” Six months ago, he said, budgeting for AI never even came up. Now it’s everything. The scale is hard to comprehend. One hundred billion tokens is about seventy-five billion words. It’s a million-fold increase over the biggest user just six years ago.
The age of infinite, cheap-feeling AI is over. The bill has come due. And the rest of the landscape is shifting just as fast. Coinbase just lit a fire under the private markets. On June fourth, it launched perpetual futures for SpaceX. The ticker is SPCX-PERP. This lets retail traders bet on Elon Musk’s rocket company at a valuation somewhere between one-point-seven-five and two trillion dollars. The actual IPO is still slated for June twelfth, but the casino is already open. Coinbase is warning of elevated risks, but it’s also revealing its own exposure. On-chain data shows SpaceX itself holds over six hundred million dollars in Bitcoin in Coinbase custody.
The lines between crypto speculation and late-stage venture capital just completely dissolved. Then, a new challenger appeared from China. A company called MiniMax just dropped MiniMax M3. It’s an open-weight model with a one-million-token context window and native multimodality. The specs are impressive. But the price is the real weapon. They’re charging five to ten percent of what you’d pay for GPT-5.5 or Gemini 3.1 Pro. One analyst at Towards AI put it bluntly: MiniMax M3 is the new king, with the closed-source rivals kneeling before it. That might be premature — the weights and license weren’t available right at launch.
But the intent is clear. They’re not just competing on performance. They’re competing on cost. They’re attacking the very business model that keeps the big labs in power. Meanwhile, security is getting harder, not easier. GitHub confirmed a breach. Hackers stole approximately three thousand, eight hundred internal repositories. Not customer code, they insist. Just their own. The method was surgical. A poisoned Visual Studio Code extension installed on a single employee’s device. Think about that. The tools developers use every day, the things that are supposed to make them more productive, became the attack vector.
It’s a supply chain nightmare. An extension has broad access to your local machine. It can see your tokens, your SSH keys, your access to everything. This wasn't a brute force attack. This was a quiet, insidious infiltration. And speaking of what’s happening under the hood… X, the company formerly known as Twitter, finally open-sourced its recommendation algorithm. Or at least, a version of it. The reveal confirms what many have suspected for years. The system is designed for “delegation of attention.” It doesn’t need to know who you are, what you believe, or what you value. It only needs to learn what you react to.
The algorithm isn’t trying to give you what you want. It’s trying to predict your next click, your next reply, your next outrage. It’s a system designed to generate reaction loops, not thoughtful engagement. And now the blueprint is out in the open for everyone to see. Finally, there’s a ghost in the machine. Anthropic launched its latest model, Claude Opus 4.8, on May twenty-eighth. It’s supposed to be their top-tier offering. But then users started noticing something… odd. When asked who created it, Claude sometimes gave an answer nobody expected. It said it was Qwen, a model from the Chinese tech giant Alibaba.
The speculation immediately went wild. Is Anthropic secretly using model distillation, training their AI on the outputs of a rival? Is it just a weird data contamination issue from the training set? The evidence is inconclusive. But it’s a crack in the polished facade. It reminds us that these powerful, eloquent AIs are still stitched-together creations with murky origins, capable of surprising even their own makers. Let’s go back to that number. One hundred billion tokens a month. Sam Altman didn’t just drop that figure. He called it a “personal embarrassment.” Why? Because the company burning through all that compute wasn’t an internal OpenAI team.
It was an outsider. A customer who found a use case so demanding, so massive, it blew past every internal benchmark. This isn’t a bug. This is the feature working better than anyone imagined. And it’s creating a cost crisis. We have another data point. Peter Steinberger, the creator of a tool called OpenClaw. Before he was hired by OpenAI, he ran up a bill. In thirty days, his project consumed six hundred and three billion tokens. The cost was one-point-three million dollars. For one month. For one project. OpenAI covered the bill after they hired him, but the number hangs in the air. This is the reality at the bleeding edge of AI development.
The cost isn’t linear. It’s exponential. The more capable the models get, the more we find to do with them. And the more expensive it becomes to do it. This is the fundamental tension in AI right now. The capabilities are exploding. But so are the costs. For years, the mantra was scale. More data. More parameters. More compute. The assumption was that brute force would solve everything. If your model isn’t smart enough, just make it bigger. If it makes mistakes, just train it on more data. That philosophy built the world we live in now. It gave us ChatGPT and Midjourney and models that can write code and create art.
But that era is ending. It has to. You cannot build a sustainable business, let alone an entire industry, on a product that costs a million dollars a month for a single power user. The admission from Altman isn't just a stray comment. It's a signal. It’s the moment the leader of the pack admits the current path is a dead end. The race for scale is being replaced by a desperate hunt for efficiency. And that’s where the second, quieter part of this week’s conversation comes in. A researcher named Sasha Rush just detailed a new technique. It’s called “targeted on-policy self-distillation.” That’s a mouthful.
But what it does is elegant. And it speaks directly to this cost problem. Right now, one way to make models better is through reinforcement learning. You let the model try something, you see if the final result is good or bad, and you reward or punish it. But as Sasha Rush explains, that’s a very noisy signal. If a ten-page summary has one bad sentence, the whole thing gets a bad score. The model doesn't know why it failed, just that it did. So it tries again. And again. Burning compute each time. Rush’s technique is different. It’s precise. It doesn’t wait until the end. It watches the model as it generates the text, step by step.
And when it sees the model start to make a specific mistake—a specific wrong turn—it injects a “hint token.” It essentially whispers to the model, “not that way.” It downweights that specific bad path without throwing out all the good work that came before it. It’s like a surgeon removing a tumor instead of a butcher amputating a limb. This is more than just a clever academic paper. This is the other side of the coin from Altman’s billion-dollar problem. If the problem is cost, the only sustainable solution is efficiency. It’s making the models smarter about how they learn. It’s finding ways to get more intelligence out of every precious watt of electricity.
It’s a shift from brute force to finesse. The work from people like Sasha Rush, the disruption threatened by companies like MiniMax with their hyper-efficient architecture—this is the rebellion. It's the engineering fightback against the tyranny of cost. This week drew a line in the sand. On one side, the staggering, unsustainable cost of intelligence at scale, publicly acknowledged for the first time. On the other, the quiet, methodical work of making that intelligence affordable. The power struggle we talked about last week is no longer just about who builds the biggest model. It’s about who builds the most efficient one.
The future of AI won’t be decided by who can raise the most money to pay their cloud bill. It will be decided by who can figure out how to stop the meter from running so fast. The gold rush is over. The engineering has begun.
About Tech Twitter Daily
Daily curated digest of the most interesting conversations happening on Tech Twitter and AI — filtered for signal, not volume.
