Lissin

Hacker News Daily · Episode 57 · 11 min · 20 May 2026

Hacker News Daily Digest: Top Tech Stories & Buzz, Curated for You

Get the sharpest takes on AI, tech trends, and lively discussions—just the highlights, no noise.

What this episode covers

Dive into the essential tech news and discussions with our daily Hacker News Digest. We sift through countless threads to bring you the top stories, compelling insights, and community buzz that truly matter, saving you valuable time. Tune in to stay effortlessly informed on the ideas worth chewing on, curated by someone who's read every thread so you don't have to.

Play this episode

11 min of audio, free in your browser — no account, no app.

Transcript

1,630 words · the script as narrated

Google's Gemini 3.5 model is estimated to have only ten to sixteen billion active parameters. Last week we were talking about the Vatican weighing in on AI ethics, which felt like the absolute peak of the culture conversation, but this week we are back on the ground, looking at the metal. Because that number—ten to sixteen billion—tells you almost everything you need to know about where the AI industry is headed in May of 2026. And it’s not the only story. The Hacker News digest today is a snapshot of an industry in transition. First up, there's the big picture from Benedict Evans. His latest presentation argues that the era of AI models as the main event is over.

We’re in mid-2026, and the models themselves are becoming infrastructure. Just plumbing. The real value, he says, is moving up the stack into the apps, the workflows, and the proprietary data you feed into them. The gold rush isn't for the biggest model anymore; it's for the smartest application. Then you have GitHub, which is having a very bad, very weird week. They announced—only on X.com, which is a choice—that they're investigating unauthorized access to their own internal code repositories. The community is, predictably, on fire. Not just about the security breach itself, but about the bizarrely limited communication. And the speculation about the motive is even wilder.

Some are asking... what if the attackers weren't trying to break things, but to fix them? It sounds insane, but stick with me. And that brings us back to Google's Gemini. A deep technical analysis posted on Hacker News is tearing through the community, using hardware specs and efficiency reports to reverse-engineer the model's size. The conclusion? Those rumors of multi-trillion parameter models from the big labs? They might be pure fantasy. The analysis pegs Gemini 3.5 at around two hundred fifty billion total parameters, with only a tiny fraction of that—ten to sixteen billion—active at any one time. It's built for speed, running at around two hundred eighty tokens per second on Google’s own hardware.

This isn't a monster truck; it's a Formula One car. Finally, there's the user sentiment on the ground. People using these new models in the wild are reporting a consistent pattern. For a single, specific coding problem? Gemini is a genius. Almost frontier-level. But ask it to handle a complex, multi-step project—what they call a long-horizon agentic task—and it just falls apart. It can't plan. It can't iterate. It’s a brilliant intern who needs a manager standing over their shoulder for every single task. So you have these four threads: AI is becoming plumbing. GitHub, a core piece of that plumbing, is springing leaks. The new models are smaller and faster than we thought.

And they're still not capable of doing real, autonomous work. You see the pattern here? The hype is finally meeting reality. Okay, let's dive deeper into that big idea from Benedict Evans. His thesis is that AI is just another platform shift. And if you want to know where we're going, you just have to look at where we've been. Think about the early internet. In 1998, just having a website was a huge deal. It was a technical marvel. The people who could build one were wizards. The value was in the infrastructure itself. Fast forward a decade. We had WordPress. We had Squarespace. Suddenly, anyone could have a website. The value wasn't in having the site anymore; it was in what you did with it.

Was it a store? A publication? A community? The value moved up the stack. Or think about cloud computing. Before AWS, if you wanted to launch a startup, you had to spend a fortune on servers. You had to buy them, rack them, manage them. It was a massive capital expense. Amazon came along and turned servers into a utility, like electricity. You just plug in and pay for what you use. The value wasn't in owning servers anymore. It was in the software you built to run on someone else's servers. The value moved up the stack. Evans is arguing that the exact same thing is happening with AI models right now. For the last two years, the story has been the models themselves.

GPT-4, Claude 3, Gemini 1.5. Who has more parameters? Who has a bigger context window? It was a spectacle. A space race measured in trillions of parameters and benchmark scores. Now, he says, that's over. The models are becoming commodities. Every major cloud provider offers you a dozen of them through an API. The ability to generate text or images is becoming... well, plumbing. It’s the new AWS. And if he's right, the question "Who is building the best AI model?" is the wrong question. The right question is "Who is building the best product using an AI model?" As one commenter on the thread put it, "A chat bot is barely a product." And outside of coding assistants, which are genuinely useful, the big AI labs haven't been very good at building actual products.

They're great at building engines, but they don't know how to build a car. This is where the analogy holds perfectly. The people who got rich in the dot-com era weren't just the ones laying fiber optic cable. It was the Amazons and the Googles who built world-changing businesses on top of that cable. The real winners of the AI era won't be the ones with the biggest model. They'll be the ones who use a good-enough model to solve a real-world business problem in a way that nobody else has. The value is migrating. It’s flowing away from the raw technology and toward the finished product. So if that's the big picture, let's zoom in on that Gemini 3.5 analysis.

Because it's the perfect piece of evidence for Evans's theory. It's the "show, don't tell" of the AI infrastructure phase. For months, the rumor mill has been churning out these insane numbers. GPT-5 is gonna be five trillion parameters! Opus 4.7 will be ten trillion! It created this narrative that progress means making these models bigger, and bigger, and bigger. Just add more layers, more data, more compute. Then this analysis drops. And it says Google's latest, greatest model... is actually pretty small. At least the active part. Ten to sixteen billion active parameters. To put that in perspective, that’s in the same ballpark as models that were state-of-the-art two or three years ago.

So what gives? Did Google fall behind? No. That's the whole point. They're not playing the "biggest model" game anymore. They're playing the "most efficient model" game. The analysis suggests the model is a Mixture of Experts, or MoE, where you have a huge library of "experts"—the 250 billion total parameters—but for any given task, you only call on a small, specialized team of them. It's optimized to run on their specific hardware, the TPU 8i. It's using mixed precision, which is a fancy way of saying it's using just enough detail to get the job done and no more. This isn't a research experiment anymore. This is a product designed for mass-market deployment.

It's engineering. The historical pattern here isn't the space race. It's the history of the automobile engine. In the beginning, it was all about raw power. Big block V8s. Displacement was king. But what happened over time? Engineers started chasing efficiency. Turbochargers, fuel injection, variable valve timing. They figured out how to get more power out of smaller, lighter, more fuel-efficient engines. A modern four-cylinder engine in a Honda Civic is an engineering marvel that would run circles around most old V8s in terms of efficiency and reliability. Google isn't building a dragster anymore. They're building the Honda Civic engine of AI. And there's a much, much bigger market for Honda Civics than there is for dragsters.

This model is designed to be cheap to run at incredible scale. It's infrastructure. But here's where the pattern-matching gets tricky. Here's where the analogy starts to break down. If the model is so cheap and efficient to run, why is Google charging so much for it? Commenters are pointing out that inference costs for Gemini 3.5 are way higher than for similarly-sized open-source models. What's going on? Are they just trying to claw back their massive training costs? Or is there a compute shortage so severe that even Google can't meet demand? This is the part of the story that doesn't make sense yet. It's the piece of the puzzle that's still missing. The engine is efficient, but the price at the pump is sky-high.

And nobody can quite figure out why. So this week, the spectacle of the AI race finally gave way to the mundane reality of building the plumbing. The shift from "how big is your model" to "how smart is your workflow" is happening right now, in real time. We’re seeing it in the strategic debates about value, and we’re seeing it in the nanometer-level engineering of the models themselves. Even the GitHub story, in its own strange way, fits this new world. The platform has become so fundamental, so much a piece of the core infrastructure for every tech company on the planet, that one of the wild theories is that the breach was caused by an AI agent tasked with improving uptime—that went a little too far.

That's probably not what happened. But the fact that it's even a plausible-sounding theory tells you everything about how deeply this stuff is now embedded in our world. The age of AI as a magical, mysterious force is ending. The age of AI as a utility—powerful, essential, and a little bit boring—has begun. And the work of building on top of it has only just started.

About Hacker News Daily

Daily digest of the best Hacker News stories and discussions — the ideas worth chewing on, filtered by someone who reads every thread.

All 155 episodes · More tech & startups shows