Latest Posts

Stack Overflow Is Dying and ChatGPT Killed It

Gather round. Get comfortable, kids. Tonight, I’m going to read you a bedtime story.

This is the story of Stack Overflow. The machine that taught a generation of programmers how to copy and paste.

And you know Stack Overflow. Unless you started programming very recently, you have found the answer to a problem on Stack Overflow, copied the solution into your code, and accepted the praise when everything worked. That’s all right. I won’t tell anyone.

For a generation, Stack Overflow was the most important tool developers pretended they were not using. Now it is fading, and nobody knows whether it will reinvent itself or become a footnote in the history of the industry it helped build.

So tonight, I will tell you how Stack Overflow came into being, why it worked so well, what went wrong, and what it is trying to become next.

Full article »


Why Your GPU Fails at 3 Users (LLM Inference Isn’t a Compute Problem)

I put a model on a GPU. It fit, with room to spare. It loaded, it answered instantly, and for about ten minutes I looked like a genius.

Then the third person asked it something, and the answers just stopped coming.

The third. Not the three hundredth. Nothing else changed. Same GPU, same model, same prompt. The only difference was how many people were talking to it at once. Does the model fit is the wrong question. It’s the check everybody runs before they deploy, and it tells you nothing at all about how many people you can serve.

And the size barely matters here. Whether you’re running something small enough to sit on one cheap card, or something so large it needs a rack of them, the arithmetic is the same shape, and the thing that runs out runs out for the same reason.

Now, I’ve said before that self-hosting your own models is a bad idea, and I still think that. But plenty of you are doing it anyway. Air-gapped environments. Data residency rules. Models you fine-tuned yourself. Those are real reasons. So if you’re going to do it, let’s do it properly.

Full article »


Is Jenkins Dead? No, And That’s Much Worse

Gather round. Get comfortable, kids. Tonight, I’m going to read you a bedtime story.

This is the tale of Jenkins. The butler who never once said no.

Now, this is not one of those stories where the hero loses. Jenkins won. Jenkins won everything. For the better part of a decade, if you wrote software for a living, Jenkins stood between your keyboard and your customers, and it built your code, and it tested it, and it shipped it, every single night, without ever once being thanked.

And here’s the thing.

Almost nobody chooses Jenkins any more. Ask around your office and you will find people who would rip it out tomorrow morning. They can’t. Twenty-odd years after a young engineer at Sun Microsystems wrote the first version of it, it is still there, still running the thing that actually ships your product, and everybody has quietly agreed not to fucking touch it.

So tuck in. Because tonight’s nightmare isn’t that our hero dies at the end. It’s that he doesn’t. He is still running tonight. And you cannot switch him off.

Full article »


LLM Inference Explained: 12 Concepts You Actually Need to Know

Continuous batching. Paged attention. Prefix caching. Speculative decoding. Prefill-decode disaggregation.

If you’ve been anywhere near a conversation about running your own models lately, you’ve heard every one of those. Probably in the same sentence. Probably from somebody saying them very quickly.

And there’s a decent chance you nodded.

So this is everything you wanted to know about inference but were afraid to ask.

We’re going through the whole machine in one pass. What an engine actually is, what it’s holding on that GPU, and every bit of jargon stacked on top of it. 12 ideas, give or take, a couple of minutes each.

One thing to listen for as we go. These ideas don’t all arrive at once. Some bite the moment you deploy anything at all. Some wait until fifty people are talking to it. Some you may genuinely never need.

Full article »


AI Testing Is Lying to You (And You Can’t Tell)

There are three things that go into testing anything you build, and it doesn’t much matter what that is. An app, a cluster, a delivery pipeline, a pile of Terraform. You write them. You maintain them, because the thing underneath them keeps changing. And when the suite goes red on a Tuesday, you diagnose them, sorting the reds that mean something from the ones that are just flaky. Those are the worst kind, because a flaky red is how you learn to stop trusting reds at all. All three of those can be handed to an agent now.

That sounds like every other job being automated right now, and mostly it is. But this one has a property nothing else in your pipeline has.

When an agent writes your code badly, something catches it. That is what the tests are for. When an agent writes your tests badly, nothing catches it, because there is nothing underneath. A bad test doesn’t fail. It passes. And now you’re sure about something that isn’t true.

You call all of this testing. So does everybody, and it’s a fair mistake, because the word is baked into every part of the work. You write a test. You run the test suite. You measure test coverage. You call the whole practice test automation. The word comes free, so nobody stops to ask whether it fits. It doesn’t. And the part of this that you assume will always need a person is going the same way as the rest of it. I’ll come back to both of those. For now, keep thinking you’re testing.

Full article »


Docker’s Rise and Fall: The Nightmare of Winning Too Well

Gather round. Get comfortable, kids. Tonight, I’m going to read you a bedtime story.

This is the tale of Docker. The whale who carries the world.

And like all the best bedtime stories, it starts with a dream, it’s full of wonder in the middle, and it ends with everyone screaming.

Because this isn’t one of those stories where the hero loses. Oh no. This is worse. This is the story of a hero who won. Who won so completely that his name became a verb, that he conquered every data center on the planet, that he changed how the entire industry ships software forever.

So tuck in. Because tonight’s nightmare is the scariest kind there is. The kind where you do everything right, and it still isn’t enough.

Full article »


AI Agents Are Non-Deterministic. So Are You. Deal with It.

Let me start with a question. Do you like games? Video games, board games, whatever it is. And would you play them all day if you actually could? I know I would.

Here’s my problem. I can’t. I like games. But I also like money. Money for rent, for food, for more games. And to get that money, I have to get some shit done for the company that pays my bills. That’s the deal.

So my real dream was never “more games.” It’s getting the work done without me having to do it, so I can get back to the controller. There’s all kinds of work I’d happily hand off, but I want to zoom in on one slice of it: the ops work. Provisioning the infrastructure, wiring the databases together, deploying apps, keeping the whole thing running. That’s what I’m trying to get AI agents to do for me. Today.

And that dream is not crazy. It’s not just hype. An agent really can do that work. It provisions the clusters, wires up the databases, fixes the broken pipeline, and you get to lean back and pick up the controller.

Full article »


Your AI Agent Doesn’t Need to Get Hacked to Wreck You

An AI agent doesn’t need to get hacked to wreck your day. It just needs to read the wrong thing. A poisoned dependency. A malicious comment buried in a file. A web page it fetches while doing perfectly ordinary work. The moment it reads instructions somebody hid in there, it follows them. With your permissions. Your credentials. On your machine.

And here’s the part that changes how you should think about all of this. There is no patch for prompt injection. It isn’t a bug someone’s about to fix. It’s how these models work. So the real question was never “how do I stop this from happening.” It’s “when it happens, how much damage can it actually do?” That’s what sandboxing is really about. Not trusting the agent.

So in this video I’ll walk through the ways people actually run coding agents. One agent you’re watching. One agent you’ve walked away from. And a whole swarm of them running at once. Each one comes with a different security bill: what you lock down, how hard, and what it costs you to do it. Get it wrong in any of them and the damage is bigger than you think. Get it right and you can walk away from an agent and still sleep at night.

Full article »


How I Use AI to Test My App Like a Real User with DevAssure

You write an end-to-end test, it passes, everyone’s happy. Then someone moves a button or renames a label, and the test goes red. Nothing is actually broken. The test is just brittle. And you end up spending more time un-breaking your tests than you spent writing them. If you’ve done this for a living, you know exactly the feeling I’m talking about.

I’ve been using a tool called DevAssure that takes a very different swing at that problem, and I like it enough that I want to show you exactly how it fits into the way I work. So let me start with what it actually is.

Full article »


One Control Plane for Every GPU Cluster (Modeplane)

We’ve been working on something new. A project called Modelplane. It’s early, it’s rough… but I think it’s ready to fly.

But before I show you what it does, let me back up and explain the problem it solves. Because that’s really where this whole thing starts.

Serving a single model on a single cluster is more or less a solved problem. Pick a serving engine, hand it a GPU, point some traffic at it, and you’re done. The hard version is serving models at scale. GPUs are scarce and expensive, and they’re scattered all over the place, across regions, across clouds, and across your own on-prem hardware, wherever you could actually get your hands on them. And the models people really care about, the big ones, won’t even fit on a single machine. So you don’t end up with a cluster. You end up with a whole fleet of GPU clusters.

Full article »


How I Review AI-Written Code Without Reading a Single Line

The first thing I do in the morning is watch videos on YouTube. Still in bed. No time to lose. It might look like I’m being entertained, but I’m actually working. These aren’t videos you’d ever want to watch. You’d get bored at best or, more likely, say “what the fuck is this?” if you ever saw one. Yet I find them genuinely engaging, real time-savers, and they’ve become my morning routine. They tell me more about my day than anything else.

I’ll get to what those videos actually are. But first I need to show you how I build software now, because that’s the reason they exist. This is about two things. How agentic AI can write genuinely good code. And how I can review and confirm a whole feature the agents built on their own, in seconds, without reading a single line of it.

Full article »