<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>One Control Plane for Every GPU Cluster (Modeplane)</title><link>https://devopstoolkit.live/infrastructure-as-code/one-control-plane-for-every-gpu-cluster-modeplane/index.html</link><description>We’ve been working on something new. A project called Modelplane. It’s early, it’s rough… but I think it’s ready to fly.
But before I show you what it does, let me back up and explain the problem it solves. Because that’s really where this whole thing starts.
Serving a single model on a single cluster is more or less a solved problem. Pick a serving engine, hand it a GPU, point some traffic at it, and you’re done. The hard version is serving models at scale. GPUs are scarce and expensive, and they’re scattered all over the place, across regions, across clouds, and across your own on-prem hardware, wherever you could actually get your hands on them. And the models people really care about, the big ones, won’t even fit on a single machine. So you don’t end up with a cluster. You end up with a whole fleet of GPU clusters.</description><generator>Hugo</generator><language>en-us</language><lastBuildDate/><atom:link href="https://devopstoolkit.live/infrastructure-as-code/one-control-plane-for-every-gpu-cluster-modeplane/index.xml" rel="self" type="application/rss+xml"/></channel></rss>