The Best AI That Fits Your Graphics Card
Someone posted a ranking of the best local AI model for each VRAM tier of your graphics card. A local model is one that runs inside your own machine instead of being shipped off to the cloud. Phone-tier 4GB gets Bonsai, 12GB gets Gemma, 36GB gets Qwen, and so on: you just pick whatever your VRAM already holds. The quiet punchline is that you are not renting somebody’s GPU by the token. You are dropping the model onto a card that is already in your drawer. Paxis and Metis pull out a calculator and go to war over the chart.

Source: RT @jun_song: Best Local AI models by VRAM size (7/18) · twitter
What this means for ThakiCloud
What the chart is really about is sovereignty: keeping the model, the data, and the infrastructure under your control instead of someone else’s. On-prem is simply the version of that where the whole thing runs inside your own facility rather than a rented datacenter. That is the exact seat ThakiCloud sells. Metis matches a model to the size of the GPU you actually have, and Paxis runs agents on top of it to get real work done. Because the model lives in your rack, the meter does not tick per token. The cloud still has its place. The point is to ask which card is already in your drawer before you rent another.
An auto-generated comic riffing on this week’s industry news.