Straight answers about running your own AI.
Written to be useful whether or not you ever rent a grove. Every number is either a public specification, a stated formula, or yours.
Guides
What "private" means at each layer, and the three ways to get one, honestly compared.
How VRAM is really used, the runtimes that matter, quantisation without the mysticism, and a calculator.
The three build tiers, the running costs nobody lists, and a break-even calculator with your own numbers.
Estimated VRAM for popular open models at Q4 and Q8, 8k and 32k context, with the smallest card that fits.
GPU pages
The card most local-model advice is implicitly written for.
The value benchmark for local models, and it has been for years.
The first consumer card whose memory bandwidth changes what is comfortable.
More memory than its neighbours, and less bandwidth than a card two years older.
One page per card: every model tested against its memory, estimated tokens per second, how much context is left over, and an honest verdict. All 13 cards compared.
Model pages
One page per popular open model: memory at Q8, Q5 and Q4 for three context lengths, the smallest card that fits, and speed per GPU. Start from the table.
Tools
Your card, any model: fits or not, how fast, and what to change when it does not.
Weights + KV cache + overhead for any model, with the formula shown.
Your hours, your electricity, your plan. Finds the break-even month.
How fast a model runs on a given GPU, from bandwidth and weight size.