Ask your assistant whether a model runs on your card and it answers with the arithmetic behind every figure on this site: the same formulas as the checker, over the same 29 models and 20 GPUs, and over any other model on Hugging Face, read from its config.json. It is free, read-only and needs no account or key.
Connect
Remote, with nothing to install. Add this URL to any client that speaks Streamable HTTP:
https://mcp.nodegrove.io/mcp Claude (claude.ai and the desktop app)
Customize → Connectors → Add → Add custom connector. Name it Nodegrove VRAM, paste the URL, choose No sign-in, and turn it on in a chat from + → Connectors. Every plan includes at least one custom connector.
Claude Code
One command (add --scope user to have it in every project):
claude mcp add --transport http nodegrove-vram https://mcp.nodegrove.io/mcp Cursor
Use the install link, or add this to ~/.cursor/mcp.json:
{
"mcpServers": {
"nodegrove-vram": {
"url": "https://mcp.nodegrove.io/mcp"
}
}
} VS Code
Use the install link, or add this to .vscode/mcp.json:
{
"servers": {
"nodegrove-vram": {
"type": "http",
"url": "https://mcp.nodegrove.io/mcp"
}
}
} Devin Desktop (formerly Windsurf)
One command, or the same entry in mcp_config.json under mcpServers with a url field:
devin mcp add -s user nodegrove-vram https://mcp.nodegrove.io/mcp LM Studio
Use the install link, or Program → Install → Edit mcp.json:
{
"mcpServers": {
"nodegrove-vram": {
"url": "https://mcp.nodegrove.io/mcp"
}
}
} Open WebUI
Admin Settings → Integrations → External Tool Servers → Add Connection. Type MCP (Streamable HTTP), the URL, authentication None.
Cline
Add this to the MCP settings file. Without the type, Cline treats a URL as the older SSE transport:
{
"mcpServers": {
"nodegrove-vram": {
"type": "streamableHttp",
"url": "https://mcp.nodegrove.io/mcp"
}
}
} Codex CLI
One command, or a [mcp_servers.nodegrove-vram] table with this url in ~/.codex/config.toml:
codex mcp add nodegrove-vram --url https://mcp.nodegrove.io/mcp Gemini CLI
One command:
gemini mcp add -s user --transport http nodegrove-vram https://mcp.nodegrove.io/mcp Zed
Add this to settings.json:
{
"context_servers": {
"nodegrove-vram": {
"url": "https://mcp.nodegrove.io/mcp"
}
}
} ChatGPT
Settings → Security and login → turn on Developer mode. Then at chatgpt.com/plugins, add one with the URL and No Authentication, and pick it in a chat from + → Developer mode.
Or run it on your own machine over stdio. It needs Node.js 20 or newer and nothing else:
{
"mcpServers": {
"nodegrove-vram": {
"command": "npx",
"args": [
"-y",
"@nodegrove/vram-mcp"
]
}
}
} What you can ask
Six tools. You ask in your own words; the assistant picks the tool. Names are matched the way people type them, so "4090", "M4 Max" and "Llama-3.3-70B-Instruct" all work, and a card that is not listed can be described by its memory.
| Tool | Ask it | It answers with |
|---|---|---|
| can_i_run | Can my RTX 4090 run Llama 3.3 70B with 32k of context? | Fits, tight or no, the memory split, a speed ceiling, the longest context that fits and, on a no, every change that would make it fit. |
| what_fits | What is the best model for my 16 GB card? | Every model checked on one card, with a recommended everyday model, the largest that fits, the best at Q8 and the first out of reach. |
| estimate_vram | How much VRAM does Qwen3 32B need at 64k? | Weights, KV cache and overhead at each quantisation, and the smallest common card that holds each. |
| estimate_from_hf_repo | How much memory does Qwen/Qwen3-Next-80B-A3B-Instruct need? | Any Hugging Face repo, read from its config.json: the attention layout, what each 1,000 tokens of context costs, and memory at every quantisation. |
| list_models, list_gpus | Which GPUs do you know? | The models and cards with their specs, ids and pages here. |
An answer
Asked whether an RTX 4090 runs Llama 3.3 70B, the server answers:
No: Llama 3.3 70B at Q4_K_M with 8,192 tokens of context needs 45.8 GB, and the RTX 4090 holds 22.8 GB after headroom. Short by 23 GB.
What would work instead:
- No context length helps: the weights alone are 43.1 GB before a single token of conversation.
- RTX 6000 Ada (48 GB) is within a whisker: 45.8 GB against 45.6 GB after headroom. With the KV cache at Q8 it needs 44.4 GB and fits.
- The smallest card here that runs it exactly as asked: A100, 80 GB usable.
- The biggest model your card does run at these settings, counting mixture-of-experts models at their dense equivalent: Qwen3 32B, 22.4 GB at ~37 tokens/s.
Alongside the sentences come the numbers as fields (memory split, budget, speed ceiling, longest context), links to the model and card pages here, a link that opens the checker on the same settings, and the assumptions behind each figure. When a model does not fit, the answer also says which GPU size a Nodegrove workspace would attach for it; Nodegrove is in early access.
How it works
The server is this site's arithmetic, published as the open-source package @nodegrove/llm-math. The pages, the calculators, the open dataset and the server all import it, so they cannot disagree, and a correction reaches all of them at once.
- Memory is weights plus KV cache plus overhead, with each model's cache counted the way it actually caches: sliding-window, hybrid and latent attention store far less than a standard transformer. The method has every constant.
- Models in the list were read from their config.json and checked against Hugging Face; the data version is 2026-09-25. Any other repo is read live by the same rules, and anything the reader cannot model is named in the answer rather than guessed.
- Speed is a ceiling from memory bandwidth, not a benchmark. The faster the figure, the further real runtimes fall below it, so figures of 100 tokens/s or more say "at most".
Limits
- Every figure is an estimate from stated formulas. Real usage moves with the runtime, batch size, flash attention and cache quantisation: read "fits" as "worth trying".
- For a mixture-of-experts model read from Hugging Face, the speed needs the active parameters from its model card; without them the answer gives memory only.
- The remote endpoint answers up to 120 requests a minute from one address. A Hugging Face lookup waits up to 8 seconds.
Privacy
- No account, no key, no cookies. The remote server runs on Cloudflare Workers. We keep no request logs and no record of what you ask: every request is answered by a fresh instance and forgotten. Cloudflare keeps standard edge logs for a short period, as for any website.
- The rate limit counts requests per address for one minute, inside Cloudflare, and keeps nothing after that.
- When you ask about a Hugging Face repo, the server fetches that repo's public config.json and metadata. Only the repo name is sent.
- The local version sends nothing anywhere, except to huggingface.co when you ask about a repo there.
The privacy notice has the same in full.
Open source
The server and the package are on GitHub at nodegrove/vram-mcp, the server is on npm as @nodegrove/vram-mcp, and it is listed in the official MCP Registry as io.nodegrove/vram-mcp. The code is MIT; the model and GPU data is CC BY 4.0, the same licence as the dataset. To use the math in your own code:
npm install @nodegrove/llm-math Found an answer that disagrees with its source? Open an issue or write to info@nodegrove.io.
- Protocol: the Model Context Protocol, served to clients on the 2025 revisions and on 2026-07-28.
- Model architecture: each model's config.json on Hugging Face, linked from every answer and every model page.
- GPU memory and bandwidth: the manufacturers' specifications, linked from every GPU page.