Can I run it? in your AI assistant

A free MCP server that tells Claude, Cursor or any AI assistant whether an open model fits your GPU, and what would fit instead, with this site's formulas.

Updated · estimates are labelled as estimates
On this page

Ask your assistant whether a model runs on your card and it answers with the arithmetic behind every figure on this site: the same formulas as the checker, over the same 29 models and 20 GPUs, and over any other model on Hugging Face, read from its config.json. It is free, read-only and needs no account or key.

Connect

Remote, with nothing to install. Add this URL to any client that speaks Streamable HTTP:

https://mcp.nodegrove.io/mcp
Claude (claude.ai and the desktop app)

Customize → Connectors → Add → Add custom connector. Name it Nodegrove VRAM, paste the URL, choose No sign-in, and turn it on in a chat from + → Connectors. Every plan includes at least one custom connector.

Claude Code

One command (add --scope user to have it in every project):

claude mcp add --transport http nodegrove-vram https://mcp.nodegrove.io/mcp
Cursor

Use the install link, or add this to ~/.cursor/mcp.json:

Add to Cursor

{
  "mcpServers": {
    "nodegrove-vram": {
      "url": "https://mcp.nodegrove.io/mcp"
    }
  }
}
VS Code

Use the install link, or add this to .vscode/mcp.json:

Install in VS Code

{
  "servers": {
    "nodegrove-vram": {
      "type": "http",
      "url": "https://mcp.nodegrove.io/mcp"
    }
  }
}
Devin Desktop (formerly Windsurf)

One command, or the same entry in mcp_config.json under mcpServers with a url field:

devin mcp add -s user nodegrove-vram https://mcp.nodegrove.io/mcp
LM Studio

Use the install link, or Program → Install → Edit mcp.json:

Add to LM Studio

{
  "mcpServers": {
    "nodegrove-vram": {
      "url": "https://mcp.nodegrove.io/mcp"
    }
  }
}
Open WebUI

Admin Settings → Integrations → External Tool Servers → Add Connection. Type MCP (Streamable HTTP), the URL, authentication None.

Cline

Add this to the MCP settings file. Without the type, Cline treats a URL as the older SSE transport:

{
  "mcpServers": {
    "nodegrove-vram": {
      "type": "streamableHttp",
      "url": "https://mcp.nodegrove.io/mcp"
    }
  }
}
Codex CLI

One command, or a [mcp_servers.nodegrove-vram] table with this url in ~/.codex/config.toml:

codex mcp add nodegrove-vram --url https://mcp.nodegrove.io/mcp
Gemini CLI

One command:

gemini mcp add -s user --transport http nodegrove-vram https://mcp.nodegrove.io/mcp
Zed

Add this to settings.json:

{
  "context_servers": {
    "nodegrove-vram": {
      "url": "https://mcp.nodegrove.io/mcp"
    }
  }
}
ChatGPT

Settings → Security and login → turn on Developer mode. Then at chatgpt.com/plugins, add one with the URL and No Authentication, and pick it in a chat from + → Developer mode.

Or run it on your own machine over stdio. It needs Node.js 20 or newer and nothing else:

{
  "mcpServers": {
    "nodegrove-vram": {
      "command": "npx",
      "args": [
        "-y",
        "@nodegrove/vram-mcp"
      ]
    }
  }
}

What you can ask

Six tools. You ask in your own words; the assistant picks the tool. Names are matched the way people type them, so "4090", "M4 Max" and "Llama-3.3-70B-Instruct" all work, and a card that is not listed can be described by its memory.

ToolAsk itIt answers with
can_i_runCan my RTX 4090 run Llama 3.3 70B with 32k of context?Fits, tight or no, the memory split, a speed ceiling, the longest context that fits and, on a no, every change that would make it fit.
what_fitsWhat is the best model for my 16 GB card?Every model checked on one card, with a recommended everyday model, the largest that fits, the best at Q8 and the first out of reach.
estimate_vramHow much VRAM does Qwen3 32B need at 64k?Weights, KV cache and overhead at each quantisation, and the smallest common card that holds each.
estimate_from_hf_repoHow much memory does Qwen/Qwen3-Next-80B-A3B-Instruct need?Any Hugging Face repo, read from its config.json: the attention layout, what each 1,000 tokens of context costs, and memory at every quantisation.
list_models, list_gpusWhich GPUs do you know?The models and cards with their specs, ids and pages here.

An answer

Asked whether an RTX 4090 runs Llama 3.3 70B, the server answers:

No: Llama 3.3 70B at Q4_K_M with 8,192 tokens of context needs 45.8 GB, and the RTX 4090 holds 22.8 GB after headroom. Short by 23 GB.

What would work instead:

  • No context length helps: the weights alone are 43.1 GB before a single token of conversation.
  • RTX 6000 Ada (48 GB) is within a whisker: 45.8 GB against 45.6 GB after headroom. With the KV cache at Q8 it needs 44.4 GB and fits.
  • The smallest card here that runs it exactly as asked: A100, 80 GB usable.
  • The biggest model your card does run at these settings, counting mixture-of-experts models at their dense equivalent: Qwen3 32B, 22.4 GB at ~37 tokens/s.

Alongside the sentences come the numbers as fields (memory split, budget, speed ceiling, longest context), links to the model and card pages here, a link that opens the checker on the same settings, and the assumptions behind each figure. When a model does not fit, the answer also says which GPU size a Nodegrove workspace would attach for it; Nodegrove is in early access.

How it works

The server is this site's arithmetic, published as the open-source package @nodegrove/llm-math. The pages, the calculators, the open dataset and the server all import it, so they cannot disagree, and a correction reaches all of them at once.

  • Memory is weights plus KV cache plus overhead, with each model's cache counted the way it actually caches: sliding-window, hybrid and latent attention store far less than a standard transformer. The method has every constant.
  • Models in the list were read from their config.json and checked against Hugging Face; the data version is 2026-09-25. Any other repo is read live by the same rules, and anything the reader cannot model is named in the answer rather than guessed.
  • Speed is a ceiling from memory bandwidth, not a benchmark. The faster the figure, the further real runtimes fall below it, so figures of 100 tokens/s or more say "at most".

Limits

  • Every figure is an estimate from stated formulas. Real usage moves with the runtime, batch size, flash attention and cache quantisation: read "fits" as "worth trying".
  • For a mixture-of-experts model read from Hugging Face, the speed needs the active parameters from its model card; without them the answer gives memory only.
  • The remote endpoint answers up to 120 requests a minute from one address. A Hugging Face lookup waits up to 8 seconds.

Privacy

  • No account, no key, no cookies. The remote server runs on Cloudflare Workers. We keep no request logs and no record of what you ask: every request is answered by a fresh instance and forgotten. Cloudflare keeps standard edge logs for a short period, as for any website.
  • The rate limit counts requests per address for one minute, inside Cloudflare, and keeps nothing after that.
  • When you ask about a Hugging Face repo, the server fetches that repo's public config.json and metadata. Only the repo name is sent.
  • The local version sends nothing anywhere, except to huggingface.co when you ask about a repo there.

The privacy notice has the same in full.

Open source

The server and the package are on GitHub at nodegrove/vram-mcp, the server is on npm as @nodegrove/vram-mcp, and it is listed in the official MCP Registry as io.nodegrove/vram-mcp. The code is MIT; the model and GPU data is CC BY 4.0, the same licence as the dataset. To use the math in your own code:

npm install @nodegrove/llm-math

Found an answer that disagrees with its source? Open an issue or write to info@nodegrove.io.

  • Protocol: the Model Context Protocol, served to clients on the 2025 revisions and on 2026-07-28.
  • Model architecture: each model's config.json on Hugging Face, linked from every answer and every model page.
  • GPU memory and bandwidth: the manufacturers' specifications, linked from every GPU page.