{
  "$type": "site.standard.document",
  "bskyPostRef": {
    "cid": "bafyreigk77qq3ec4kn377hmd27tve7ouwbjx7befzupgflb5jkh2zhppsu",
    "uri": "at://did:plc:pgryn3ephfd2xgft23qokfzt/app.bsky.feed.post/3mnk2sra6k552"
  },
  "path": "/t/fitllm-open-source-calculator-will-this-llm-fit-my-gpu-mac-accurate-on-gemma-4-qwen-3-6/176567#post_1",
  "publishedAt": "2026-06-05T11:10:41.000Z",
  "site": "https://discuss.huggingface.co",
  "tags": [
    "https://fitllm.run",
    "GitHub - click6067-ship-it/fitllm-engine: Accurate LLM memory & speed calculator — models sliding-window/linear/MoE attention where naive VRAM/RAM calculators are 4–11× off. Apple Silicon + NVIDIA RTX · GGUF · one MIT file. Powers fitllm.run. · GitHub"
  ],
  "textContent": "Live: https://fitllm.run\nEngine (MIT, one file): GitHub - click6067-ship-it/fitllm-engine: Accurate LLM memory & speed calculator — models sliding-window/linear/MoE attention where naive VRAM/RAM calculators are 4–11× off. Apple Silicon + NVIDIA RTX · GGUF · one MIT file. Powers fitllm.run. · GitHub\n\nMost calculators use the textbook KV-cache formula and are multiples off on modern models:\nGemma 4 (sliding-window), Qwen 3.6 (linear attention), MoE. FitLLM reads each model’s\nconfig.json and models each layer type separately. Apple Silicon + NVIDIA RTX. Paste any HF\nmodel id. Estimator, not ground truth.",
  "title": "FitLLM – open-source calculator: will this LLM fit my GPU/Mac? (accurate on Gemma 4 / Qwen   3.6)"
}