{
"$type": "site.standard.document",
"bskyPostRef": {
"cid": "bafyreigk77qq3ec4kn377hmd27tve7ouwbjx7befzupgflb5jkh2zhppsu",
"uri": "at://did:plc:pgryn3ephfd2xgft23qokfzt/app.bsky.feed.post/3mnk2sra6k552"
},
"path": "/t/fitllm-open-source-calculator-will-this-llm-fit-my-gpu-mac-accurate-on-gemma-4-qwen-3-6/176567#post_1",
"publishedAt": "2026-06-05T11:10:41.000Z",
"site": "https://discuss.huggingface.co",
"tags": [
"https://fitllm.run",
"GitHub - click6067-ship-it/fitllm-engine: Accurate LLM memory & speed calculator — models sliding-window/linear/MoE attention where naive VRAM/RAM calculators are 4–11× off. Apple Silicon + NVIDIA RTX · GGUF · one MIT file. Powers fitllm.run. · GitHub"
],
"textContent": "Live: https://fitllm.run\nEngine (MIT, one file): GitHub - click6067-ship-it/fitllm-engine: Accurate LLM memory & speed calculator — models sliding-window/linear/MoE attention where naive VRAM/RAM calculators are 4–11× off. Apple Silicon + NVIDIA RTX · GGUF · one MIT file. Powers fitllm.run. · GitHub\n\nMost calculators use the textbook KV-cache formula and are multiples off on modern models:\nGemma 4 (sliding-window), Qwen 3.6 (linear attention), MoE. FitLLM reads each model’s\nconfig.json and models each layer type separately. Apple Silicon + NVIDIA RTX. Paste any HF\nmodel id. Estimator, not ground truth.",
"title": "FitLLM – open-source calculator: will this LLM fit my GPU/Mac? (accurate on Gemma 4 / Qwen 3.6)"
}