{
"$type": "site.standard.document",
"bskyPostRef": {
"cid": "bafyreidzcbxqbsddsmgoppnyyvxnmqvrckmllmeqxpeuroliud7gjxjhku",
"uri": "at://did:plc:pgryn3ephfd2xgft23qokfzt/app.bsky.feed.post/3mpb2qsqnpgr2"
},
"path": "/t/i-built-a-novel-triple-hybrid-llm-mamba-attention-32-expert-moe-from-scratch-for-50-titan-v1-complete-titan-v2-first-cycle-done-expanding-dataset-now/177063#post_6",
"publishedAt": "2026-06-27T04:48:45.000Z",
"site": "https://discuss.huggingface.co",
"textContent": "Hi Mateusz, this is a really useful build log. The parts that stood out are the repeated training cycles, the autoresume checkpoint at step 47,000, and keeping the work viable on L4/A100-level budget instead of a cluster.\n\nI’m Rahul, founder of VaultLayer. VaultLayer is live now and helps individual builders run training/fine-tuning jobs on reliable, available, and affordable GPUs with monitoring, checkpointing, and recovery around the run.\n\nIf you want to compare it against the current AWS/L4 workflow for Titan v2 checkpoints or the 57M ablation runs, happy to share details.\n\nRahul\nFounder, VaultLayer",
"title": "🧠 I built a novel triple-hybrid LLM (Mamba + Attention + 32-expert MoE) from scratch for ~$50 — Titan v1 complete, Titan v2 first cycle done, expanding dataset now"
}