Long-form, honest accounts of the things I build, with the numbers, the receipts, and the failures kept in.
My GPU has 16 GB of VRAM. The model needs 234 GB just to sit on disk. The story of Tarfa, an inference engine built in 23 days that streams exactly the experts each token asks for, bit-identical BF16, SHA-256-verified, and why "exact" turned out to be the whole point.
Read the post →