Three and a Half Weeks, 689M Tokens, No Meter
689M tokens. 25 days. ~$40 electricity. The math, the caveats, and what changed when the meter stopped existing.
689M tokens. 25 days. ~$40 electricity. The math, the caveats, and what changed when the meter stopped existing.
Three days into making Qwen3.8 27B my daily local coding driver on dual 3090s — the hybrid linear-attention architecture, an OOM that masqueraded as a backend error, and what the …
How self-hosted AI became the final piece of my homelab puzzle, delivering true parallel processing for multi-user setups and unlocking the real superpower of knowledge management.
How I went from a blank Docker template to 116+ tok/s with speculative decoding, FlashInfer, and a 160k context window on dual 3090s.
A real-world comparison of Qwen3.5 27B Q8 and 35B-A3B Q8 running locally on a dual RTX 3090 homelab — which one actually belongs in your daily workflow?
Pydantic is one of those libraries I underestimated until the day it saved me four hours of debugging. Here's what it actually does, where it hurts, and why Pydantic-AI has me …
Set up your own AI chatbot locally using Meta's Llama model and Docker in just two commands