Three and a Half Weeks, 689M Tokens, No Meter
689M tokens. 25 days. ~$40 electricity. The math, the caveats, and what changed when the meter stopped existing.
689M tokens. 25 days. ~$40 electricity. The math, the caveats, and what changed when the meter stopped existing.
Three days into making Qwen3.8 27B my daily local coding driver on dual 3090s — the hybrid linear-attention architecture, an OOM that masqueraded as a backend error, and what the …
What this proves: Resilience — enterprise-grade AI infrastructure at a single operator's desk. A custom routing plane in front of a three-server GPU fleet that switches models and …
I hooked up multiple local models behind a single proxy that routes traffic by complexity. Here is what I found deploying it on a cheap Intel GPU.
How I went from a blank Docker template to 116+ tok/s with speculative decoding, FlashInfer, and a 160k context window on dual 3090s.
The secret to a great AI assistant isn't the model — it's what you tell it about yourself. Here's the framework for building a system prompt that makes AI actually useful.
A real-world comparison of Qwen3.5 27B Q8 and 35B-A3B Q8 running locally on a dual RTX 3090 homelab — which one actually belongs in your daily workflow?
A year ago, I predicted CLI AI tools would transform development. Here's what actually happened—the good, the unexpected, and the 'wait, what?' moments.
The art of crafting perfect search queries has evolved into the skill of prompt engineering. Learn how mastering AI interactions is the next generation of being a master Googler.
Skip the theory overload—build working ML models fast with TensorFlow and deploy them in production.
By Laurence Moroney
Showing 1-10 of 16 items (Page 1 of 2)