Qwen3.8 27B on Dual RTX 3090s: A vLLM Field Report
Three days into making Qwen3.8 27B my daily local coding driver on dual 3090s — the hybrid linear-attention architecture, an OOM that masqueraded as a backend error, and what the …
Three days into making Qwen3.8 27B my daily local coding driver on dual 3090s — the hybrid linear-attention architecture, an OOM that masqueraded as a backend error, and what the …
What this proves: Resilience — enterprise-grade AI infrastructure at a single operator's desk. A custom routing plane in front of a three-server GPU fleet that switches models and …
How self-hosted AI became the final piece of my homelab puzzle, delivering true parallel processing for multi-user setups and unlocking the real superpower of knowledge management.
I build and operate AI infrastructure right now, today.