Vllm

Qwen3.8 27B on Dual RTX 3090s: A vLLM Field Report featured image

Qwen3.8 27B on Dual RTX 3090s: A vLLM Field Report

Three days into making Qwen3.8 27B my daily local coding driver on dual 3090s — the hybrid linear-attention architecture, an OOM that masqueraded as a backend error, and what the …

Derek Armstrong - Payments Engineer · AI · Infrastructure
Derek Armstrong
Read more
Highly Available AI Inference Cluster featured image

Highly Available AI Inference Cluster

What this proves: Resilience — enterprise-grade AI infrastructure at a single operator's desk. A custom routing plane in front of a three-server GPU fleet that switches models and …

Read more
Self Hosted AI: Actually Running Local LLMs for a Multi-User Household featured image

Self Hosted AI: Actually Running Local LLMs for a Multi-User Household

How self-hosted AI became the final piece of my homelab puzzle, delivering true parallel processing for multi-user setups and unlocking the real superpower of knowledge management.

Derek Armstrong - Payments Engineer · AI · Infrastructure
Derek Armstrong
Read more
Self-Hosted AI Inference Cluster featured image

Self-Hosted AI Inference Cluster

I build and operate AI infrastructure right now, today.

Read more