← Alpha Ranch
Inference Lab
local iron vs the cloud — real runs, real dollars
Methodology: prompt suite = short QA, ~800-word longform, code generation; N reps per model per prompt;
TTFT = time to first streamed token; tok/s = output tokens / generation time (medians shown).
Local models run on the Alpha Ranch rig (AMD Strix Halo, ROCm); cloud models are called from the same host.
$/1k = provider list price per 1k output tokens; local inference has no per-token cost. No synthetic numbers — every row is a real measured run.