← Alpha Ranch

Inference Lab

local iron vs the cloud — real runs, real dollars
// loading…
Methodology: prompt suite = short QA, ~800-word longform, code generation; N reps per model per prompt; TTFT = time to first streamed token; tok/s = output tokens / generation time (medians shown). Local models run on the Alpha Ranch rig (AMD Strix Halo, ROCm); cloud models are called from the same host. $/1k = provider list price per 1k output tokens; local inference has no per-token cost. No synthetic numbers — every row is a real measured run.