Blog
Engineering notes on evaluation, local inference, and building useful AI systems.
What I learned building memory-aware inference on Apple Silicon
OCTOBER 4, 2026
Memory budgets, fixture correctness, and why bounded native dispatch does not establish a whole-model speedup.