Writing

Writing about LLM inference, long-context serving, GPU and CPU architecture, and the systems work that sits between a model and the hardware it runs on.

2026