I’m Mohammad Ummair. I work on machine learning systems — the layer where model execution meets the hardware underneath it.

This site is where I write up things I’ve measured, read, or argued about: inference performance, long-context serving, memory hierarchies, and the tradeoffs that only show up once something is actually running.

Posts here are working notes rather than finished papers. If something looks wrong, I would like to know.