Ritesh Yadav
Research Notes
Inference Engineering
Notes from production inference work, serving models, especially large language models, efficiently under real traffic: latency, throughput, and cost per token.
Topics will cover prefill vs decode, KV cache, continuous batching, paged attention, speculative decoding, and quantization.
Notes and deep-dives for this category will land here.
About the author
Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.