Ritesh Yadav
Ritesh Yadav
/perf

Performance

GPUs exist when performance on general-purpose hardware isn't enough. Correctness is necessary but not sufficient, latency, throughput, and efficiency are the product.

This category collects the vocabulary of GPU performance: bottlenecks, roofline thinking, occupancy, coalescing, and related ideas.

Notes and deep-dives for this category will land here.

About the author

Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.