README
██████╗ ██╗████████╗███████╗███████╗██╗ ██╗ ██╔══██╗██║╚══██╔══╝██╔════╝██╔════╝██║ ██║ ██████╔╝██║ ██║ █████╗ ███████╗███████║ ██╔══██╗██║ ██║ ██╔══╝ ╚════██║██╔══██║ ██║ ██║██║ ██║ ███████╗███████║██║ ██║ ╚═╝ ╚═╝╚═╝ ╚═╝ ╚══════╝╚══════╝╚═╝ ╚═╝ ██╗ ██╗ █████╗ ██████╗ █████╗ ██╗ ██╗ ╚██╗ ██╔╝██╔══██╗██╔══██╗██╔══██╗██║ ██║ ╚████╔╝ ███████║██║ ██║███████║██║ ██║ ╚██╔╝ ██╔══██║██║ ██║██╔══██║╚██╗ ██╔╝ ██║ ██║ ██║██████╔╝██║ ██║ ╚████╔╝ ╚═╝ ╚═╝ ╚═╝╚═════╝ ╚═╝ ╚═╝ ╚═══╝
Independent research notes by Ritesh Yadav on ML performance, infrastructure, and systems: CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.
These pages collect working notes on the GPU stack and production inference: definitions, references, and the reasoning I use when reading kernels, profilers, and serving systems. They are written for careful reading, not as product documentation.
Research Topics
This is the map of the notes. Each area links to its index page and calls out the questions it is meant to answer.
| # | Research area | What it covers |
|---|---|---|
| 01 | Device Hardware | Microarchitecture, streaming multiprocessors, and the memory hierarchy |
| 02 | Device Software | CUDA programming, kernels, tiling, and device-side execution |
| 03 | Host Software | Driver APIs, runtimes, CUTLASS, and CuTe |
| 04 | Performance | Roofline analysis, memory coalescing, and bank conflicts |
| 05 | Inference Engineering | Quantization, vLLM, and low-latency serving |
| 06 | Large Language Models | Attention kernels, parallelism, and mixture-of-experts systems |
Browse the Notes
The table of contents on the left is the primary index on desktop. On smaller screens, open the menu to move between areas. The previous and next arrows at the bottom of each page follow the same reading order.
About the author
Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.