Ritesh Yadav
Ritesh Yadav
/readme

README

██████╗ ██╗████████╗███████╗███████╗██╗  ██╗
██╔══██╗██║╚══██╔══╝██╔════╝██╔════╝██║  ██║
██████╔╝██║   ██║   █████╗  ███████╗███████║
██╔══██╗██║   ██║   ██╔══╝  ╚════██║██╔══██║
██║  ██║██║   ██║   ███████╗███████║██║  ██║
╚═╝  ╚═╝╚═╝   ╚═╝   ╚══════╝╚══════╝╚═╝  ╚═╝
██╗   ██╗ █████╗ ██████╗  █████╗ ██╗   ██╗
╚██╗ ██╔╝██╔══██╗██╔══██╗██╔══██╗██║   ██║
 ╚████╔╝ ███████║██║  ██║███████║██║   ██║
  ╚██╔╝  ██╔══██║██║  ██║██╔══██║╚██╗ ██╔╝
   ██║   ██║  ██║██████╔╝██║  ██║ ╚████╔╝
   ╚═╝   ╚═╝  ╚═╝╚═════╝ ╚═╝  ╚═╝  ╚═══╝

Independent research notes by Ritesh Yadav on ML performance, infrastructure, and systems: CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.

These pages collect working notes on the GPU stack and production inference: definitions, references, and the reasoning I use when reading kernels, profilers, and serving systems. They are written for careful reading, not as product documentation.

Research Topics

This is the map of the notes. Each area links to its index page and calls out the questions it is meant to answer.

#Research areaWhat it covers
01Device HardwareMicroarchitecture, streaming multiprocessors, and the memory hierarchy
02Device SoftwareCUDA programming, kernels, tiling, and device-side execution
03Host SoftwareDriver APIs, runtimes, CUTLASS, and CuTe
04PerformanceRoofline analysis, memory coalescing, and bank conflicts
05Inference EngineeringQuantization, vLLM, and low-latency serving
06Large Language ModelsAttention kernels, parallelism, and mixture-of-experts systems

Browse the Notes

The table of contents on the left is the primary index on desktop. On smaller screens, open the menu to move between areas. The previous and next arrows at the bottom of each page follow the same reading order.

About the author

Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.