Ritesh Yadav
Research Notes
Large Language Models
Vocabulary for designing and operating LLM systems, architecture choices and how tokens move through a serving stack: transformers, attention, context windows, embeddings, autoregression, MoE, and fine-tuning.
Notes and deep-dives for this category will land here.
About the author
Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.