Ritesh Yadav
Ritesh Yadav
/large-language-models

Large Language Models

Vocabulary for designing and operating LLM systems, architecture choices and how tokens move through a serving stack: transformers, attention, context windows, embeddings, autoregression, MoE, and fine-tuning.

Notes and deep-dives for this category will land here.

About the author

Ritesh Yadav works as an AI/ML Engineer. He writes independent research notes on ML performance, infrastructure, and systems, covering CUDA, low-latency inference, generative AI, distributed training, Kubernetes, and LLMOps.