Book 9 — Transformers & LLMs¶
The foundational guide to modern language models — from self-attention and positional encodings to foundation model families, distributed pretraining, alignment, and parameter-efficient fine-tuning.
Architecture¶
- Self-attention (Q/K/V)
- Multi-head attention
- Positional encoding
- Encoder & decoder
- Causal attention & masking
- Feed-forward, layer norm, residuals
Comprehensive Deep Dive:
Model Families¶
Comprehensive Deep Dive:
Training¶
- Pretraining & Distributed Systems
- Fine-tuning & instruction tuning (SFT)
- RLHF / DPO
- Quantization
- LoRA / QLoRA / PEFT
Comprehensive Deep Dive:
Mastery Roadmap
Each topic includes comprehensive mathematical derivations, architectural diagrams, PyTorch implementations from scratch, common debugging pitfalls, staff-level interview questions, and a 10-level mastery ladder.