Deep dives per domain.
Tokenizers, attention variants, long-context positional encodings, MoE routing, decoding strategies, and state-space alternatives.
60 problems10 sections★ 32 high-frequency
Sections (10)
- Tokenization
- Core attention and Transformer blocks
- Attention variants
- Long-context positional encodings
- Activations and FFN variants
- Autoregressive generation and decoding
- Mixture-of-Experts
- Training tricks
- State-space and alternatives
- Closing tips
Noise schedules, DDPM and DDIM, score matching, classifier-free guidance, ODE solvers, flow matching, rectified flow, and consistency.
69 problems16 sections★ 23 high-frequency
Sections (16)
- Noise schedules and SDE coefficients
- DDPM forward and posteriors
- DDPM training
- DDPM sampling (reverse process)
- DDIM and accelerated samplers
- Score-based formulation
- Classifier-free guidance and conditioning
- Higher-order ODE solvers
- Latent diffusion
- Flow matching
- Rectified flow
- Normalizing flows
- Mean flow and consistency
- Architecture pieces and training tricks
- Inversion, editing, and image-to-image
- Closing tips
Inference efficiency end to end: KV cache, quantization, distillation, LoRA and PEFT, pruning, speculative decoding, batching, and serving.
55 problems13 sections★ 18 high-frequency
Sections (13)
- KV cache fundamentals
- KV-cache quantization
- Knowledge distillation
- Quantization fundamentals
- Activation-aware quantization
- PEFT: LoRA family
- Other PEFT methods
- Pruning
- Speculative decoding
- Batching and serving
- Export and runtime
- Distributed inference
- Closing tips
Cameras and rays, volume rendering, NeRF training, Gaussian splatting, spherical harmonics, 4D/dynamic scenes, and SLAM.
60 problems14 sections★ 15 high-frequency
Sections (14)
- Camera, rays, and geometry
- Volume rendering
- NeRF building blocks
- NeRF training
- Tri-plane / TensoRF
- Gaussian Splatting fundamentals
- Spherical harmonics
- 3DGS training mechanics
- 2DGS, surfaces, and meshes
- Dynamic / 4D Gaussian Splatting
- Compression and serving
- Metrics and tools
- SLAM and 3DGS
- Closing tips
Video tokenizers and 3D VAEs, video diffusion, vision-language-action policies, and world models.
40 problems6 sections★ 15 high-frequency
Sections (6)
- Video tokenization and VAEs
- Video diffusion
- VLA: Vision-Language-Action models
- World models
- Multimodal generation tricks
- Closing tips