ML Knowledge Base

39 reference documents · converted from Markdown

Master References3

Cross-cutting merges compiled from the whole knowledge base

Transformers & LLM Architecture6

Attention, tokenization, MoE, KV cache

Reasoning, RL & Agents7

Policy optimization, rewards, agentic systems, test-time compute

Agentic Intelligence — Technologies & Tricks Updated July 2026 with 2025–2026 SOTA additions — new entries marked ★. R-Agentic_Intelligence_SOTA_Updated.md Policy Optimization — All Variants & Tricks in RL Updated July 2026 with 2025–2026 SOTA additions — new entries marked ★. R-Policy_Optimization_SOTA_Updated.md Prompt · Context · Harness · Graph Engineering — and Self-Improving AI The engineering stack around a frozen model: how you prompt it, what context you feed it, the harness you wrap it in, the graphs you impose on it — and how systems improve themselves. R-Prompt_Context_Harness_Graph_Engineering_SOTA_Updated.md RL Training Strategies & Practical Recipes — Every recipe and engineering trick that turns a policy-optimization paper into a model that actually trains 22 sections · classical RL through 2026 RLHF R-RL_Training_Strategies_Recipes_CheatSheet_with_Links.md Reasoning Technologies — in Modern AI Updated July 2026 with 2025–2026 SOTA additions — new entries marked ★. R-Reasoning_Technologies_SOTA_Updated.md Reward Functions — All Variants & Design Patterns Updated July 2026 with 2025–2026 SOTA additions — new entries marked ★. R-Reward_Functions_SOTA_Updated.md Training-Free & Test-Time Optimization / Training — SOTA Cheat Sheet Methods that improve a frozen (or barely-touched) model at inference — no training run required. Decoding, activation steering, weight merging, speculative decoding, guidance; test-time compute & search; tes… R-Test_Time_and_Training_Free_Optimization_SOTA_Updated.md

Efficiency, Data & Evaluation6

Quantization, pruning, distillation, PEFT, data, metrics

Diffusion Models2

Formulations, derivations and samplers

Foundation & Generative Models5

Scaling laws, video, world models, VLA

3D Vision & Neural Rendering8

NeRF, Gaussian splatting, SfM, avatars, driving

Computer Vision Deep Dives2

Principal-level breadth and mathematics