Learn the playbook
This map is a navigation and learning aid derived from the playbook's own lifecycle — the chapters themselves are the normative source.
The lifecycle
Wherever a project enters, work proceeds through this loop. Each stage links to the chapter that owns its procedure.
- 1. Decision context
Turn a fuzzy request into a profile, a stated decision, and first actions.
01 · Project Intake and Decision Context - 2. Execution-system definition
Pin what “the same system” means; measure reproducibility boundaries before comparing anything.
02 · Execution System Model - 3. Trusted evaluation substrate
Build the versioned eval with integrity gates and measured ceilings — before optimization.
03 · Evaluation Foundation - 4. Experimental design
Frozen contracts, MDE, consequence-bearing tolerances, INCONCLUSIVE as an outcome.
04 · Experiment Design and Statistics - 5. Baseline & candidates
Simplest credible baseline first; bounded candidate sets; runtime regimes, not winners.
05 · Model, Runtime, and Harness Selection - 6. Failure diagnosis
Classify observed failures in the canonical taxonomy; instrument defects first.
03 · Evaluation Foundation07 · Optimization and the Intervention Ladder - 7. Evidence-selected intervention
Climb the intervention ladder with diagnostic gates — retrieval/tools/routing before training.
07 · Optimization and the Intervention Ladder08 · Retrieval, Tools, Workflows, and Routing09 · Training and Data - 8. Performance & capacity
Characterize serving performance and capacity fit against real constraints.
06 · Inference Performance and Capacity - 9. Economics
Rent vs buy vs API as a measured decision: demand ledgers and pre-committed triggers.
11 · Economics, Hardware, and Cloud - 10. Deployment
Shadow and canary with restraint; rollback as a designed, rehearsed path.
10 · Deployment and Operations - 11. Observability & learning
Telemetry floor for experimentation; harvest production evidence back into the eval (→ 03).
12 · Observability, Learning, and Promotion
Production evidence flows back into the evaluation substrate (12 → 03) — the loop never ends.
Cross-cutting chapters
How chapters depend on each other
The book's own reading contracts — the ordering rules that prevent expensive mistakes.
Major concepts
The glossary's own seven thematic clusters — each opens the filtered glossary.
Coming later: agent prompt recipes
The same canonical chapters and templates that render these reading pages can be assembled into bounded, agent-ready stage prompts — “Start a Project”, “Freeze an Eval Suite”, “Run a Performance Autopsy”. That layer is deliberately not built yet; the content pipeline already exposes every section it would need, so it can be added without a second source of truth.