Knowledge graph

Explore the book

Browse by part and chapter, compare the 55 practices developed in the monograph with the 137 companion-only practices, and trace the evidence to 368 sources.

Parts
6
Chapters
18
Practices
192
Developed
55
Companion-only
137
References
368

Book explorer

Taught in chapter Carried in companion

Part 1

Evaluation measurement and experiment design

Chapter 1: Run-to-run variance, statistical power, and paired comparisons

13 practices

Chapter 2: Baselines, ablations, and cost-accuracy tradeoffs

8 practices

Chapter 3: Benchmark contamination, oracle strength, and workload validity

22 practices

Part 2

Evaluation and grading systems

Part 3

Containment, durable execution, and recovery engineering

Chapter 7: Agent isolation, injection defenses, and independent verification

9 practices

Chapter 8: Persistent agent state, durable workflows, and idempotent retries

14 practices

Chapter 9: Replayable traces and fault-injection recovery testing

9 practices

Chapter 10: Human-auditable failure analysis and taxonomy development

10 practices

Part 4

Context engineering: retrieval, budgets, and memory

Chapter 11: Measuring and designing repository retrieval

8 practices

Chapter 12: Localization funnels, repository indexes, and freshness checks

11 practices

Chapter 13: Usable context budgets, consolidated-spec restarts, and file-based tool output

10 practices

Chapter 14: Cross-session memory, raw traces, and compaction policies

9 practices

Part 5

Human review and accountability engineering

Chapter 15: Efficient verification interfaces and risk-based human escalation

11 practices

Chapter 16: Autonomy calibration, provenance, effective gates, and accountability

17 practices

Part 6

Research agenda: work allocation and cost engineering