L0 FREE · One passing LLM run is one observation.
A green check can hide unstable behavior.
Variable output can remain schema-valid while its meaning, grounding, or consistency changes. This free PDF introduces five failure patterns, compact detection signals, and the limits you should record before treating an evaluation as release evidence.
- 01 Five reviewable patterns: output variation, context degradation, unsupported claims, format instability, and semantic drift
- 02 Compact Python reference snippets with explicit parameters, assumptions, and review triggers
- 03 Guidance for choosing thresholds and sample plans from your decision, risk, and labeled evidence rather than a universal number