An output can fail in more than one way at once, and each failure mode is distinct enough to check separately:
| Dimension | What it catches |
|---|---|
| Accuracy | Are the facts, figures, and claims actually correct? |
| Completeness | Does the output address every part of what was asked — not just the easiest part? |
| Consistency | Does it hold together internally, and does it stay aligned with any source material it was supposed to be grounded in? |
| Relevance | Does it answer the actual question, rather than a nearby one? |
| Audience suitability | Is the tone, complexity, and format right for who will actually read it? |
These are independent failure modes. An output can be accurate but incomplete — every stated fact checks out, but half the request went unaddressed. It can be complete but inconsistent — every part got answered, but a number stated in paragraph one contradicts one in paragraph three. It can be consistent but irrelevant — internally coherent and well-argued, but answering a question adjacent to the one actually asked. Checking only one dimension (usually accuracy, since it feels the most "factual") lets the other four slip past.
Audience suitability loops back to Module 2: evaluating an output partly means checking whether the audience specificity dial from the original prompt was actually honored in the result.