Skip to content

Build trustworthy workflows

Evaluate AI Before Trusting It

Last updated:

An AI evaluation tests whether a system meets a defined requirement on representative tasks. Start with a small set of examples you can check, rather than an impressive demonstration.

Define success before running the test

For a notes assistant, success might mean answering from a supplied passage, naming the source, and declining questions the passage cannot answer. Our suggested rubric gives one point for each of those checks. It is a learning exercise, not an industry benchmark.

Create ten questions: four straightforward, three needing evidence from more than one passage, two without an answer in the notes, and one containing distracting instructions. Write expected answers before testing. Use fictional or public material.

Record failures you can act on

Keep the model ID, prompt, source passages and date with each result. Label a failure as retrieval, unsupported claim, formatting or unauthorized action. A fluent answer with a wrong source is still a failure. Have a person inspect borderline examples; an AI judge can make mistakes too.

After changing the prompt or model, rerun the same set. Keep some examples aside so you do not improve only the questions you already know. For agents, inspect tool calls as well as the final response.

Source and practice

Anthropic's evaluation guide explains success criteria and test design. Build your own scorecard before moving on to automation.

Progress is saved in your browser only — no account, nothing sent anywhere.