1. Define the claim before measuring it
Every run records the system, source collection, model roles, retrieval and prompt configuration, operating mode, sample size, and time of evaluation. The result applies to that recorded combination—not to every deployment of the underlying model.
- Name the model or assistant under test
- Fix the approved source boundary
- Record retrieval, prompt, runtime, and model-role settings
- State whether the run is complete or sampled
- Preserve dates, exclusions, connection errors, and operator decisions