Methodology

A number should survive the challenge.

VeriTrooper records source-linked references, deterministic results, diagnostic opinions, and human decisions separately. Review the checks, configured model roles, and limits behind each result.

01

1. Define the claim before measuring it

Every run records the system, source collection, model roles, retrieval and prompt configuration, operating mode, sample size, and time of evaluation. The result applies to that recorded combination—not to every deployment of the underlying model.

  • Name the model or assistant under test
  • Fix the approved source boundary
  • Record retrieval, prompt, runtime, and model-role settings
  • State whether the run is complete or sampled
  • Preserve dates, exclusions, connection errors, and operator decisions
02

2. Build source-linked reference answers

Scout creates or imports questions and expected answers tied to the material the customer approved. Reference answers need review. Items labeled unanswerable test refusal behavior within the recorded evidence scope; they do not prove that no answer exists anywhere in the corpus.

  • Reference answers remain connected to source evidence
  • Question categories exercise different reasoning and retrieval demands
  • Unanswerable items test refusal behavior and unsupported certainty
  • Review test quality and record any exclusions or reference defects
03

3. Compare recorded configurations

Baseline and pipeline results use the same eligible questions across different recorded configurations, with symmetric exclusions. Their difference reflects the complete configurations, including retrieval and evidence presentation; it does not isolate the benefit of an individual control.

  • Baseline: reference retrieval or the selected deployed endpoint
  • Pipeline: question-linked evidence and the configured checks
  • Delta: the measured difference on this test set
  • Inspect sample size, paired gains and regressions, category results, and exclusions
04

4. Show how each answer was judged

Deterministic checks evaluate arithmetic, numeric precision, units, refusal behavior, and other supported rules against stored references. Diagnostic opinions and optional verification are recorded separately. Passing does not establish that a reference is correct.

  • Deterministic checks settle exact, normalized, refusal, and other rule-resolvable cases
  • A diagnostic reviewer explains suspected failures and assigns the failure class
  • A separate-vendor model or deterministic coded judge can review contested decisions
  • Model roles and shared models are disclosed; using one model in multiple roles is not independent model review
  • Advisory disagreements and review referrals remain visible alongside deterministic headline verdicts
05

5. Preserve the adjudication trail

A final score is backed by the item-level record. Reviewers can inspect the question, expected answer, model answer, source evidence, initial decision, diagnosis, verifier decision, failure category, and any human disposition.

  • Empty, hedged, unsupported, and malformed answers receive conservative treatment
  • False positives and bad tests remain distinguishable from confirmed model failures
  • Overrides and referrals are visible
  • Human signoff is an explicit action bound to the reviewed result set
06

6. Interpret the result within its boundary

The reported accuracy and delta describe measured behavior under the recorded conditions. They can support a deployment decision, remediation plan, or broader validation program; they do not promise that installation alone changes serving accuracy.

  • A point-in-time audit is not a permanent model rating
  • An absence of findings is limited to what was actually examined
  • Changes to the model, prompt, retrieval, sources, or production conditions can invalidate the inference
  • The evidence supports governance and conformity work but is not legal advice or a conformity certification
07

7. Make the result independently checkable

Successful, seal-eligible paid runs retain readable reports, applicable machine records, scope, and signature material. Evaluation provides a limited proof set. The portable checker verifies file integrity; reviewers assess the conclusions separately.

  • Hashes bind the listed artifacts to the manifest
  • The detached signature binds the manifest
  • The checker identifies changed, missing, or added files
  • Human signoff, when recorded, is hash-bound to the reviewed result set
  • RFC 3161 timestamping is recorded only when configured and successfully obtained
08

8. Re-audit when the system changes

Assurance expires when its factual basis changes. Watchtower extends the method into production by assessing captured answers or grounded probes on a cadence, while a new Scout run re-establishes a pre-release benchmark after material changes.

  • Re-audit after model, prompt, retrieval, corpus, or endpoint changes
  • Use Watchtower to detect production drift and recurring failure patterns
  • Retain prior packages so changes can be compared without rewriting history
Next step

Test the claim on your own system.

Discuss an assessment