Methodology

A number should survive the challenge.

VeriTrooper separates source evidence, deterministic judgment, independent review, and human responsibility so no model quietly grades its own homework.

01

1. Define the claim before measuring it

Every run records the system, source collection, model roles, retrieval and prompt configuration, operating mode, sample size, and time of evaluation. The result applies to that recorded combination—not to every deployment of the underlying model.

  • Name the model or assistant under test
  • Fix the approved source boundary
  • Record retrieval, prompt, runtime, and model-role settings
  • State whether the run is complete or sampled
  • Preserve dates, exclusions, connection errors, and operator decisions
02

2. Build ground truth from the approved sources

Scout creates or imports questions whose expected answers are tied to the material the customer approved. The test set includes supported questions and, where selected, questions the sources cannot answer so confident invention is measured rather than ignored.

  • Reference answers remain connected to source evidence
  • Question categories exercise different reasoning and retrieval demands
  • Unanswerable items test refusal behavior and unsupported certainty
  • Malformed or unusable tests are identified explicitly rather than silently scored
03

3. Compare like with like

When Scout runs a comparative audit, the baseline and VeriTrooper pipeline answer the same scored questions under the recorded conditions. Exclusions are applied symmetrically so one arm cannot improve merely by discarding harder cases.

  • Baseline: the model answering through the recorded reference retrieval
  • Pipeline: the same model with the VeriTrooper retrieval and control layer
  • Delta: the measured difference on this test set
  • Sample size, category results, confirmed failures, and exclusions accompany the headline score
04

4. Use the strongest available decision method

Clear cases are resolved with reproducible code and verified ground truth. Cases that cannot be settled mechanically move through diagnosis and, when configured, independent review.

  • Deterministic checks settle exact, normalized, refusal, and other rule-resolvable cases
  • A diagnostic reviewer explains suspected failures and assigns the failure class
  • A separate-vendor model or deterministic coded judge can review contested decisions
  • Same-vendor selections are disclosed because they weaken independence
  • Unresolved cases are referred to a person instead of being forced into pass or fail
05

5. Preserve the adjudication trail

A final score is backed by the item-level record. Reviewers can inspect the question, expected answer, model answer, source evidence, initial decision, diagnosis, verifier decision, failure category, and any human disposition.

  • Empty, hedged, unsupported, and malformed answers receive conservative treatment
  • False positives and bad tests remain distinguishable from confirmed model failures
  • Overrides and referrals are visible
  • Human signoff is an explicit action bound to the reviewed result set
06

6. Interpret the result within its boundary

The reported accuracy and delta describe measured behavior under the recorded conditions. They can support a deployment decision, remediation plan, or broader validation program; they do not promise that installation alone changes serving accuracy.

  • A point-in-time audit is not a permanent model rating
  • An absence of findings is limited to what was actually examined
  • Changes to the model, prompt, retrieval, sources, or production conditions can invalidate the inference
  • The evidence supports governance and conformity work but is not legal advice or a conformity certification
07

7. Make the result independently checkable

A completed, seal-eligible package carries the readable reports, canonical machine record, detailed evidence, configuration, logs, manifest, signature material, and a portable checker.

  • Hashes bind the listed artifacts to the manifest
  • The detached signature binds the manifest
  • The checker identifies changed, missing, or added files
  • Human signoff is hash-bound to the exact result set
  • RFC 3161 timestamping is recorded only when configured and successfully obtained
08

8. Re-audit when the system changes

Assurance expires when its factual basis changes. Watchtower extends the method into production by assessing captured answers or grounded probes on a cadence, while a new Scout run re-establishes a pre-release benchmark after material changes.

  • Re-audit after model, prompt, retrieval, corpus, or endpoint changes
  • Use Watchtower to detect production drift and recurring failure patterns
  • Retain prior packages so changes can be compared without rewriting history
Next step

Test the claim on your own system.

Plan a guided pilot