Actual Scout output · sealed public sample

Portable evidence for independent review.

Inspect the recorded questions, answers, findings, and evaluation scope. Download the complete sample and check its signed files without a VeriTrooper account, installation, or network connection.

These reports come from the 26 August 2026 Scout run of gpt-5.6-sol against public SEC 10-K material: 88.08% with VeriTrooper versus a 72.89% vanilla-RAG baseline across 889 scored questions. This is a scoped result, not a general model rating.

889 scored questions+15.19 percentage points12 signed payload filesOffline checker included
Executive summaryLeadership · PDF
Open separately ↗
Public, minimized, and verifiable

Evidence you can retain, share, and inspect independently.

The ZIP includes seven customer-facing reports, a minimized run record, three reviewer aids, a signed manifest, signer certificate, fingerprint, and standalone verification tools. It omits credentials, protected implementation material, full configuration, and internal logs.

Integrity verification proves that the listed files match the signed manifest. It does not prove substantive correctness or compliance. Confirm the published fingerprint through a separate trusted channel.

VERIFIED SAMPLE/├── READ ME FIRST.txt├── Audit_Report.pdf├── Individual files/├── Transparency/All_QA_Pairs.pdf├── System files/public_sample_run_record.json└── Integrity/   ├── manifest + signature   ├── signer certificate + fingerprint   └── offline checker
Recorded evaluation history

Documented repeatability. Traceable results.

Our evaluation archive preserves months of results, failures, corrections, and retesting. Two ten-run studies use a 238-question panel drawn from the Uniform Code of Military Justice (UCMJ). The study traces each result to its saved run record and explains the scoring basis.

July 25, 2026 · Adopted scoring basis

236 of 238 verdicts stayed the same.

The same recorded pipeline verdict appears in all ten runs for 236 question IDs, or 99.16% of the panel.

  • 97.48% mean pipeline accuracy, with a 97.06% to 97.90% range.
  • Matching saved configurations and sample manifests across the ten runs.
  • Documented scoring corrections and pre-correction records preserved.
May 26, 2026 · Original two-host study

97.48% raw sweep accuracy in every run.

The original study records ten runs, with five on each of two reported hosts.

  • The raw pipeline sweep metric is identical across all ten runs.
  • The historical adjusted headline used a different, subsequently superseded scoring rule.
  • Original summary files, configurations, and individual results retained.

These are internal evaluations of a fixed question panel under recorded conditions. The PDF separates the two studies and includes all twenty run references, the July correction history, and the limits of the findings. Repeatability alone does not establish correctness or performance on other workloads.

3 pages · Prepared 22 September 2026 · Supporting preserved records can be discussed during an evidence review.

Your system is the real test

Run the same evidence discipline on one important workflow.

The 10-business-day assessment uses your approved sources, boundary, and acceptance test.

Request an assessment