VERITROOPER — veritrooper.com Patent pending — non-provisional patent filed 05/22/2026

Deploy Accuracy Anywhere

Your AI Never Tells You When It’s Wrong.

VERITROOPER Scout measures how accurately any AI answers from your own data, pinpoints exactly where and why it goes wrong, and hands your team the fixes — every contested verdict verified by a different vendor’s model, and the whole result sealed in a signed, independently verifiable evidence package.

Point Scout at any AI. It finds every confident wrong answer on your data — and proves it. Claude Opus 4.8 on 1,002 IRS tax-code questions: 93.01% → 99.20%, independently audited — and the same architecture holds across safety, drug-labeling, and financial filings.
Any Data.
Tax, safety, financial filings — every domain lifted.
Any AI.
From a 7B laptop model up to flagship frontier.
Results you can use.
Per-question diagnosis + plain-English fixes.
Independently checked.
Disputed answers reviewed by a second vendor, or by deterministic code when you run air-gapped.
EU AI Act evidence.
Annex IV and Article 15 testing evidence, on every run — nothing to switch on.
Hover or tap a point to see how VERITROOPER protects you.

Integrity, Honesty & Transparency — by Design.

A result is only worth as much as the process behind it. Every step that produces one is built to be defensible — to your auditors, your buyers, and your own engineers: independent cross-vendor verification, scoring that rounds against us, built-in hallucination traps, every failure shown in full, tamper-evident sign-off. Integrity here isn’t a claim — it’s the mechanism.

Don’t take our word for it.

Four domains. One result.

We didn’t measure this once and call it proof. VERITROOPER has been run end-to-end across four unrelated regulated worlds — U.S. tax code, OSHA workplace-safety regulation, FDA drug labeling, and SEC 10-K financial filings — each the same 1,000-question audit, the same seven models (a 7B on a gaming GPU up to flagship frontier), the same cross-vendor verification. One domain could be luck; four behaving the same way is a pattern. And the numbers are your model’s — VERITROOPER carries no score of its own, it inherits the model’s floor and ceiling. Frontier models land in the high 90s; the weaker the model, the more accuracy the audit recovers. Llama 3.1 70B’s honest 80.84 on FDA drug labeling is the proof, not an outlier — a real audit has to be able to return a low number when the model earns one.

IRS Tax Code

Federal income-tax regulations
ModelBase­lineAuditedΔ
Claude Opus 4.893.0199.20+6.19
GPT-5.591.0299.20+8.18
Gemini 2.5 Pro92.7197.11+4.40
Qwen 2.5 72B86.3398.20+11.87
Llama 3.1 70B86.0397.41+11.38
Gemma 3 27B88.6397.99+9.36
Qwen 2.5 7B86.1294.87+8.75

OSHA Safety

29 CFR — general industry, construction, hazmat
ModelBase­lineAuditedΔ
Claude Opus 4.894.8999.00+4.11
GPT-5.592.4098.80+6.40
Gemini 2.5 Pro93.2098.00+4.80
Qwen 2.5 72B93.2098.30+5.10
Llama 3.1 70B79.9095.80+15.90
Gemma 3 27B89.5098.40+8.90
Qwen 2.5 7B89.6097.00+7.40

FDA Drug Labels

High-alert & common prescription drugs
ModelBase­lineAuditedΔ
Claude Opus 4.893.4198.60+5.19
GPT-5.590.0298.20+8.18
Gemini 2.5 Pro90.9296.81+5.89
Qwen 2.5 72B91.3298.30+6.98
Llama 3.1 70B69.2680.84+11.58
Gemma 3 27B87.1396.81+9.68
Qwen 2.5 7B87.2394.51+7.28

SEC 10-K Filings

Apple, NVIDIA, JPMorgan, Coca-Cola — financial filings
ModelBase­lineAuditedΔ
Claude Opus 4.886.6996.30+9.61
GPT-5.582.6096.10+13.50
Gemini 2.5 Pro88.1097.40+9.30
Qwen 2.5 72B89.5394.76+5.23
Llama 3.1 70B87.0895.52+8.44
Gemma 3 27B82.7994.96+12.17
Qwen 2.5 7B85.6590.89+5.24

Baseline = the model with vanilla-RAG retrieval (BM25 top-5), the way real deployments serve. Audited = the same model’s same answers, re-measured against the correct source evidence, every contested verdict verified by a different vendor’s model. Δ = the accuracy your model is leaving on the table — every point of it traced to a specific question and root cause your team can fix. Every figure reproducible from timestamped logs.

See the full audit & download the run records → See sample output →

Run it again. The wording changes. The verdict doesn’t.

Reproducibility is the line between a measurement and a guess. So we ran the exact same audit ten times in a row — same 238 questions, same model, same machine, no reused answers. A language model almost never writes the same sentence twice, and this one didn’t. The audit still landed in the same place. Here is every run, untouched.

Ten runs, back to back

RunAuditedBaselineΔ
197.48%93.70%+3.78
297.48%93.70%+3.78
397.48%94.54%+2.94
497.06%93.70%+3.36
597.90%93.70%+4.20
697.48%94.12%+3.36
797.48%94.96%+2.52
897.06%92.86%+4.20
997.90%93.70%+4.20
1097.48%94.96%+2.52

What that adds up to

 Across all ten
Audited score97.06% – 97.90%
Most common result97.48% — 6 of 10 runs
Answers the model reworded125 of 238
Same verdict in all ten runs236 of 238 — 99.2%
Same answer graded two waysNever — 0 in 4,760

The spread you see is the model’s, not the audit’s. Even at low temperature the model rewrote most of its own answers between runs — only 113 of the 238 came back word-for-word identical all ten times. The audit did not follow it around: across 4,760 graded answers, the same answer was never graded two different ways. Two questions out of 238 changed verdict at all, and both changed because the model gave a materially different answer, not because the audit changed its mind. An LLM auditor still writes the per-question diagnosis, but only a code verdict is allowed to move the headline — which is why the score moves with the model and nothing else. The 0.84-point spread sits well inside the measurement’s own margin of error (±1.99 points at this sample size), so the ten runs are indistinguishable as measurements.

How it works.

Raw data in, detailed after-action out. VERITROOPER's safety parachute pinpoints where your LLM fails on your data — in plain English you can hand to an engineer. Hover a stage.

What is VERITROOPER.

VERITROOPER is an accuracy engine for any AI working over your own data — it measures how accurately the model answers, proves it, and tells you exactly what to fix. It doesn’t replace your AI or change how it serves; it grades it against ground truth and hands back the evidence. You hand it a data set — tax code, safety regulations, drug labeling, financial filings, gaming rules, anything written down — and it generates a question set with verified ground truth, then runs the model two ways: once on realistic retrieval (the baseline) and once with the correct source evidence in front of it (the audit). The gap between them is the accuracy your model is leaving on the table. Failures are routed to specialist diagnostic modules, and every contested verdict is confirmed by an independent cross-vendor verifier that can override it — the model under test never gets the final say on its own answers. The run ends in a plain-English report on what the model got wrong, why, and exactly what to fix to recover it.

On regulated material it does what naive retrieval can’t: when a question depends on a cross-referenced section, the pipeline pulls that referenced section’s text into the evidence — multi-hop resolution the baseline can’t do. For financial filings it parses label/value/period tables and checks every calculation against the source numbers with a built-in financial calculator, so a fabricated figure or a wrong-year value gets caught.

The failure mode it's built to catch is hallucination — when an LLM confidently produces a wrong answer. Models don't crash when they hallucinate; there's no flag, no warning, no error code. They just sound certain about something that isn't true. That's what breaks naked LLM deployment in any setting where the answer actually matters. VERITROOPER catches it across any domain, with any model.

What you get out the other side: a final verified accuracy score against ground truth — with a confidence interval and a clear PASS / CONDITIONAL / FAIL disposition against an acceptance standard you set, not a naked percentage — a per-question categorized list of every failure (with patterns and clusters identified), a per-category and per-data-source accuracy breakdown, a failure-recovery rate (the share of the model’s baseline failures the audit recovered), and concrete engineering recommendations that would close the specific gaps the LLM showed. It all lands as one portable evidence package: role-specific documents plus a machine-readable record, sealed with a cryptographic signature and a trusted timestamp and shipped with a standalone verifier, so a reviewer can prove nothing changed after sign-off. Every verdict is reproducible from timestamped logs — no black box.

Deploying into Europe? One toggle adds the EU AI Act conformity evidence to the same run: declared accuracy and robustness testing (Article 15), a drop-in Annex IV technical-documentation record, recurring accuracy-drift monitoring (Article 72), a dated, signed human-review audit trail (Article 14), and a per-category gap diagnostic (Article 10). VERITROOPER produces the evidence a conformity file relies on — it does not replace the provider's conformity assessment or confer compliance.

Four regulated-domain data sets tested across seven different LLMs, phone-tier 7B to flagship frontier. Measured against the correct evidence with cross-vendor verification, every model — from a $2,000 RTX 4090 running a 7B up to flagship frontier APIs — lands in the 80.8–99.2% band on the same 1,000 IRS Tax Code questions. That's the punchline: this isn't a frontier-only luxury. Even a 7B on a laptop comes within shouting distance of a frontier model when grounding is solved — so the audit shows the gap on your data is mostly recoverable, and pinpoints exactly which failures to fix to close it.

For the story behind Scout and the cast, visit the home page.

Get in touch.

For enterprise pilots, technical evaluation, partnerships, and licensing.

contact@veritrooper.com

Enterprise pilots: the best way to see what Scout finds is to point it at your own AI. Scout ships as a self-contained app — it runs on your hardware, inside your network, against your model and your data; nothing has to leave your environment. You approve the question set; Scout returns the full sealed evidence package: every wrong answer with its source evidence, a per-category accuracy breakdown, the failure-recovery rate, a concrete plain-English fix list your engineers can act on the same day, and a signed, independently verifiable record — plus optional EU AI Act conformity evidence in one toggle. Low-lift on your side, and the result is on data you already trust.

Request a guided pilot →   or run the free trial yourself →

The company: VERITROOPER is a registered Delaware LLC in good standing that owns the patent application and the codebase outright, with clean, assigned title. More about the company →

Public results, sample run records, and the methodology need no NDA. Raw logs, the full dataset, and the patent package are shared under NDA. Live walkthroughs by request.

Download Technical Proof Packet (PDF) →