veritas
benchmark integrity auditor

Check whether frontier models are benchmark maxxing.

Veritas stress-tests reported scores across prompt wording, answer order, templates, repeated samples, and fresh data—then separates robust capability from protocol dependence, exposure-consistent signals, evaluator weakness, and uncertainty.

Behavioral evidence is not proof of contamination. Veritas reports what the evidence supports and marks unavailable tests as not run.

frontier model auditschema v2
canonical scorebaseline
validated variantspaired
fresh-set transferwhen available
exposure evidencecalibrated or not run
choose a model, benchmark, and request budget →
what Veritas maps

A benchmark score is the starting point. The audit asks which explanations survive controlled comparisons and keeps conflicting or missing evidence visible.

Prompt and template dependence

paired variants

Measure whether a score survives meaning-preserving wording, formatting, and instruction-template changes.

Choice and position sensitivity

mapping checked

Permute options and remap labels while preserving the correct answer and parent-item lineage.

Exposure-consistent behavior

never proof alone

Record exact, near-reference, likelihood, completion, and error-reproduction evidence only when requirements are met.

Capability that transfers

controls required

Use fresh, temporal, transformed, and distribution-shift controls to distinguish robustness from benchmark familiarity.

Built-in benchmark catalog

plus JSONL custom sets
MMLUMMLU-ProGPQAARC-ChallengeHellaSwagTruthfulQAHumanEvalSWE-bench Verified

Presets preserve upstream source, split, revision, license, access constraints, and deterministic sampling in provenance.

one reproducible workflow

Point Veritas at a benchmark and an API, local open-weight model, subprocess, or replayed response set. The same application service powers inspection, execution, and reporting.

  1. 01

    Evaluate the canonical set

    Capture every prompt, raw response, parsed answer, token count, latency, model setting, and cache key.

  2. 02

    Stress-test the score

    Apply seeded, validated prompt, template, choice-order, identifier, notation, and task-specific transformations.

  3. 03

    Map the evidence

    Compare paired scores with uncertainty, detector findings, alternatives, unavailable evidence, and an audit hash.

provider-neutral

HTTP APIs, vLLM, local Transformers, subprocesses, mocks, and replay

private-aware

external transfer requires an explicit acknowledgement and redacted provenance

resumable

content-addressed responses and staged checkpoints prevent duplicate inference

evidence-first

effect sizes, uncertainty, assumptions, alternatives, and not-run outcomes stay visible

Veritas does not convert behavioral anomalies into a contamination verdict. Results distinguish exact or semantic exposure evidence, protocol dependence, evaluator weakness, distribution shift, legitimate capability, unavailable evidence, and inconclusive outcomes. The original biological leakage auditor remains supported through the legacy report path.