Prompt and template dependence
paired variantsMeasure whether a score survives meaning-preserving wording, formatting, and instruction-template changes.
Veritas stress-tests reported scores across prompt wording, answer order, templates, repeated samples, and fresh data—then separates robust capability from protocol dependence, exposure-consistent signals, evaluator weakness, and uncertainty.
Behavioral evidence is not proof of contamination. Veritas reports what the evidence supports and marks unavailable tests as not run.
A benchmark score is the starting point. The audit asks which explanations survive controlled comparisons and keeps conflicting or missing evidence visible.
Measure whether a score survives meaning-preserving wording, formatting, and instruction-template changes.
Permute options and remap labels while preserving the correct answer and parent-item lineage.
Record exact, near-reference, likelihood, completion, and error-reproduction evidence only when requirements are met.
Use fresh, temporal, transformed, and distribution-shift controls to distinguish robustness from benchmark familiarity.
Presets preserve upstream source, split, revision, license, access constraints, and deterministic sampling in provenance.
Point Veritas at a benchmark and an API, local open-weight model, subprocess, or replayed response set. The same application service powers inspection, execution, and reporting.
Capture every prompt, raw response, parsed answer, token count, latency, model setting, and cache key.
Apply seeded, validated prompt, template, choice-order, identifier, notation, and task-specific transformations.
Compare paired scores with uncertainty, detector findings, alternatives, unavailable evidence, and an audit hash.
HTTP APIs, vLLM, local Transformers, subprocesses, mocks, and replay
external transfer requires an explicit acknowledgement and redacted provenance
content-addressed responses and staged checkpoints prevent duplicate inference
effect sizes, uncertainty, assumptions, alternatives, and not-run outcomes stay visible
Veritas does not convert behavioral anomalies into a contamination verdict. Results distinguish exact or semantic exposure evidence, protocol dependence, evaluator weakness, distribution shift, legitimate capability, unavailable evidence, and inconclusive outcomes. The original biological leakage auditor remains supported through the legacy report path.