SPAR Spring 2026
Test smells: a causal decomposition of eval-identifiability in LLM benchmarks
Research Fellow · Advised by Qiyao Wei
Which benchmark cues make a model recognize an evaluation? I built a cue bank and paired experiments that inject or remove these “test smells,” separating removable formatting signals from signals intrinsic to the task.
Study figures below are preliminary and use 258 templates. The linked public dataset card documents 255 templates and ordinal judgment scores; a versioned report reconciling the release and study figures is not yet linked.
- 258
- templates in the preliminary study analysis (6 benchmarks)
- +0.192
- causal lift in a judge's eval-probability from injected cues (11/11 cells)
- 7.9×
- larger effect than matched neutral rewrites on capability benchmarks
- 89.8%
- BBH task accuracy preserved after removing the cues
Read full abstractHide full abstract for Test smells: a causal decomposition of eval-identifiability in LLM benchmarks
Frontier models behave differently when they sense they are being evaluated, which makes safety and capability numbers hard to trust. I ask a narrower, causal question: which surface features of a transcript actually move an LLM judge's “is this an eval?” belief, and can they be removed without changing the underlying task? Borrowing a term from software engineering, I call these features “test smells” — answer markers, formatting and style constraints, item identifiers, scoring rubrics, fictional placeholders, and explicit harm probes — and extract a bank of 258 canonical templates from six widely used benchmarks (Needham, Hawthorne, MASK, AgentHarm, BBH, MALT), grouped into six categories.
The result is decompositional. Forward injection — inserting bank cues into deployment-style transcripts — raises the judge's eval-classification by a mean of +0.21 across all 11 cross-benchmark pairs; under matched paraphrase controls the bank-specific causal shift is +0.192, positive in 11/11 cells. Reverse intervention behind a five-criterion preservation gate lowers eval-probability by -0.391 on the source pool and by -0.176 (p < 10⁻¹³) on a never-seen held-out pool, while preserving the task (92% behavioral parity, 89.8% BBH accuracy, 80.8–100% refusal-class preservation on two safety benchmarks). Neutral-rewrite controls localize the effect: on capability benchmarks the cues move the judge 7.9× more than matched neutral edits, whereas on safety benchmarks the signal is more entangled with the task itself. Activation probes across six open-weight models support the direction, but only when pooled across benchmarks. The central claim is that eval-identifiability has a removable surface component plus a large intrinsic one, and that cues do not behave like independent additive features. I release the template bank, paired forward/reverse datasets, and analysis code. A preprint is forthcoming.
