Artificial Analysis published research introducing three diagnostic probes to quantify benchmark optimization and fitting in automatic speech recognition models across VoxPopuli and LibriSpeech datasets.
Aug 21, 2026
10d agoKey Details
- Evaluated 11 open-source automatic speech recognition models using consensus disagreement, masked entity retrieval, and orthographic switching probes
- Observed that six of 11 models reproduced erroneous reference transcripts from VoxPopuli rather than faithfully transcribing the actual spoken audio
- Demonstrated that top-scoring models on LibriSpeech reproduced silenced numbers in roughly 30% to 40% of test examples despite the audio being removed
- Found models rely on acoustic cues to select benchmark-specific spelling conventions, exceeding random choice with up to 90% switch accuracy
- Introduced a Benchmark fitting tab to the Open-ASR Leaderboard to identify models over-optimized for public benchmarks