Voice Benchmarks Hear, but Don’t Listen
By Roman Smirnov and Garreth Lee
There’s a new benchmark released for voice agents and speech-to-speech models every other week. They look different on the surface, but underneath they all feature the same four dimensions:
- The voice input what the model actually hears
- The scenarios and tasks what the model has to accomplish
- The harness what runs the interaction and tracks state
- The metrics how a model’s performance is scored
At Mundo, we do applied research on the frontier of multimodal AI, particularly on ways we can improve the evaluation of speech-to-speech interactions. Designing our own benchmark meant that we had to take a hard look at how the current ones are put together – starting with the most fundamental of those four dimensions: the input data being fed into the models.
A core design decision we analyzed is the choice of human vs synthetic (generated) data to evaluate these models. It is intuitive to prefer human data as it is closer to the distribution of downstream tasks (after all, humans are the end user, not other voice models - for now at least!), but human data is difficult to collect at scale and difficult to collect in advance for dynamic environments that voice agents are evaluated against. So modern benchmarks lean on synthetic data, which is cheaper to scale and faster to iterate on.
This convenience, however, rests on the assumption that synthetic and human speech are similar enough to be interchangeable. (Otherwise, why benchmark with non-human voices at all?) We compared recorded human speech with audio from six commercial TTS systems and tested that assumption from three angles:
- How different are human and synthetic speech? Can a simple classifier tell human and synthetic speech apart, and what drives the separation?
- What noise generators add, and what they don't What do the noise generators that many benchmarks today ship with actually change in the recorded audio?
- Does the difference change a model's output? Does a model that sees the difference between human and synthetic inputs produce different outputs?
How different are human and synthetic speech?
How we compared human and synthetic speech
We built a corpus of human speech – taken from our proprietary datasets and recording samples from Full-Duplex-Bench v3[1] – and synthetic speech, which is generated from the same transcripts of the human speech (doing this ensured that any difference is from the audio itself, not the spoken content) from six commercial TTS systems (ElevenLabs, Gemini 2.5 Pro Preview TTS, Gemini 3.1 Flash TTS Preview, GPT-4o mini TTS, Grok TTS, and Qwen3-TTS)[2] sampling a random voice each time and a set of “designed” voices for extra accent coverage. In total, there were around 900 clips from 19 human speakers and 780 clips from 88 default voices and 15 designed voices.1
To ask whether human and synthetic speech come from the same distribution, we used a Classifier Two-Sample test[3]: if a classifier can separate the two types of speech better than chance, they aren’t from the same distribution. We trained a classifier against the labels, ensuring that the training and test splits are shuffled and grouped by voices to avoid cheating by memorizing speaker identity.
We used three approaches to prepare features, chosen to cover the various types of audio representation:
Hand-crafted statistics (MFCC)
Each of the frequency data sequences was compacted to mean and standard deviation features only with the following PCA projection to 100 principal components.
Gemma 4 E2B Instruct embeddings
Gemma 4 E2B is a model that is capable of audio understanding, it is using dense projections through the USM audio encoder and projection layer to insert to the LLM (Gemma) context. We applied the same encoding and projections using the weights from the original Gemma to receive the embeddings the model actually uses as the input for audio. Sequence of multidimensional embeddings was compacted to mean and standard deviation for each of the embedding dimensions with the following PCA projection to 100 principal components.
Qwen3-TTS discrete audio tokens (Qwen3-TTS)
Similar to Gemma, Qwen is capable of audio processing, but it is using sparse discrete audio tokens that are embedded in RVQ manner. We applied the same approach the original Qwen is using to encode audio into the LLM embeddings. Sequence of multidimensional embeddings was compacted to mean and standard deviation for each of the embedding dimensions with the following PCA projection to 100 principal components.
For each, we summarized every clip into a fixed-length vector and reduced it with PCA.2
Running the test with these three feature extraction approaches gave the following results:
| MFCC | Gemma 4 E2B | Qwen3-TTS | |
|---|---|---|---|
| AUC | 0.83 | 0.99 | 0.97 |
How do we know the classifier is finding a real difference and not a quirk of our setup? We ran the same test again with the answer key scrambled. Every speaker's clips were tagged human or synthetic at random, so there was nothing true left to learn, which means that the classifier should score around 0.5 AUC, which is the same as guessing. We got exactly this score with the same setup as before, only with the labels scrambled. This supports the finding that the classifier learned a difference in the recordings. The checks below examine which properties of those recordings could explain it.
The difference is visible even when we use only high-level audio statistics as features, though it becomes much more pronounced with the audio embeddings. This suggests that real human speech and synthetic speech come from measurably different distributions, and that the audio encoders used by modern models make this difference even easier to detect.
Where the difference comes from
This result invites two obvious objections. First, perhaps the classifier is simply picking up artifacts from the distinct recording conditions of our proprietary dataset. Second, Qwen3-TTS is both one of our synthetic data generators and feature extractors, so perhaps the classifier trained on Qwen features is simply recognizing its own output. We ran two checks to test these explanations, followed by a third to investigate where the remaining difference might come from. The chart below shows each check beside its result.
| All data | Without proprietary data | Without Qwen3-TTS clips | Human vs human | |
|---|---|---|---|---|
| MFCC | 0.83 | 0.71 | — | 0.97 |
| Gemma 4 E2B | 0.99 | 0.99 | — | 0.92 |
| Qwen3-TTS | 0.97 | 0.98 | 0.99 | 0.98 |
A rotating three-dimensional scatter plot of the same 8,410 clips coloured by which TTS system produced them, with the human recordings in blue. The six synthetic groups overlap one another and remain offset from the human recordings, with some overlap.
While the human to human experiment has fewer than 20 speakers in total and should be treated as exploratory, it points to an important source of variation. This suggests that some of the separation comes from the recording conditions themselves. Change the speaker, microphone, room, or recording setup, and the resulting audio changes too. Synthetic speech does not naturally contain all of this variation. Therefore, recording/call conditions are an important part of the input distribution that a benchmark needs to represent.
What noise generators add, and what they don't
Noise increases overlap, but a difference remains
The last check suggests that recording conditions (noise, artifacts, and the general messiness of real-world audio) may explain some of the difference between human and synthetic speech. We wanted to test how much of the remaining gap could be explained by noise. The human recordings contain background noise and recording artifacts that the synthetic clips do not. Here, “clean” means the original clip before we add benchmark noise; it does not mean a noise-free recording. If these differences are driving the classifier, adding similar degradation to the synthetic clips should reduce its ability to separate the two.
Conveniently, most benchmarks today have noise generators integrated within to test robustness. So we took three real pipelines and applied them to both human and synthetic clips, then re-ran the test:
- τ-voice[4] (code) One of the most complex methods: it includes 4 continuous and 5 burst noises, and a set of channel degradation techniques, including frame drops and a telephony effect. However, it uses a very limited collection of noise sounds.
- EVA-Bench[5] (code) 7 continuous noises and channel degradation only.
- DeepEval background mixer (code) Noises taken from ESC-50[6], a collection of 2,000 environmental recordings.
Long story short: adding noise increased the overlap between human and synthetic speech, but there's still a fundamental distributional difference. We ran two comparisons. First, we added noise to both the human clips and the synthetic clips. All three feature sets still told them apart. Second, we compared noised synthetic clips against clean human clips. This tests whether adding noise makes synthetic speech match the human recordings. The classifier still told them apart.
| Noised synthetic vs noised human | Noised synthetic vs clean human | |
|---|---|---|
| MFCC | 0.79 | 0.87 |
| Gemma 4 E2B | 0.99 | > 0.99 |
| Qwen3-TTS | 0.93 | 0.98 |
A rotating three-dimensional scatter plot of 8,410 clips. The 4,510 human clips and the 3,900 synthetic clips form two overlapping but clearly offset clouds.
Here is the same view for MFCC features (Qwen embeddings gave a similar picture). The colors show the different noising approaches:3
A rotating three-dimensional scatter plot of 8,410 clips coloured by noising approach. The τ-voice points sit furthest from the clean points, and the τ-voice points with the telephony effect form a separate arm.
The analysis and images both tell the same story: adding noise increases the overlap between human and synthetic speech, but there’s still a distributional difference between them.
That leaves an important question: does any of this change what a model does with the audio?
Does the difference change a model's output?
Measuring changes in model output
So far we have looked at the audio as a model represents it. The Gemma 4 E2B embeddings separated human from synthetic speech almost perfectly, which tells us the encoder can see the difference, but it doesn't tell us whether the model acts on it. To find out, we needed a task where the same model produces an output for every clip and we can score that output without a scenario or a harness in the way.
Transcription is the simplest such task. Since it's the step closest to the raw audio and every clip has a known script, transcription errors are the first place an input problem shows up. We asked Gemma 4 E2B Instruct to transcribe every clip, clean and noised, and measured how much the transcript changed after noise was added. We use character error rate (CER) against the clean transcript to measure how noise affects the model's output.
For each human recording, we generated a synthetic version from the same transcript, then applied every configuration of each noise pipeline. For ESC-50, which contains a larger set of possible noises, we sampled 10 noises per input.
- One transcript
Speech
- Human recording
- Synthetic clip
Noise pipelines
- τ-voice
- EVA-Bench
- DeepEval + ESC-50
- Clean audio
- Add noise
- Gemma 4 E2B Instruct transcript
- CER against Gemma’s clean-audio transcript
Alongside the source and noise pipeline, we measured loudness, band energy, and SNR across five frequency ranges covering voicing, vowels, consonants, and sibilants. We trained a CatBoost[7] regressor to predict CER from these features, making sure to run cross-validation grouped by each speaker and clean file.
Ultimately, the type of noise pipeline barely mattered. The strongest predictor was the loudness of the clean speech, which makes sense: with the same amount of noise, quieter speech has a worse signal-to-noise ratio. But the next strongest signal was whether or not the speech was human or synthetic. By SHAP importance[8], this ranked 2nd out of 14 features, while the noise pipeline ranked 13th.
All human speakers
- Loudness of the clean speech: 0.125
- Human or synthetic: 0.107
- SNR, consonant band: 0.091
- Loudness of the noised clip: 0.065
- SNR, voicing band: 0.058
- SNR, lower vowel band: 0.057
- SNR, 1,000–2,000 Hz band: 0.049
- Noise loudness: 0.042
- SNR, sibilant band: 0.026
- Overall SNR: 0.016
- Burst noise: 0.015
- Channel degradation: 0.006
- Noise pipeline: 0.005
- TTS provider: 0.000
A strange group of speakers
The first SHAP graph also pointed us toward something unexpected in the human data.
A subset of human speakers behaved differently from everyone else. The difference was not obvious from high-level audio statistics. Instead, these speakers had unusually low energy between 1,000 and 2,000 Hz, one of the frequency ranges associated with vowels.
| Series | Voicing | Vowels, low | 1,000–2,000 Hz | Consonants | Sibilants | Mean CER after noising |
|---|---|---|---|---|---|---|
| Human | 0.4002 | 0.4381 | 0.0113 | 0.017 | 0.1334 | 0.316 |
| Human | 0.2473 | 0.5351 | 0.1596 | 0.0338 | 0.0242 | 0.017 |
| Human | 0.1596 | 0.8338 | 0.004 | 0.0015 | 0.0012 | 0.284 |
| Human | 0.1861 | 0.5881 | 0.17 | 0.0379 | 0.0179 | 0.022 |
| Human | 0.1707 | 0.6664 | 0.1157 | 0.0337 | 0.0135 | 0.03 |
| Human | 0.23 | 0.7154 | 0.0377 | 0.0102 | 0.0067 | 0.048 |
| Human | 0.1738 | 0.8167 | 0.0037 | 0.0032 | 0.0026 | 0.312 |
| Human | 0.0699 | 0.9161 | 0.0094 | 0.0039 | 0.0007 | 0.441 |
| Human | 0.2572 | 0.6825 | 0.0543 | 0.0034 | 0.0026 | 0.085 |
| Human | 0.4257 | 0.4996 | 0.0576 | 0.0148 | 0.0024 | 0.029 |
| Synthetic | 0.2879 | 0.5624 | 0.1127 | 0.0265 | 0.0105 | 0.066 |
| Synthetic | 0.0005 | 0.1652 | 0.7075 | 0.1268 | 0 | 0.042 |
| Synthetic | 0.0004 | 0.2277 | 0.6203 | 0.1515 | 0 | 0.06 |
| Synthetic | 0.2083 | 0.605 | 0.1687 | 0.0139 | 0.0042 | 0.014 |
| Synthetic | 0.0007 | 0.2343 | 0.6936 | 0.0714 | 0 | 0.05 |
| Synthetic | 0.0003 | 0.1162 | 0.8536 | 0.0298 | 0 | 0.225 |
| Synthetic | 0.7376 | 0.1656 | 0.0519 | 0.0115 | 0.0333 | 0.004 |
| Synthetic | 0.1054 | 0.5075 | 0.3047 | 0.0752 | 0.0071 | 0.003 |
| Synthetic | 0.0009 | 0.137 | 0.7272 | 0.1347 | 0.0003 | 0.016 |
| Synthetic | 0.4848 | 0.3671 | 0.0842 | 0.0188 | 0.045 | 0.034 |
| Synthetic | 0.2894 | 0.4514 | 0.1857 | 0.0561 | 0.0173 | 0.018 |
| Synthetic | 0.7031 | 0.2376 | 0.0224 | 0.0133 | 0.0236 | 0.012 |
| Synthetic | 0.4302 | 0.4647 | 0.075 | 0.0157 | 0.0143 | 0.011 |
| Synthetic | 0.6814 | 0.267 | 0.0384 | 0.0059 | 0.0073 | 0.016 |
| Synthetic | 0.3912 | 0.5184 | 0.0746 | 0.0101 | 0.0057 | 0.043 |
| Synthetic | 0.7217 | 0.2514 | 0.0142 | 0.0104 | 0.0022 | 0.046 |
| Synthetic | 0.4087 | 0.5047 | 0.069 | 0.0121 | 0.0055 | 0.081 |
| Synthetic | 0.3357 | 0.5143 | 0.1162 | 0.016 | 0.0176 | 0.014 |
| Synthetic | 0.3195 | 0.5661 | 0.0682 | 0.0376 | 0.0086 | 0.007 |
| Synthetic | 0.4761 | 0.4647 | 0.0466 | 0.0051 | 0.0074 | 0.024 |
Because we found this group after looking at the first results, we treat this analysis as exploratory. But removing them—roughly a quarter of the human speakers—changed the feature importance substantially: whether the speech was human or synthetic fell from 2nd to 9th out of 14 features, while the per-band SNR features moved up.
All human speakers
- Loudness of the clean speech: 0.125
- Human or synthetic: 0.107
- SNR, consonant band: 0.091
- Loudness of the noised clip: 0.065
- SNR, voicing band: 0.058
- SNR, lower vowel band: 0.057
- SNR, 1,000–2,000 Hz band: 0.049
- Noise loudness: 0.042
- SNR, sibilant band: 0.026
- Overall SNR: 0.016
- Burst noise: 0.015
- Channel degradation: 0.006
- Noise pipeline: 0.005
- TTS provider: 0.000
Unusual speakers removed
- Loudness of the clean speech: 0.290
- SNR, lower vowel band: 0.137
- SNR, voicing band: 0.116
- SNR, sibilant band: 0.112
- Loudness of the noised clip: 0.080
- SNR, 1,000–2,000 Hz band: 0.054
- Noise loudness: 0.045
- SNR, consonant band: 0.044
- Human or synthetic: 0.039
- Overall SNR: 0.031
- Burst noise: 0.023
- Channel degradation: 0.016
- Noise pipeline: 0.003
- TTS provider: 0.000
This gives us another explanation for the gap. Some of the human speakers had frequency characteristics that were not present in the synthetic speech, and removing those speakers made the human/synthetic label considerably less important to the model.
Human speech varies along many axes at once: of course the noise profile plays a part but so do accents, recording setups, the frequency shape of individual voices, and many more. The population of speakers here wasn't merely noisier, but had entirely different spectral profiles, which is crucial for speech benchmarks in order to adequately represent the full diversity of human speech.
Noise can also be used against the model
While modeling changes in Gemma's transcripts after adding noise, we found that one feature was particularly predictive: the similarity between the model's representation of the clean audio and its representation of the noised audio. The further the noised audio moved from the original in Gemma's embedding space, the larger the change from its clean-audio transcript tended to be.
That gave us a simple experiment: instead of adding noise at random, could we choose noise that deliberately pushes the model's representation away from the original?
We started with a short clip from Full-Duplex-Bench v3 of someone saying:
“A B C 1 2 3”
We then optimized a noise vector to move Gemma's audio embedding as far as possible from the original, while limiting the noise to 3% of the RMS of the speech.
Original
- Transcript
- “A B C 1 2 3”
The speech still sounds somewhat similar, but Gemma heard something else. This type of adversarial audio is well established, so the result itself is not surprising.[9]
The interesting part was making the constraint realistic. Rather than optimizing an arbitrary noise vector, we limited the search to environmental sounds from ESC-50. We used a straight-through estimator[10] to optimize which sounds to include while keeping the final selection discrete.
The resulting audio was transcribed by Gemma as:
“A C 4.2 grade.”
Original
- Human
- “A B C 1 2 3”
ESC-50 optimized noise
- Human
- “A B C 1 2 3”
- Gemma
- “A C 4.2 grade.”
The final perturbation was a mixture of 50 ESC-50 recordings at equal weight. A member of our team listened to the resulting clip and still heard the original phrase clearly.
These experiments are preliminary: we tested only a small number of examples and haven’t tested whether the effect holds across different clips, speakers, noise constraints, or models. But this gives us something useful to test next. Two noises at the same level may not be equally difficult for a model. When evaluating robustness, we therefore care about both the amount of degradation and the particular noise being applied.
What a reliable input module needs
This “input” dimension in benchmarks decides what everything downstream hears. Get it wrong and every error propagates, so it is important to check our assumptions and make sure they are accounted for. From these results, a reliable input module needs:
- Recorded human speech from many speakers and many recording setups, with coverage that is documented.
- Per-band SNR and energy distribution for each clip, since this is where human and synthetic speech differ.
- Noise as a measured variable. Run the same audio clean and degraded, and report how performance moves across different signal-to-noise levels.
- Real channel conditions, such as telephony codecs and bandwidth limits, not only resampling.
This post covers one angle on the input dimension: whether the audio that reaches the model looks like recorded human speech, and how noise generators that modern benchmarks ship with change it. It is not a full audit of where a benchmark's inputs can fail. Transcript content, turn-taking, language and accent coverage, and how prompts are scripted are all factors that can affect this dimension, and we did not test them here. We leave this for future work.
This is the first of four parts. Next, we'll look at scenarios and tasks, harnesses, and metrics. For each, we'll ask the same question: what failures does the benchmark capture, and what does it miss?
We want evaluations to reflect what models actually encounter in the real world. If you’re building speech or multimodal models and want to evaluate them on real human data, work with us. If you want to help us build better datasets, evaluations, and benchmarks, we’re hiring.
Footnotes
- The exact counts are 902 human clips and 780 synthetic clips. Each clip also has four noised copies, one per noise configuration, so the point clouds below show 8,410 points.
- The point clouds below show the first three principal components. They explain 23.6% of the variance for the Gemma embeddings and 61.3% for the MFCC features.
- We ran the τ-voice pipeline twice, with and without its telephony effect, so the legend lists the two runs separately.
References
- Lin, G.-T., Chen, C., Chen, Z., Lee, H.-y., “Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency”, 2026. Link
- Hu, H., Zhu, X., He, T., Guo, D., Zhang, B., Wang, X., Guo, Z., Jiang, Z., Hao, H., Guo, Z., Zhang, X., Zhang, P., Yang, B., Xu, J., Zhou, J., Lin, J., “Qwen3-TTS Technical Report”, 2026. Link
- Lopez-Paz, D., Oquab, M., “Revisiting Classifier Two-Sample Tests”, 2016. Link
- Ray, S., Dhandhania, K., Barres, V., Narasimhan, K., “τ-Voice: Benchmarking Full-Duplex Voice Agents on Real-World Domains”, 2026. Link
- Bogavelli, T., Gauthier Melançon, G., Stankiewicz, K., Bamgbose, O., Riols, F., Nguyen, H. H., Mehndiratta, R., Brin, L. D., Marinier, J., Subramani, H., Madamala, A., Nemala, S. K., Sunkara, S., “EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents”, 2026. Link
- Piczak, K. J., “ESC: Dataset for Environmental Sound Classification”, 2015. Link
- Prokhorenkova, L., Gusev, G., Vorobev, A., Dorogush, A. V., Gulin, A., “CatBoost: unbiased boosting with categorical features”, 2018. Link
- Lundberg, S. M., Lee, S.-I., “A Unified Approach to Interpreting Model Predictions”, 2017. Link
- Carlini, N., Wagner, D., “Audio Adversarial Examples: Targeted Attacks on Speech-to-Text”, 2018. Link
- Bengio, Y., Léonard, N., Courville, A., “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation”, 2013. Link
Citation
Smirnov and Lee, “Voice Benchmarks Hear, but Don’t Listen”, Mundo AI, 2026.
@misc{smirnov2026benchmarks,
title={Voice Benchmarks Hear, but Don’t Listen},
author={Roman Smirnov and Garreth Lee},
year={2026},
publisher={Mundo AI},
url={https://mundoai.world/research/benchmarks-dont-hear-humans},
}