The Naturalness Problem
By Simon Williams and Garreth Lee
We recently partnered with Sesame and others on TurnBench, a benchmark for measuring turn-taking in spoken dialogue. Our team recorded and annotated the conversations behind it, while also defining the taxonomy of conversational events the benchmark scores, including turns, interruptions, backchannels, and mid-turn pauses. The dataset includes conversations performed by 106 professional voice actors in 53 pairings across six types of conversation, including casual talk, argument, and storytelling.
Throughout the recording process, we found that conversational data can meet every technical requirement and still feel off. At first, the task seemed straightforward: recruit voice actors, record their conversations, and add the labels. But making those conversations feel authentic was much harder. Simply telling people to “be natural” or “just argue” doesn’t work. The more directly you specify the behaviour you want to capture, the more that specification shapes how voice actors produce it, often making the resulting conversation feel less natural.
How specification changes the data
- Specify the behaviour you want. Example: Act angry for an hour.
- Voice actors act out the instruction. Example: Raise your voice and deliver insults.
- The conversation stops feeling natural. Example: The anger sounds performed, not felt.
Two recurring mistakes shaped the lessons that follow: prescribing and isolating.
The prescription problem
At first, our initial instinct was to give precise instructions about how people should act emotionally and propose specific heuristics for them to follow.
For example, when we decided to gather long, convincing angry conversations, we ran into this problem immediately. Anger is familiar, but almost nobody is comfortable being needlessly angry at a stranger, face to face, for an hour. The voice actors reached for the surface of anger instead. They raised their voices, delivered empty insults, and fell back on the phrase “I am angry with you!” Annotators found the conversations unconvincing and difficult to tag without context. Explaining our theory of what anger looks like only made things worse: the voice actors were now acting out the theory rather than the emotion.
If you tell people exactly what to do, the recordings just reinforce your existing heuristics, and models trained on them absorb those same, often biased, assumptions. We found that instead of prescribing a list of keywords or phrases, creating authentic situations worked much better. For example, when we needed argument-heavy interruptions, we paired people who genuinely disagreed on a specific topic and gave them a real issue to debate.
The conditions were set, and the behaviour followed.
The isolation problem
Even if we never prescribe anything, a second failure remains: the behaviours we want to capture one at a time do not exist one at a time.
Consider what actually makes a conversation feel frustrated, happy, or collaborative: it’s never just one thing. Tone, timing, word choice, shared history, and even the listener’s experience all play a role. We intuitively recognize these qualities when they're woven together in context, but if you extract one on its own, it loses its meaning and quickly feels unnatural. None of these facets are independent variables; they make sense together, not in isolation. Trying to separate one always leads to unexpected effects on the others.
Interruptions best illustrate the challenge, and they’re also the most important behavior for a turn-taking benchmark. For an interruption to feel real, it needs to be motivated — maybe there’s built-up context, frustration, or something so important that it’s worth cutting in. If you remove that context, every attempted interruption feels forced and unnatural.
In our early sessions, when we asked voice actors to interrupt, most gradually returned to polite conversation. But when we explicitly told them to talk over each other, the result was a loud, chaotic jumble. Neither approach produced usable data.
We tried it with comedy as well. We hired professional comedians and handed them a set of heuristics: they were told not to make the other person the butt of the joke, not to use real-world references, to make the other speaker laugh regardless of content, and to keep it clean. The rules were meant to isolate what made something funny so we could measure it, but the recorded sessions ended up being dry and any semblance of humor was forced. We had seen one of the comedians do stand-up before the session, so we knew the performer was not the problem.
Instead of asking people to produce a “backchannel-heavy conversation”, we worked backwards from the kinds of conversations in which backchannels already happen, and recorded those. If you want backchannels, casual talk is the conversation to collect. If you want interruptions, you need argument, which means finding people who actually disagree. Storytelling is where the long turns and pauses show up. The TurnBench corpus spans six conversation types chosen this way, and the mix of behaviours in each one then makes sense.
Naturalness under constraint
So you can’t script the behaviour, and you can’t collect it piece by piece. The temptation, at that point, is to stop controlling altogether: just put two people in a room and let them talk. That doesn’t work either. Ask people to chat in an unfamiliar studio and what comes out rarely sounds like how they talk at home.
Given enough time, people do eventually settle; anyone who has listened to a longform podcast knows the exact moment, a few hours in, when the presence of microphones fade into the background of the conversation. While we might have that luxury with podcasts, it makes for a poor collection method because the required behaviours need to appear often enough, and with enough variety, to make the recording time useful.
- Wait it out
- You get natural speech, but only if you can wait hours and live with whatever conversation happens to occur.
- Ask for it
- You get the behaviours you need, but only if you accept a prescribed performance.
What actually works is hiring people trained to produce authentic behaviour under conditions that are openly artificial. That profession is acting. Good actors can hold natural behaviour intact inside an unnatural environment. It takes thoughtful character direction and, above all, voice actors who commit to becoming the people who would plausibly have the conversation.
That is how we created the conversations for TurnBench. We gave voice actors situations that naturally encouraged the kinds of interactions we wanted to capture. Mid-turn pauses, backchannels, and barge-ins emerged as part of the conversation, giving the benchmark examples grounded in the flow of a real interaction.
There is an irony to all of this. Put anyone in a recording studio and ask them to have a conversation, and they will behave a little differently than they would in everyday life. Perfect naturalness, in that sense, is impossible. What matters is creating the conditions for people to behave naturally despite the artificial setting. Actors are trained to do exactly that.
Trust takes time
Creating those conditions depends on more than the performers. It also depends on how the technical and creative teams work together.
Technical teams are used to making requirements precise and outcomes repeatable. Performance works differently. A voice actor may interpret the same direction differently from one take to the next, and that variation can look like noise when you are trying to build a controlled dataset. But much of what makes a performance feel natural comes from exactly that freedom to interpret and respond.
Getting the balance right took time. Researchers needed to communicate what the data had to capture without prescribing the performance itself. Directors and voice actors needed enough context to understand the goal, then enough freedom to decide how to get there. As the teams worked together, they developed a shared language for translating technical requirements into direction that performers could actually use.
Having that shared understanding was crucial because naturalness is difficult to specify in advance. You often know it when you hear it, and learning how to produce it consistently requires iteration between the people defining the data and the people performing it. Trust, in that sense, becomes part of the collection method.
The ingredients of naturalness
Everything above comes down to one idea: naturalness is a property of a whole system of people, and you collect it by building that system.
What the system needs
- Find the right voice actors.
- Match people whose perspectives and working styles create the right conditions.
- Give them direction they understand and believe in.
That takes longer than writing a list of instructions. However, taking the time to build that system organically has been the most reliable way we've found to produce high-quality data.
TurnBench reflected the same challenge in its scores. Models struggled most on behaviours whose meaning depends on the surrounding conversation. A short response might be a backchannel or the start of a new turn. Overlapping speech might be an interruption or simply part of the exchange. Capturing those distinctions requires preserving the context that gives them meaning.
We have not solved naturalness. TurnBench conversations are recorded in a studio, on a schedule, with voice actors working from direction. That is not the same as speech at home. Annotators still flag sessions that sound stiff. Models still miss behaviours that only make sense in context.
But we are closer than we were in our early sessions. Each round of work teaches us how to get closer still.
Build models that understand human nuance
Natural human behaviour is the raw material of perceptual intelligence. Collecting it well requires thoughtful planning and patience to achieve data that is in distribution.
If you’re passionate about building models that understand the nuance of human behaviour, work with us. If you’re excited about performing, directing, or creating the conditions for natural behaviour, we’re hiring.