All posts

Perception is the Next Frontier

By The Mundo Team

The Missing Half of Intelligence

Over the past few years, AI has made extraordinary progress.

Models can write software, solve mathematical problems, answer complex questions, and reason through increasingly difficult tasks. Every few months, another model pushes the frontier forward.

But intelligence is more than reasoning.

Before we can reason about the world, we have to perceive it. Humans do this almost instinctively.

We recognize when someone is confused before they ask for help. We know when a conversation has naturally ended and it’s our turn to speak. We can tell the difference between genuine excitement and polite enthusiasm. We notice when someone gestures toward an object instead of naming it. We understand that the same sentence can mean entirely different things depending on tone, facial expression, timing, environment, or social context.

Current AI systems are becoming remarkably capable at reasoning over structured information, yet much of the world remains unstructured. Conversations overlap. Background noise obscures speech. Cameras capture incomplete views. People interrupt each other, hesitate, change their minds, communicate implicitly, and express meaning through signals that never appear in a transcript.

Understanding those signals is a fundamentally different challenge.

We call this perceptual intelligence: the ability for AI to deeply understand the richness of real-world sensory experiences—from speech and video to gestures, environments, and human interaction—and use that understanding to interact naturally, and this will be a defining capability of the next generation of AI systems.

From internet-scale data to real-world experiences

Just as the first generation of foundation models learned from internet-scale data, the next generation will learn from increasingly rich forms of human experience. Audio, video, and entirely new modalities demand new kinds of data – not just more of it. They also require new ways to measure progress, because existing benchmarks often fail to capture the capabilities that matter.

Today, our datasets and evaluations are used by leading AI labs to build the next generation of multimodal systems. From natural speech-to-speech interactions and fine-grained video understanding to emerging modalities that don’t yet have established learning methods, our data helps researchers build, measure, and improve perceptual intelligence.

Our belief is simple: advances in AI won’t come from better models alone. They’ll come from a continuous feedback loop between better evaluations and better data. As models evolve, the datasets and benchmarks that shape them must evolve alongside them.

Join us in advancing perceptual intelligence

This investment gives us the opportunity to significantly expand our team and accelerate our growth. We’re hiring across research, engineering, and operations. If you’re excited about building the foundations that will teach the next generation of AI to perceive the world more naturally, come build with us.

To everyone who’s helped us get here—our customers, partners, and investors—thank you. We’re just getting started.