Fueling the World's Top Frontier AI Labs
and Fortune 100 Enterprises with Frontier AI Data

Built By Engineers From

Backed By

  • Orange Collective

Benchmarks

Converse Benchmarks

We believe voice will be the natural interface between people and AI, but human speech is uniquely hard to get right. Unlike code or math, there is rarely one right answer: a good conversation depends on the setting, and every utterance carries emotion, tone, pace, accent, and language on top of its words.

Converse tests whether voice-native frontier models can understand, speak, and act on real-world speech the way people do. The first benchmark in the family, Converse-STT, measures how accurately they transcribe real, unscripted conversations.

Word error rate · lower is better

Show 10 more
Audio

Converse-STT

Conversational Speech-to-Text Benchmark

Evaluates frontier AI and speech-to-text models on Real World Conversational Data and on agent-generated, publicly available Pipecat audio. It measures word error rate (WER) on each dataset and shows how each model's score shifts between them.

Category

Speech-to-Text

Languages

English
In collaboration withCekura
View all

The Problem

AIcansolveolympiadproblems.Itstilllackshumannuance.

We'reinaTechnologicalRenaissance.Modelshavememorizedtheinternet.Theycanwriteessays,passbarexams,andprovetheorems.Butaskonetonegotiateadeal,comfortagrievingpatient,orspeakwiththewarmthandtimingofarealhumanvoice—andtheillusionbreaks.

Humanexpertiseisstaggeringlycomplex.Itspanseverymodalityandeveryculture—howasurgeonseestheoneshadowonascanthatchangeseverything,howatraderhearsriskinapausebetweenwords,howatherapistreadsafacebeforeasinglesentenceisspoken,howmeaningshiftsbetweenlanguages,accents,anddialectsthatnomodelwastrainedtounderstand.

Noneofthiswaseverinthetrainingdata.Scalingcomputewon'tconjureit.Syntheticdatawon'tapproximateit.Thebottleneckwasneverintelligence—it'stherichnessoflivedhumanexperience.

Our Solution

Weencodehumanexpertiseintomodelsthatworkfortherealworld.

We'reanappliedresearchlabbuildingthedatainfrastructureandhumanexpertisenetworktoencodereal-worldknowledgeintofrontiermodels—acrosseverymodality,language,anddomain.

Wepartnerwitheliteprofessionalstocapturewhattheyactuallydo.Thereasoningbehindadiagnosis.Theinstinctinanegotiation.Thecadenceofanativespeaker.Theengineeringinsightinadesigndecision.Themicro-expressionsamachinehasneverbeentaughttosee.

ThisflowsthroughourDataFoundry—apurpose-builtenginethattransformsrawexpertiseintostructuredtrainingdata,alignmentsignals,andrigorousevaluationsatscale.FromPhDmathematiciansandvoiceactorstoconstitutionallawyersandlinguists,everydiscipline,accent,anddialectgetsitsownpipeline.

Expert-Level Training Data
Datasets
Alignments
Evals
Benchmarks
Data Foundry
Tasks, Tools, RL Environments, & Rubrics
Elite Expert Network
Domain Experts
Linguists
Researchers
Global Workforce

Conversational AI

Models can speak. Teaching them how to sound human is the real work.

Closing the gaps our benchmarks expose takes data recorded the way people actually talk: full-duplex captures, emotional tagging, prosodic markers, scenario-anchored conversations, human preference data, and the evaluation loops that turn raw audio into training signal. This is the catalogue we ship to the labs building the next generation of conversational AI.

Full-Duplex Conversational Datasets

Two-speaker conversations captured at 48 kHz with isolated channels, overlap, backchannels, and barge-in preserved verbatim — the training audio behind real-time, conversational voice agents.

    Listening…

    Domain-Specific Speech Datasets

    Task-anchored sessions across medical intake, customer support, technical interviews, and emergency calls — tagged by scenario, role, and intent for vertical voice agents.

    The odor of spring makes young hearts jump.

    Scripted Voice Datasets

    Single-speaker performance reads from voice actors and trained narrators, phonetically balanced with controlled emotion ranges and multiple takes per line — production-grade material for TTS, voice cloning, and speech-to-speech.

    Transcription00:00:12

    Yeah, so I was thinking, <breathe/> maybe we could push the release until [hesitation] next Thursday? [laughter]

    That's not a bad idea, actually. [agreement] Let me check the calendar.

    Annotation & Evaluation Datasets

    Word-level transcripts, diarization, prosodic markers, scenario and role labels, continuous emotional tagging, and human preference scores — the training signal that turns raw audio into controllable, evaluable speech.

    Available in 40+ languages

    • American English
    • English
    • Spanish
    • French
    • German
    • Italian
    • Portuguese
    • Dutch
    • Polish
    • Russian
    • Ukrainian
    • Czech
    • Slovak
    • Hungarian
    • Romanian
    • Bulgarian
    • Serbian
    • Croatian
    • Greek
    • Swedish
    • Norwegian
    • Danish
    • Finnish
    • Icelandic
    • Irish
    • Turkish
    • Hebrew
    • Arabic
    • Persian
    • Urdu
    • Hindi
    • Bengali
    • Mandarin
    • Japanese
    • Korean
    • Vietnamese
    • Thai
    • Indonesian
    • Malay
    • Filipino
    Ready to bring AI into the real world?