Request data

Tell us the input your model gets wrong.

Most teams do not need more data. They need the twelve minutes, the one shelf, the one form their model keeps failing on — the accent, the noise floor, the lighting, the phrasing. Say which one, and we go and collect it.

Capacity

The bench that would run your collection is already running. These are hours we have delivered in speech, not capacity we could hire.

4,000+
Hours of audio delivered
650
Transcribers and aligners on the bench
50 hrs
Delivered every working day
95%+
QA-sample accuracy, checked against WER
Live localesde_DEfr_FRit_ITes_ESpt_BRhi_INzh_CNko_KRja_JP— anything else we recruit to order.
What you can ask for

13 things, priced by the unit they are delivered in.

Grouped by modality. Speech carries delivered hours; the rest run on the same bench and are scoped per request — the badge says which is which.

Speech & audio

Delivered

Two-party conversations

2–4 weeks

Recorded calls between a caller and an agent, both sides on separate channels.

  • Scripted scenarios or free-form, your choice of sector and outcome mix.
  • Both channels recorded separately so a model can be trained on one side.
  • Consent captured on the recording and stored with the file.
You receiveStereo audio + verbatim transcript + turn boundariesPriced per hour of usable audio

Spontaneous speech

2–3 weeks

Unscripted speech from real speakers — the hesitations, restarts and overlaps scripts never produce.

  • Prompted topics, no read text. Fillers and false starts are transcribed, not cleaned.
  • Speaker metadata: age band, region, first language, device.
You receiveMono audio + verbatim transcript with disfluency markupPriced per hour of usable audio

Accent & language coverage

3–5 weeks

Targeted read and prompted speech to fill the accents your model drops.

  • You name the locales; we recruit to a quota and report against it weekly.
  • Code-switching sets available where the sector actually code-switches.
You receiveAudio + transcript + speaker profilePriced per speaker (30–60 min)

Real acoustic conditions

2–4 weeks

The same script recorded in a car, a warehouse, a call centre and a kitchen.

  • Matched pairs, so you can measure degradation instead of guessing it.
  • Far-field, hands-free, headset and handset variants of the same utterance.
You receiveParallel audio per condition + condition labelsPriced per condition-hour

Transcription & diarization

3–10 days

Your audio, transcribed and speaker-attributed to a written standard.

  • Two independent passes with a senior adjudicator on disagreement.
  • Verbatim or clean-read conventions, versioned in the guide we agree on.
You receiveTranscript, speaker turns, timestamps, agreement reportPriced per audio hour

Entity & intent labeling

1–2 weeks

The tokens that carry the transaction — account numbers, amounts, drug names, SKUs, intents.

  • Schema authored with your team, not pulled off a shelf.
  • Gold tasks injected at 3–5% throughout the batch.
You receiveSpan labels + intent tags + schemaPriced per audio hour

Preference & rating sets

1–2 weeks

Side-by-side judgments on model replies: which one a customer would accept.

  • Raters are sector operators, not general crowd workers.
  • MOS and MUSHRA runs follow ITU-R BS.1534 procedure.
You receivePairwise preferences, rubric scores, MOS/MUSHRA where relevantPriced per judgment

Adversarial & red-team audio

2–4 weeks

Callers who lie, push, spoof and hand the phone to someone else mid-call.

  • Severity tiers agreed before the run; nothing is published without review.
  • Includes replay, impersonation and authority-pressure attempts.
You receiveAttack audio + severity labels + expected refusalPriced per engagement

Video

Open

Egocentric task video

scoped per pilot

First-person video of people doing everyday tasks in real rooms, from a chest or head mount.

  • GoPro-class or phone rig with a timestamped IMU stream.
  • Task families and environments recruited to a quota, protocol agreed first.
You receiveRaw video + task and environment metadata + QC statusPriced per finished hour

Image

Open

Product & scene imagery

scoped per request

Photographs to a spec — shelves, interiors, objects — on the devices and in the light your users actually have.

  • Same subject across devices and lighting when you need matched pairs.
  • Locales recruited the way we recruit accents.
You receiveImages + capture metadata + labels where orderedPriced per image

Documents & forms

scoped per request

Receipts, forms and IDs photographed the way people actually photograph them, redacted to spec.

  • Synthetic or consented documents only — no scraped material.
  • Field schema authored with your team.
You receiveImages + field extraction where orderedPriced per image

Text & conversation

Open

Written dialogues & prompts

1–3 weeks

Multi-turn conversations and instructions authored by people who work in the sector, in its register.

  • Scenario briefs agreed first; authors pass the same calibration gate as annotators.
  • Available in every live locale.
You receiveDialogues in JSON + author metadataPriced per dialogue

Conversation transcripts

1–3 weeks

Transcripts paired with the audio they came from, or written to a scenario where no audio exists.

  • Verbatim or clean-read, versioned in the guide.
  • Entity and intent spans on request.
You receiveTranscript + turn boundaries + optional audio pairingPriced per conversation
How a request runs

Four steps, and the first one is a phone call.

  1. 01
    Tell us the failure

    Not the dataset — the call your model gets wrong. A clip is enough to start.

    day 0
  2. 02
    We scope and quote

    Quantity, locales, conditions, annotation schema, and the contributor rate you will be paying.

    2–3 days
  3. 03
    Pilot batch

    5% of the volume, delivered first, so the schema breaks early instead of at the end.

    1 week
  4. 04
    Delivery

    Weekly drops into your bucket, each with an agreement report and the reviewer chain attached.

    ongoing
What arrives

The delivery spec, published.

These are the defaults for speech, and the pattern for everything else. Every one of them is negotiable in the written spec we agree before the pilot — but nobody should have to ask what a delivery contains.

Audio

16-bit PCM WAV at 16 kHz by default, 48 kHz on request. Mono, or two-channel where both sides were recorded separately. Original device capture is kept alongside anything we resample.

Transcript

UTF-8 JSON per clip, with a plain-text and CSV mirror. Verbatim by default: fillers, repeats and false starts stay, and [pause], [overlap] and noise tags follow one written convention agreed in the spec.

Alignment

Start and end timestamps on every word, so accuracy and latency can be scored apart. Speaker turns and diarization labels ride in the same file.

Manifest

One row per clip: duration, locale, speaker profile, device, acoustic condition, consent reference and the reviewer chain that signed it off.

Quality report

Per batch: QA-sample accuracy, WER against the adjudicated reference, inter-annotator agreement on the double-passed portion, rework rate, and what was rejected and why.

Handover

Weekly drops to your S3 or GCS bucket, or ours with time-boxed credentials. Data stays in the region agreed at kickoff.

Before you ask

Ownership, consent, residency.

Who owns what we deliver?

You get a perpetual, non-exclusive licence for internal model development and evaluation. Exclusive and time-limited-exclusive terms exist — ask, and they are priced on the order form.

Read the terms
How was it consented?

Every contributor consents on the recording, before they speak, to exactly this use. Consent is stored with the file and referenced in the manifest, and it can be withdrawn — which is why the consent is taken up front rather than after.

Read the privacy policy
Where does it live?

A collection is pinned to one region at kickoff — EU, US or India — and the audio, transcripts and labels stay in it. Graders stream the clip rather than copy it.

Data residency
What if the pilot is wrong?

The pilot batch is 5% of the volume and it exists to break the schema early. If it comes back wrong, we fix the spec and re-run it before the main volume starts — that is what the pilot is for.

Not sure which of these you need?

Send one input your model handles badly. We will tell you what would fix it, and what it would cost — before anyone signs anything.

Start a request