Two-party conversations
2–4 weeksRecorded calls between a caller and an agent, both sides on separate channels.
- Scripted scenarios or free-form, your choice of sector and outcome mix.
- Both channels recorded separately so a model can be trained on one side.
- Consent captured on the recording and stored with the file.
You receiveStereo audio + verbatim transcript + turn boundariesPriced per hour of usable audio
Spontaneous speech
2–3 weeksUnscripted speech from real speakers — the hesitations, restarts and overlaps scripts never produce.
- Prompted topics, no read text. Fillers and false starts are transcribed, not cleaned.
- Speaker metadata: age band, region, first language, device.
You receiveMono audio + verbatim transcript with disfluency markupPriced per hour of usable audio
Accent & language coverage
3–5 weeksTargeted read and prompted speech to fill the accents your model drops.
- You name the locales; we recruit to a quota and report against it weekly.
- Code-switching sets available where the sector actually code-switches.
You receiveAudio + transcript + speaker profilePriced per speaker (30–60 min)
Real acoustic conditions
2–4 weeksThe same script recorded in a car, a warehouse, a call centre and a kitchen.
- Matched pairs, so you can measure degradation instead of guessing it.
- Far-field, hands-free, headset and handset variants of the same utterance.
You receiveParallel audio per condition + condition labelsPriced per condition-hour
Transcription & diarization
3–10 daysYour audio, transcribed and speaker-attributed to a written standard.
- Two independent passes with a senior adjudicator on disagreement.
- Verbatim or clean-read conventions, versioned in the guide we agree on.
You receiveTranscript, speaker turns, timestamps, agreement reportPriced per audio hour
Entity & intent labeling
1–2 weeksThe tokens that carry the transaction — account numbers, amounts, drug names, SKUs, intents.
- Schema authored with your team, not pulled off a shelf.
- Gold tasks injected at 3–5% throughout the batch.
You receiveSpan labels + intent tags + schemaPriced per audio hour
Preference & rating sets
1–2 weeksSide-by-side judgments on model replies: which one a customer would accept.
- Raters are sector operators, not general crowd workers.
- MOS and MUSHRA runs follow ITU-R BS.1534 procedure.
You receivePairwise preferences, rubric scores, MOS/MUSHRA where relevantPriced per judgment
Adversarial & red-team audio
2–4 weeksCallers who lie, push, spoof and hand the phone to someone else mid-call.
- Severity tiers agreed before the run; nothing is published without review.
- Includes replay, impersonation and authority-pressure attempts.
You receiveAttack audio + severity labels + expected refusalPriced per engagement