Data collection

We go and record the thing your model has never heard.

Real people, their own devices, the rooms they actually live and work in. Speech is what we have shipped — thousands of hours, nine locales, consent on every file. Video, image and text run on the same recruiting and QA machinery.

Start a request
Delivered so far

Every number here is speech work already shipped. New modalities are listed the same way — after they ship, not before.

4,000+
Hours of audio delivered
650
Transcribers and aligners on the bench
50 hrs
Delivered every working day
95%+
QA-sample accuracy, checked against WER
Modalities

What we collect, and where each one stands.

One line says delivered because it is. The rest say open because the bench is ready for them and the first orders are being scoped — we will list their hours the same way once they ship.

Speech & audio

Delivered

Two-party calls, spontaneous speech, read prompts, accents to quota, and the same script in a car, a kitchen and a warehouse.

  • Two-party conversations, each side on its own channel
  • Spontaneous and prompted speech with fillers kept
  • Accent and locale coverage recruited to quota
  • Matched acoustic conditions for degradation studies
4,000+ hours delivered · 9 locales live · 50 hours a dayRequest speech data

Video

Open

Task video from a phone or an action camera — third-person and first-person — plus the embodied capture on the Physical Intelligence pages.

  • Everyday tasks filmed in real homes and workplaces
  • Egocentric capture from chest and head mounts
  • Scripted scenarios with the same person across conditions
  • Depth and gripper demonstrations, to a written protocol
No volumes listed until the first order ships.See Physical Intelligence

Image

Open

Photographs to a spec — products, scenes, documents, faces with consent — across the devices and lighting your users actually have.

  • Product and shelf photography in real shops
  • Scenes and interiors across locales
  • Documents, receipts and forms, redacted to spec
  • Same subject across devices and lighting
No volumes listed until the first order ships.

Text & conversation

Open

Written dialogues, prompts and chat transcripts authored by people who work in the sector, in the register the sector actually uses.

  • Multi-turn dialogues to a scenario brief
  • Prompts and instructions in nine locales
  • Sector-specific written material by practitioners
  • Transcripts paired with the audio they came from
No volumes listed until the first order ships.
A contributor recording a speech session on her phone at home
Illustrative render of the capture setting — not customer data.
A worker recording speech on a phone while walking through a warehouse aisle
Illustrative render of the capture setting — not customer data.
How a collection runs

Four steps, and the first one is a phone call.

  1. 01
    Tell us the failure

    Not the dataset — the input your model gets wrong. A clip, a photo, a paragraph is enough to start.

    day 0
  2. 02
    Spec and quote

    Quantity, locales, devices, conditions, annotation depth — and the contributor rate you will be paying, stated as its own line.

    2–3 days
  3. 03
    Pilot batch

    Five percent of the volume, delivered first, so the spec breaks early instead of at the end.

    1 week
  4. 04
    Weekly delivery

    Drops into your bucket with a quality report and the reviewer chain attached to every batch.

    ongoing
What arrives

What lands in your bucket.

Defaults, negotiable in the written spec — but nobody should have to ask what a delivery contains.

Full delivery spec and terms
  • Original device capture, plus any agreed derivative
  • Manifest per item: locale, device, condition, consent reference
  • Transcript or labels to the agreed convention, where ordered
  • Quality report per batch — accepted, rejected, and why
  • Delivery into your region, on your bucket or ours
Coverage

9 locales live. The next one is a recruiting problem, not a research one.

Contributors record on their own phones, in their own rooms, in the accent they speak. Blue is where the bench works today; anything else we recruit to order against locale quotas.

de_DE Germanfr_FR Frenchit_IT Italianes_ES Spanishpt_BR Portuguesehi_IN Hindizh_CN Chineseko_KR Koreanja_JP Japanese
Record & earn

Your voice is worth something. We pay for it.

Pick a session, talk for twenty minutes on your phone, get paid once it passes review. No studio, no interview, no shift you have to show up for.

The base rate is $20 per approved audio hour. Conversation sessions, noisy-room sessions and locales we are short on pay above it — and whatever a session says it pays is what you are paid, shown before you start.

$20
Base rate per approved audio hour
20 min
Typical session length
7 days
From approval to payout
$10
Minimum balance to cash out
Questions

Before you send a brief.

Can you collect a modality you have not delivered yet?
Yes, and we say so on this page rather than pretending otherwise. Recruiting, consent, calibration and QA are the same machinery for a photo as for a call; what changes is the protocol and the kit, which is what the pilot batch is for.
Who are the contributors?
People recruited for the locale, the setting or the sector the spec calls for, paid per approved unit at a rate shown before they start. Anyone can join the speech pool; sector-specific work is recruited against the credential.
How is consent handled?
Captured on the file, before the contributor speaks or films, naming what the data trains. It can be withdrawn before payout, and the consent reference ships in the manifest of every item.
Where does the data live?
A collection is pinned to one region at kickoff — EU, US or India — and stays there. Graders stream, they do not copy.

Not sure which modality you need?

Send one input your model handles badly — a clip, a photo, a paragraph. We will tell you what would fix it, and what it would cost, before anyone signs anything.

Start a request