FAQ

The questions we get before a first call.

From buyers, robotics teams and the people who work on the bench. If yours is not here, it is one message away.

Ask a question
1 of 5

Getting started

What does Perit actually sell?
Training data and the means to judge it. We collect speech, video, image and text to your specification, annotate on Foundry, license datasets we have already recorded, run speech benchmarks and RL environments, and capture embodied data — egocentric, depth, teleoperation, gripper, hand pose and tactile — for robotics teams.
How fast can we start?
A written spec and quote come back within 48 hours of a brief. A sample episode in your format follows within five working days, and a pilot batch with its QA report ships in two weeks. Volume runs on a weekly cadence after that.
What do we need to give you to get a quote?
The failure, not the dataset: a clip, a photo, a paragraph or a task description is enough. We come back with quantity, locales, devices, conditions, annotation depth and the contributor rate as its own line.
Where is pricing?
Every program is priced on its scope — locale scarcity, annotation depth, turnaround, exclusivity and, for embodied data, the rig and environment. Contact us with the brief and the quote comes back within 48 hours with every assumption listed.
Is there a minimum engagement?
No fixed minimum. The pilot batch — five percent of a collection, or ten to thirty episodes for embodied data — sets the smallest sensible engagement and is the cheapest way to find out whether the spec is right.
2 of 5

Data collection and annotation

Which modalities do you collect?
Speech and audio, video, image, and text and conversation, on one recruiting, consent and QA machinery. What changes per modality is the protocol and the kit, which is what the pilot batch is for.
Which locales and languages?
Nine locales today — Spanish, German, French, Italian, Brazilian Portuguese, Japanese, Mandarin, Hindi and Korean — with contributors recruited in-country. Anything else is a recruit-to-order and is scoped in the same quote.
Who are the contributors?
People recruited for the sector or the locale — operators, nurses, claims handlers, shopkeepers — not a general crowd. Every one of them clears a calibration test before paid work and is re-checked against the same set on a schedule.
How is annotation quality measured?
Delivered batches are sampled and scored against the written guide and, for audio, an adjudicated reference for WER. Both numbers, plus inter-annotator agreement and the rework rate, ship with every batch.
Can you annotate data we already have?
Yes. Most annotation work is on customer audio, images or video. It is pinned to your region and annotators stream it on Foundry rather than download it.
Do you license existing datasets?
Yes, on a perpetual, non-exclusive licence for internal model development and evaluation by default. Exclusive terms exist and are priced on the order form. Contributor consent covers that use and is referenced on every file.
3 of 5

Physical Intelligence

Where can a robotics team get real-world training data?
From Perit's Physical Intelligence lines: egocentric video, depth, teleoperation, handheld-gripper demonstrations, hand-pose ground truth, tactile streams, capture on your own rig, and video annotation — recorded in real homes, shops, workshops, farms and light-industrial sites across nine locales.
What hardware do you capture with?
Action cameras, LiDAR and Android phones, custom monocular and stereo rigs, RealSense-class depth cameras, Quest- and Pico-class headsets, UMI-style wrist-camera grippers, tactile gloves and ALOHA-class leader–follower arms — or your proprietary rig, operated by our crews under NDA.
Which environments can you record in?
Kitchens, living rooms and bedrooms, offices, retail and small shops, warehouses and light industrial, workshops and labs, agriculture and outdoor, and client-defined or confidential sites. Western homes are real homes in Spain, Germany, France and Italy, not sets.
Can we see a sample before we commit?
Yes. Sample episodes for egocentric, depth, gripper and teleoperation capture are on the samples page to watch and download, each with its file list. Hand-pose, tactile, BYOR and annotation samples are cut to your spec on request.
What formats do you deliver in?
Raw device capture plus MP4 for video, 16-bit PNG or NPY depth, CSV or HDF5 for IMU, pose and tactile streams, LeRobot- or RLDS-style layouts for teleoperation, JSON per episode for metadata and labels — or your schema.
Can you run a confidential program?
Yes. NDA'd crews with limited-visibility workflows, rigs kept in locked cases between sessions, capture on your site or in a space recreated to your drawings, and storage pinned to your region from the first upload.
4 of 5

Quality, consent and security

How is consent handled?
Captured before anyone speaks or films, naming what the data trains, withdrawable before payout, and referenced in the manifest of every item delivered. Bystanders are excluded or blurred to the protocol.
Where does the data live?
In the region agreed at kickoff — your bucket and your enclave if you want them. Annotators stream a clip rather than copy it, so review in one region does not move a file out of another.
Which certifications do you hold?
SOC 2 (report under NDA), ISO 27001:2022, ISO 27701:2019 and TPN Gold, with GDPR, CCPA and HIPAA controls documented on the compliance page. Certificates and the data processing agreement come with the control summary.
What ships with every batch?
The manifest, consent references and reviewer chain on every item, a quality report — sample accuracy, WER where it applies, agreement and rework rate — and, for embodied data, QC status per episode against the published checklist.
5 of 5

Working with us

How do I join the bench?
Every open role is on the careers page with its pay. Apply through the listing; within a few hours you get Foundry credentials for the training and testing modules, and paid work starts once you clear the calibration set.
How do I get paid to record?
Through the contributor workspace. The base rate is $20 per approved audio hour; conversation sessions, noisy-room sessions and locales we are short on pay above it, and whatever a session says it pays is what you are paid.
Are you independent?
Yes. Backed by Y Combinator and owned by nobody who trains a frontier model. Benchmark audio is held out and never delivered to any model provider.

Still a question?

Send it with whatever context you have. A written answer, or a spec and quote if it was really a brief, comes back within 48 hours.

Ask a question