Egocentric video, depth, teleoperation and handheld-gripper demonstrations — recruited, consented and QA-reported on the same bench that has shipped four thousand hours of speech. The people on camera are doing work they actually do.
Perit has shipped 4,000+ hours of speech through this bench. Physical-AI capture runs on the same recruiting, consent and QA machinery, and the first collections are being scoped with early-access partners now. Nothing on these pages is a delivered volume — the first hours will be listed here the way our speech hours are, after they ship.
Human video for breadth, depth for geometry, teleoperation when the embodiment has to match, handheld grippers for grasps in rooms a robot has never seen. Hover a card to see the capture.
First-person RGB video of real people doing real tasks.
You receive: raw video per episode, original device capture keptRGB-D, stereo and LiDAR capture for 3D scene understanding.
You receive: rgb video and aligned depth per episodeHuman-guided robot demonstrations, one episode at a time.
You receive: synchronised camera streams per episodeIn-the-wild wrist-gripper demos for grasp generalisation.
You receive: fisheye video per episodeTasks are recorded where they actually happen — not in a studio dressed to look like a kitchen. Hover a room to see the kind of setting the protocol calls for.
Recruiting to a locale, consent on the file, a calibration gate before paid work, QA sampling and a report per batch — none of it cares whether the file is audio or a first-person video.
These are speech figures. Physical-AI hours will be listed separately once the first collections ship.
Ten to thirty episodes before anything scales — because the protocol always breaks somewhere, and it should break cheaply.
The task family, the environments, the rig, the sensors, the episode count. One page, agreed before anything is bought or anyone is recruited.
A written capture protocol — framing, lighting, where a task starts and ends, what counts as a failed take — and the kit list. Operators are trained and tested on it before a single episode is paid for.
A small batch, ten to thirty episodes, delivered with QC status and a quality report. This is where the protocol breaks — on purpose, and cheaply.
You review the pilot, we fix the protocol, then the run scales on a weekly delivery cadence with the same report attached to every batch.
One page is enough to scope a pilot. If the protocol is wrong, the pilot is where we find out — before the volume.