Egocentric capture with a depth channel — stereo headsets, LiDAR-equipped phones or custom stereo rigs — so a model can learn where things are, not only what they look like.
Perit has shipped 4,000+ hours of speech through this bench. Physical-AI capture runs on the same recruiting, consent and QA machinery, and the first collections are being scoped with early-access partners now. Nothing on these pages is a delivered volume — the first hours will be listed here the way our speech hours are, after they ship.
The same tasks and rooms as egocentric capture, with a second stream: a per-frame depth map aligned to the RGB image, metric where the sensor supports it. Hand–object distance, container interiors, the gap between a plate and the shelf it is going onto — the geometry a monocular camera has to guess.
Rigs range from a consumer stereo headset to a LiDAR phone to a calibrated stereo pair; calibration files ship with every episode so the depth can be trusted downstream.
Tasks are recorded where they actually happen — not in a studio dressed to look like a kitchen. Hover a room to see the kind of setting the protocol calls for.


Labelled on the same bench as our speech work, to a written guide, behind the same calibration test.
The task family, the environments, the rig, the sensors, the episode count. One page, agreed before anything is bought or anyone is recruited.
A written capture protocol — framing, lighting, where a task starts and ends, what counts as a failed take — and the kit list. Operators are trained and tested on it before a single episode is paid for.
A small batch, ten to thirty episodes, delivered with QC status and a quality report. This is where the protocol breaks — on purpose, and cheaply.
You review the pilot, we fix the protocol, then the run scales on a weekly delivery cadence with the same report attached to every batch.
The task family, the environment, the rig and the episode count — one page, and we come back with a protocol and a quote.