Anatomy of a capture
What one Rayform capture holds, layer by layer, and why every layer is aligned to the same clock and the same frame.
Every capture in Rayform carries the same five layers. This post walks through them using one example: a few minutes in a kitchen, with a person cooking and a robot arm at the counter.
Scenes
The first layer is the space itself. Before anyone moves, the kitchen is reconstructed in three dimensions: a dense mesh of the room, the counters, the appliances, and everything on them. Behind the mesh sit the raw frames it was built from, depth and LiDAR, so you can go back to the sensor data when you need to.
The scene layer is what most 3D datasets stop at. For us it is the stage.
Motion
The second layer is everything that moves through the stage. The person's path from the fridge to the stove. The robot's camera as it looks down at the counter. The cup as it is lifted, carried, and set down. Each of these is a trajectory: a pose, meaning position and orientation, on every frame.
Because every trajectory lives in the same coordinate frame as the scene, you can ask simple questions and get precise answers. How far was the cup from the edge of the counter when it was put down? Did the person's path cross the robot's workspace? Where was the robot's camera when it saw the drawer open?
Interaction
The third layer is what the motion means. Grasps, placements, doors, drawers, tools. When the person opens the drawer, that is marked as an interaction with a start, an end, and the parts involved. Hands and moved objects are segmented, so a model can learn what a hand does to a handle, not just that both were present.
This is the layer that turns a recording into training data for behaviour.
Robot state
The fourth layer is the robot's own account of what it did. Joint positions and velocities, gripper opening, commanded versus measured. These are recorded in sync with every sensor frame, so a frame of depth video and the joint state at that instant are two views of the same moment.
If you are training a policy, this is the layer that closes the loop between what the robot saw and what the robot did.
Context
The fifth layer is everything around the capture. Where it was, when, what the lighting was like, how cluttered the counter was, what the surfaces were made of. Context is what lets you filter the corpus: every kitchen in low light, every capture on a reflective floor, every scene with more than one person.
Why alignment is the point
Each layer is useful on its own. Together, aligned to one clock and one frame, they are something else: a record of a place behaving, with the robot's actions and the world's responses on the same timeline.
That alignment is where most of our effort goes, because it is where most of the value is. A dataset where the layers almost line up teaches almost the right thing.
If you want to see what a capture looks like in your own pipeline, join the waitlist and tell us what you are training.