Rayform

Why robots need the world in four dimensions

A scan tells a robot what a room looks like. A recording tells it how the room behaves. The difference is the whole problem.

Rayform2 min read

Most 3D data about the real world is a snapshot. A scanner walks through a building, and the output is a mesh or a point cloud: millions of points, frozen at one instant. For architecture, surveying, and games, that is exactly what you want.

For a robot, it is the beginning of the problem, not the end.

What a snapshot leaves out

A robot does not need to know what a kitchen looks like. It needs to know what happens in a kitchen. Which cupboard doors swing out and which slide. How a chair moves when it is pushed. Where people walk, and how quickly they appear from around a corner. What a cup does when it is picked up by the rim rather than the handle.

None of that is in a snapshot. All of it is in a recording.

When we say 4D we mean exactly this: three spatial dimensions plus time, captured together. Not a sequence of unrelated scans, but a continuous record of a space as it changes, with every frame placed in the same coordinate system as the last.

Three things time gives you

Causality. In a recording, actions have consequences you can see. A hand closes on a handle, the door opens, the light in the room changes. A model trained on this learns that the world responds, and how. A model trained on snapshots learns only that doors exist.

Dynamics. Objects have mass, friction, and momentum. A ball rolls off a table differently from a box. A recording captures the trajectory, not just the start and end. For a policy that has to catch, place, or avoid things, the trajectory is the lesson.

Affordances. The same object affords different actions in different contexts. A drawer that is full behaves differently from a drawer that is empty. Time reveals this: you watch the drawer being used, and the way it is used tells you what it is for.

Why this is hard to collect

Recording the world in four dimensions means multiple sensors, each with its own clock and its own frame of reference, all running at once in a place that was not designed for it. Depth cameras disagree with LiDAR. Timestamps drift. A robot's joint encoders report in one coordinate system and its camera in another.

Making these agree, frame by frame, is most of the work. It is also what makes the data usable. A capture where the depth frames are a few milliseconds out of step with the robot's joint state is a capture that teaches the wrong thing.

What we are doing about it

Rayform records every environment with calibrated multi-sensor rigs, time-stamps every stream at the source, and registers everything into one coordinate frame before it enters the corpus. Then we do it again in another place, under different conditions, with different people and different robots, because a model that has seen one kitchen has not seen kitchens.

The result is a dataset where time is a first-class dimension. If that is what your model needs, join the waitlist.

Be first to train on the real world.

Private preview. The waitlist gets first access.

More from Rayform