Rayform

The physical world, recorded for robots.

Private preview. The waitlist gets first access.

Robots learn from the world. Almost none of it has been recorded for them.

Every kind of model has learned from a record of the world. Robots are the first without one.

Language models
The whole internet.
Text on every topic, in every language, written by billions of people.
Image models
Every photograph uploaded.
Billions of pictures, captioned by the people who took them.
Robot models
A handful of lab datasets.
A few rooms, a few objects, usually one gripper, usually one camera.

Rayform records the physical world the way robots need it: in three dimensions, over time, across the places where robots will actually work.

What's in the dataset.

Every capture carries the same five layers, aligned frame by frame, whether it came from a kitchen, a factory floor, or a mine.

Scenes
Dense 3D reconstructions of real spaces: meshes, splats, and the raw depth and LiDAR frames behind them.
Motion
Camera, object, and body trajectories through those spaces, with a full pose on every frame.
Interaction
People and robots acting in the scene: grasps, placements, doors, drawers, tools. Hands and moved objects are segmented.
Robot state
Joint positions, velocities, and gripper state, recorded in sync with every sensor frame.
Context
Where, when, and under what conditions each capture happened: lighting, clutter, surface, weather.

How Rayform is built.

Three steps turn a capture in the field into a batch in your trainer, whatever the device and whatever you plan to train.

  1. Captured in the field.

    Rayform records real environments with calibrated multi-sensor rigs: depth, LiDAR, RGB, and the state of any robot in the loop. Every stream is time-stamped at the source, so frames from different sensors line up.

    • Homes, factories, mines, open terrain
    • Depth, LiDAR, and RGB together
    • Robot state recorded in the loop
  2. Aggregated into one corpus.

    Captures are registered into shared coordinate frames, deduplicated, and labeled with one vocabulary, so a kitchen in one city and a kitchen in another are directly comparable.

    • Shared coordinate frames
    • Synchronized timelines
    • One label vocabulary across sources
  3. Retrieved on your terms.

    Ask for exactly what you need to train on. Every door opened. Every grasp of a cylindrical object. Every scene in low light. Stream it straight into your pipeline.

    • Query by scene, object, action, or condition
    • Stream batches to your trainer
    • Delivered in open formats

Built for every team training robots.

One corpus, many robots. The same captures serve a humanoid learning to grasp and a drone learning to survey a mine.

Foundation models

Pretrain world models and policies on real geometry and real motion, not renders.

Humanoids and manipulation

Human and robot interactions with everyday objects, from the first reach to the final placement.

Navigation and autonomy

Spaces revisited over time, with furniture moved, doors closed, and the light changed.

Factories and mines

Industrial environments as they actually are: dust, glare, moving machinery, narrow passages.

Drones

Flight through real structures, indoors and out, with the geometry to check every trajectory against.

Research labs

A consistent, documented corpus to build on, with the raw frames available when you need them.

What makes it different.

Six commitments that hold for every capture in the dataset.

ScannedRendered
  • Real, not rendered.

    Every frame comes from a sensor in a real place. Simulation is a tool. It is not the ground truth.

  • Time, not snapshots.

    4D means the world as it changes: motion, interaction, and consequence, not a single frozen scan.

  • Variety by design.

    Captures are planned across places, conditions, and tasks, so a model learns the world and not one room.

  • Aligned across sensors.

    Depth, LiDAR, RGB, and robot state share one clock and one coordinate frame.

  • Licensed for training.

    Clear terms for commercial and research use, set with you during the preview.

  • Versioned releases.

    Each release is numbered and documented, so a result you publish can be reproduced.

Is the data real or synthetic?
Real. Every capture comes from sensors in a physical place. Nothing is rendered or simulated.
Can we ask for specific environments or tasks?
Yes. Waitlist members tell us what they are training, and that shapes what we capture next.
How is it licensed?
Per team, for training and evaluation. Terms are set with you during the private preview. Academic and commercial terms differ.
When does it launch?
Rayform is in private preview now. The waitlist gets first access.

Be first to train on the real world.

Join the waitlist and tell us what you’re building. We’ll show you what Rayform already holds for it.

Private preview. The waitlist gets first access.