Language models learned from the internet's text. But the physical world doesn't come annotated — a robot's first problem isn't reasoning, it's knowing where things are, in metres, right now. That's spatial intelligence, and it's the substrate every embodied capability is built on.
Pictures of space vs space itself
RGB cameras give you pictures of space; depth sensors give you space itself. A depth frame is already structured data — distances, volumes, trajectories — with no inference step between the pixel and the metre. That's why a 40 × 30 dToF array can drive obstacle avoidance that a 4K camera can't, at a hundredth of the compute: the hard part is already done in the physics of the sensor, not in a network you have to train, power, and trust.
The data problem underneath
Text is abundant and self-annotating; the physical world is neither. A robot policy needs observation-action data — what the machine saw, what it did, and what happened — captured with the kind of time alignment that only comes from purpose-built data capture systems. This is the quiet reason so much of embodied AI is bottlenecked on data collection rather than model architecture, and why we build the capture rigs and the VisionLibra Data Services alongside the sensors.
The stack we're betting on
Our bet with VisionLibra is that the winning stack for Physical AI looks like the winning stack for GPUs: hardware, a developer-loved SDK, a model ecosystem, and a platform that operates fleets. Depth is where it starts, not where it ends — from a Spatial Mini in a smart lock to a Spatial Robot on an autonomous mobile robot, the same SpatialAI SDK runs the whole range. Get the spatial substrate right and everything above it gets simpler.
Start with the SpatialAI SDK — pip install spatialai and you're reading
depth in five lines. See the quickstart.
