01 /Products
Products built for teams processing massive amounts of video.
Scalable infrastructure and the expertise to go with it, so your team builds intelligence from video faster: from training the model to testing it in the world.
01 /Video primitives
The video infrastructure underneath everything.
Transcoding, frame tiling, real-time ingestion and high-throughput pipelines, built in house. Ask for one frame or one window and get exactly that, with no cutting and no re-encoding.
- In-house transcoding
Every format in, one addressable form out. Built for random access, not playback.
- Frame tiling
A whole window in one image, at the resolution the question needs.
- Frame-addressable streaming engine
Any window of any episode, addressed by frame, streamed to a model or a person in one request.
- Real-time ingestion
Cameras, streams and screens enter the same pipeline as an archive.
- High-throughput pipelines
100,000 hours turned into scene-level samples in two weeks, files left in place.
02 /Machine annotation
Dense annotation of every hour you have, at archive scale.
Connect your data sources and data providers. Run any model as an analyzer: segmentation, dense captions, events, the fields you define. 100K+ hours in weeks, with human QC through our partners.
- Any model as an analyzer
Ours, open-weight, proprietary, or the one you trained last week. One versioned index per analyzer; a better model re-reads the archive.
- Segmentation and dense annotation
Objects, hands, activities, on-screen text, speech, outcomes: per frame, per scene, per episode.
- Data sources and data providers
Your buckets, your fleet, and the collection partners who record for you, landing in one index.
- Human QC, with partners
Low-confidence and rare events go to reviewers with the clip. The verdict lands in the same index.
03 /Episode retrieval
Ask your training runs anything.
World-leading visual search over every episode you have. Moments and frames in 500 ms. Build agents with deep search on top.
- Leads the chart on benchmarks
Supports queries that no one else can. Hand your agent the best tool to probe interesting scenarios. Read the paper
- 500 ms to the moment
Ask in plain English. The window is decided by the question: a slip, a regrasp, a whole shift.
- From hits to training set
Filter, balance, export a reproducible manifest. Train the next model on it.
- Agents with deep search
Deep search, aggregation, detail and watch agents, from Claude Code, Cursor, MCP or the SDKs.
04 /Realtime ingestion
Connect 1000s of cameras. Analyze in real time.
Set up alerts and reactive systems. Real runs and cameras join the same pipeline as your archive.
- Search while it streams
Cameras, live streams and screen recordings, indexed while they run.
- Alerts with the clip attached
Events described in plain language. Every alert arrives with the moment.
- The robot's memory
Long-term memory is the index. Short-term memory is the evidence stream.
- Every agent run is evidence
Computer-use agents recorded all the time. Ask what the agent clicked before the error.
05 /Connected to your robot data
MCAP, LeRobot and RLDS in. Manifests and loaders out.
Your files stay the source of truth for what the robot did. VideoDB takes the video part, the part nobody could search.
- In
MCAP, LeRobot, RLDS, RTSP, and files in S3, GCS or Azure. Simulator and world-model output too.
- Out
Parquet manifests, a PyTorch loader adapter, MCAP shards, evidence streams, scheduled agent results.
- Keeps the links
Every result carries episode id, timestamp, camera, source type and model version.
- Told apart
Real, simulated and generated video stay distinguished, in the index and in every export.