<!-- source: /index.md -->
# VideoDB — Infrastructure for AI that understands the visual world

> VideoDB turns files, live streams, and camera feeds into persistent, searchable context. AI applications and agents retrieve exact moments and act on them in real time.

**Status:** v2.4 · RTStream is generally available

---

## Hero

Infrastructure for AI that understands the visual world.

VideoDB turns files, live streams, and camera feeds into persistent, searchable context. AI applications and agents retrieve exact moments and act on them in real time.

CTAs: [Talk to us](/company#contact) · [Build with VideoDB](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page)

---

## The problem

Video was built for people to watch. AI needs it as context.

The playback stack delivers files to players. VideoDB gives AI applications one backend to see, understand, and act on video, from archives to live streams.

- **The playback stack** — Built to store, encode, and deliver files to people on a player. Storage (S3 / GCS / buckets / archives), encoding + delivery (ffmpeg / transcoding / HLS / CDN / players), human playback (DRM / thumbnails / analytics / watch pages), AI bolted on (ASR / OCR / VLMs / embeddings / vector DBs), and custom glue (metadata stores / queues / webhooks / brittle pipelines). Built for watching, extended for AI, hard to scale.
- **The AI stack** — One system ingests every video source and connects to any model. Continuous footage becomes persistent memory, and applications retrieve the exact moments they need and act on them instantly. Built for AI, persistent memory, deploy anywhere.

---

## How VideoDB works

From continuous video to context, memory, and action. One backend runs the full loop. Connect any source, bring any model, deploy in any cloud. [Read the Platform Spec](/platform)

1. **Ingest any source** — Files, meetings, live streams, cameras, and screens.
2. **Understand every moment** — Speech, scenes, people, objects, actions, and events, indexed with time.
3. **Remember everything** — Understanding accumulates into persistent, queryable visual memory.
4. **Retrieve exact moments** — Search in natural language and get playable clips and context back.
5. **Act and create** — Trigger events, webhooks, and agents, or generate clips, captions, overlays, and streams.

---

## How teams use VideoDB

How organisations are solving real problems with VideoDB.

1. **[LIVE CAMERA] Real-time intelligence across thousands of cameras** — VideoDB ingests live RTSP feeds, runs custom CV models on every frame, and returns playable evidence the moment something matters. [See how teams use it](/live-camera-intelligence)
2. **[PROGRAMMABLE MEDIA] Create videos from search results** — Search your video library for scenes in natural language, then combine the best moments into new clips with captions, voice, and overlays, all through code. [See how teams use it](/programmable-media)
3. **[AGENTIC PERCEPTION] Realtime vision for AI agents** — Agents are creating content, running marketing, recording meetings, and using browsers for you. VideoDB gives them realtime visual context, so they can see, remember, and act on what happened. [See how teams use it](/agentic-perception)
4. **[TRAINING DATA] Turn raw footage into training samples** — Search across massive video archives to find the exact scenes, actions, objects, and edge cases needed for model training. [See how teams use it](/training-video-data)

---

## Why teams pick VideoDB

Built differently for the AI video era. VideoDB brings together encoding, indexing, retrieval, editing, and streaming in one purpose-built backend, powered by a media processing engine built over years for AI-native video applications.

- **10x** — Lower total cost. One bill replaces 10+ vendors.
- **100x** — Faster video retrieval. ~120 ms across PB archives.
- **5 min** — To production. pip install, POC same day.
- **1 API** — For realtime + files. Same SDK, same mental model.

---

## Production — already live

VideoDB powers real video workloads across healthcare, safety, media, AI agents, and enterprise workflows.

- **10 TB** uploaded every month — across cloud, VPC, and edge.
- **40K+** custom indexes built each month — bring your own model.
- **25K+** searches per month — every result a playable clip.

### Case patterns

- **OTT engagement:** "Custom search and editing on 2,500 hours of premium content. Pilots that used to take six months ran in six weeks." — streaming customer, 2,500 hrs catalog, programmable editing.
- **Healthcare deployment:** "1,000+ cameras, custom CV models, sub-second alerts in production. VideoDB collapsed nine months of integration into eight weeks." — healthcare customer, 1,000+ feeds, custom CV models.

---

## Loved by builders

Video infrastructure veterans, AI builders, and media technologists see VideoDB as a new backend for searchable, programmable, and agent-ready video.

- **Oliver Cameron** — Co-Founder & CEO, Odyssey ML
- **Tom Mason** — previously CTO at stability.ai
- **Johnny Boufarhat** — Founder, Hopin
- **Yohei** — General Partner, Untapped Capital
- **tooz** (@adarshsolanki), **Nour** (@Defi_Nour), **Strakyo** (@Strakyo)

---

## Ecosystem

Built for modularity. Works with every modern stack. Agent runtimes, model labs, camera systems, robotics pipelines, sovereign clouds — VideoDB is the common video data layer underneath. [Become a Partner](/company#partners)

- **Agent platforms** · **AI models integration** · **Cloud partners** · **Integration partners**
- **Camera systems** — fleet management, NVRs, security platforms.
- **Robotics stacks** — real-world perception data for control and policy training.
- **Media products** — catalog search and programmable editing for OTT and creator tools.
- **Sovereign clouds** — AI-native video layer for regional and government hosting.
- **World model pipelines** — structured, provenance-aware video for frontier model labs.

---

## Deployment

One SDK. Managed or in your cloud. Start managed, move to your own cloud later — no code changes. Edge GPU and on-prem extensions plug into either tier.

| Tier | Use | Notes |
|---|---|---|
| **[FULLY MANAGED]** Fastest path | Run on VideoDB Cloud | Multi-region (US, EU, IN), auto-scaling indexing & retrieval, usage-based, POC live in <48 hours. Best for media teams, agentic startups, ML product teams. |
| **[BRING YOUR OWN CLOUD]** Sovereign / regulated | Run inside your AWS, GCP, or Azure | Single-tenant control plane in your account, private link, no data egress, optional edge GPU for sub-second alerts. Best for healthcare, fintech, defense, telco, regulated EU buyers. |

---

## Security & controls

Enterprise-grade security and controls. SOC 2 Type II, GDPR, HIPAA, and ISO 27001, with customer-managed keys, Zero Data Retention, and SSO across your IDP of choice.

- **Zero Data Retention** — Configurable ZDR for sensitive workloads. Queries and originals can be automatically purged on your schedule.
- **SOC 2 Type II** — Independent audit of access control, change management, and incident response, renewed annually.
- **Single Sign-On** — SAML and OIDC across your IDP. SCIM provisioning and team-level RBAC on annual plans.

---

## Closing

The next wave of AI will understand the visual world. Give your applications and agents the infrastructure to see, remember, and act. Start with a file, a stream, or a camera.

CTAs: [Talk to us](/company#contact) · [Read Documentation](https://docs.videodb.io/pages/getting-started/welcome)

---

## Sitemap

- [/](/) — Homepage
- [/platform](/platform) — Platform overview
- [/agentic-perception](/agentic-perception) — Agents that see, hear, remember
- [/live-camera-intelligence](/live-camera-intelligence) — Cameras, sensors, live ops
- [/programmable-media](/programmable-media) — Query, edit, stream by code
- [/training-video-data](/training-video-data) — Training data for physical AI
- [/developers](/developers) — Quickstart, SDKs, OSS agents, pricing
- [/company](/company) — Mission, partners, team, investors, contact

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /platform.md -->
# Platform — VideoDB

> One backend. Any source. Six primitives — ingest, index, remember, retrieve, edit, stream — across files, live streams, cameras, and screens.

---

## Hero

One backend. Any source.

Six primitives: ingest, index, remember, retrieve, edit, stream. Across files, live streams, cameras, and screens.

CTAs: [Try VideoDB](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page) · [Read the docs](https://docs.videodb.io/)

---

## Trust strip

Fully Managed · Cross-Cloud · Real-time + Batch · BYO Models · Sovereign-Ready · SOC 2 Type II

---

## Six primitives

Six building blocks. One SDK. Any video moment. Static files and live feeds enter one pipeline. Index layers compound into a queryable memory. Retrieval surfaces the right moments. A new stream comes out the other side, on demand.

1. **Ingest** — Any source, any format, normalized on the way in. A single ingestion path with a built-in transcoding engine: any codec, any container. Files and live feeds land in VideoDB’s internal format, ready to be indexed. [See ingest reference](/developers#ingest)
2. **Index** — Stack understanding layers over every frame: scenes, speech, objects, embeddings, brands, people, custom domain events. Indexes are additive and reusable; new questions do not require re-processing the video. [See index reference](/developers#index)
3. **Remember** — Scene-level memory that persists. A 4-hour archive, a 1,000-camera fleet, an agent session all reduce to the same memory model. Scoped per user, agent, workspace, or fleet. Re-indexable forever. [See remember reference](/developers#memory)
4. **Retrieve** — Sub-second retrieval across petabytes. Semantic and structured queries across every video layer. Every hit is a streamable clip URL. Not a timestamp, not a row. Verifiable for users, consumable by agents. [See retrieve reference](/developers#retrieve)
5. **Edit** — Programmable timeline, cuts as code: cut, compose, dub, subtitle, reframe, add branding. All by code, not drag-and-drop. An agent-ready skill for media editing. The output is a clean stream. [See edit reference](/developers#edit)
6. **Stream** — Output fits the consuming system: HLS for players, WebSockets for agents, Webhooks for operations, agent-ready URLs for tools. One feed or a thousand. Every output is a stream URL. [See stream reference](/developers#stream)

One pipeline · realtime + batch · same SDK. PB-scale memory · ms-scale retrieval.

---

## How VideoDB works

One abstraction for the full video lifecycle. Ingest once, index deeply, retrieve exact moments, and stream results without stitching together storage, transcoding, search, and delivery services.

- **You ship indexes.** The platform owns compute and storage — chunking, replication, encoding, GPU scheduling. Your team ships indexes, not infrastructure.
- **Information access is managed for AI.** Queries run through segregated indexes scoped by tenant, workload, or domain. Results route automatically: no federation, no glue, no leaks.
- **Built-in admin console.** Define who reads what, where keys live, and what data can leave the boundary. One surface, no per-endpoint access logic.
- **Modular.** Choose the components you want and plug them into your system. Take search, memory, clips, alerts, or actions piece by piece, or adopt the whole stack.

---

## Why teams choose VideoDB

Built differently for the AI video era. VideoDB brings together encoding, indexing, retrieval, editing, and streaming in one purpose-built backend, powered by a media processing engine built over years for AI-native video applications.

- **10x** — Lower total cost. One bill replaces 10+ vendors.
- **100x** — Faster video retrieval. ~120 ms across PB archives.
- **5 min** — To production. pip install, POC same day.
- **1 API** — For realtime + files. Same SDK, same mental model.

---

## MP4 stack vs VideoDB abstraction

| | MP4 stack | VideoDB abstraction |
|---|---|---|
| Storage | Files in S3 / GCS; you manage chunking, replication, lifecycle (manual) | Managed by the platform, content-addressed, dedup'd (0 ops) |
| Edits | Re-encode on every edit; ffmpeg pipelines to run | Chunk-level slice to a new HLS link, no re-encode |
| Indexes | Vector DB + Postgres + metadata store, 3+ sources of truth | Indexes live with the media, one source of truth |
| Query | Cross-system join, custom code on the hot path (~12 s p95) | Indexes run retrievals, composed/filtered/scoped at query time (~120 ms p95) |
| Output | HLS adapter bolted on as a separate downstream service | HLS is native; every output is a stream URL |

---

## Technical solutions

Reference technical solutions, built on the same six primitives. Each has a deeper page with code, architecture, and dataflow.

1. **Personalized streaming** — Pull the right moments from memory, compose them with branding and subtitles, and ship a fresh HLS feed per user. Built on retrieve + edit + stream. [See architecture and code](/platform/personalized-streaming)
2. **Realtime alert** — Attach a camera, define what counts as an event in natural language, and get sub-second alerts with the matching clip attached. Built on stream + index + retrieve. [See architecture and code](/platform/realtime-alert)

---

## What teams say

- "Composable indexes are the right abstraction for video AI. VideoDB is the first team I've seen build the data-infrastructure layer instead of yet another model wrapper." — Mira Sundaram, Director of Video AI, Major OTT
- "We replaced four pipelines with one platform and shipped a feature in a sprint that was on our roadmap for two quarters. The team owns the indexes, and that's the unlock." — Daniel Park, VP Engineering, AI Infra Startup
- "What sold us was the search engine on top: ranking, recall, sub-second latency. Most products give you embeddings and call it search. VideoDB actually retrieves." — Priya Anand, Head of Search, Media Group

---

## Deployment

One SDK. Managed or in your cloud. Start managed, move to your own cloud later — no code changes.

- **[FULLY MANAGED]** Fastest path — Run on VideoDB Cloud. We manage the compute, storage, and scale; you build with the SDK. Best for media teams, agentic startups, ML product teams. [Build with VideoDB](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page)
- **[BRING YOUR OWN CLOUD]** Sovereign / regulated — Run inside your AWS, GCP, or Azure. Keep videos in your storage, use the same SDK. Best for healthcare, fintech, defense, telco, regulated EU buyers. [Talk to us](/company#contact)

---

## People who back us

Investors and advisors who've built this before — search engineers from Apple, computer vision investors from Voxel AI, veteran AI infrastructure VCs.

- **Kranti Parisa** — Advisor, Search. Former Search Lead at Apple; has built search systems at planetary scale.
- **Vaibhav Viswanathan (Vai)** — Computer Vision Lead at Voxel AI and investor at VideoDB. Deep operating experience across real-time perception and large-scale video pipelines.
- **Robin Vasan** — Investor, AI Infrastructure. Three decades backing developer platforms and data systems, from open source to enterprise-scale infra.

---

## Closing

Build AI that understands the visual world. Start with the SDK, or talk to us about an enterprise architecture review.

CTAs: [Start building](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page) · [Talk to us](/company#contact)

---

© 2026 VideoDB, Inc. · videodb.io

---

<!-- source: /developers.md -->
# Developers — VideoDB

> One command. Your agent gets a video backend. Skill-first install. SDKs for Python and JavaScript. OSS reference agents. The fastest path from media to structured machine context.

**v2.4 SDK**

CTAs: [Read the docs](https://docs.videodb.io) · [Join Our Discord](https://discord.com/invite/py9P639jGz)

---

## At a glance

- **2.4 M** — Hours indexed
- **240 ms** — P95 query latency
- **99.97%** — Uptime, last 90d
- **10x** — Cheaper than MP4

---

## Quickstart — from zero to a playable clip in five minutes

One command bootstraps every primitive into your agent runtime — or use the SDK directly. Live quickstart, v2.4.0.

```bash
# python
$ pip install videodb

# node.js
$ npm install videodb

# skill
$ npx skills add video-db/skills
```

```python
from videodb import connect
vdb = connect(api_key="vdb_...")

video = vdb.upload(url="https://.../keynote.mp4")
video.index_spoken_words()
video.index_scenes()

hits = video.search("the moment they show the demo")
print(hits.shots[0].stream_url)
```

- **Skills, not SDKs.** One command bootstraps every primitive. Claude Code · OpenAI · Cursor · n8n · Zapier — any agent that speaks tools.
- **Python & JavaScript.** Same primitives in both. Server- and browser-side covered.
- **Same loop, every scale.** From one clip to a corpus. The SDK that handles a 60-second screen capture handles a multi-million-hour archive.

---

## Concepts — four total, that's the whole API surface

If you've built on a database, you've built on this.

1. **Media** — The source. Files, streams, cameras, captures, datasets.
2. **Indexes** — Reusable layers. Scenes, speech, objects, events, brands, domain logic.
3. **Memory** — What you keep. Scoped, queryable, retention-aware.
4. **Events** — What you act on. Discrete moments with context, evidence, and a clip.

---

## Built on VideoDB — fork, run, ship

Production-shaped reference apps across the four solution surfaces. Open source. Pick one, clone it, ship.

| Repo | Language | What it is |
|---|---|---|
| [pair-programmer](https://github.com/video-db/pair-programmer) | python | Screen-aware coding agent. Continuous capture, indexed memory, tool functions for the agent. |
| [call.md](https://github.com/video-db/call.md) | typescript | Meeting copilot, no bots. Index calls into scenes, decisions, action items with playable evidence. |
| [agentic-streams](https://github.com/video-db/agentic-streams) | python | Live agent perception. Stream-in / stream-out perception so an agent can intervene as things happen. |
| [bloom](https://github.com/video-db/bloom) | python | Generative video. Research → script → assemble → publish. End-to-end. |

---

## Pricing — generous free credits to explore VideoDB

Start with enough credits to upload videos, build indexes, search moments, generate clips, and stream results. Try the SDK and ship a prototype before thinking about pricing.

- **Play around for free.** Get started with the SDK, sample apps, notebooks, and demo workflows. No sales call needed. [Get an API key](https://docs.videodb.io)
- **Building something special?** Apply for more credits and tell us what you are working on. We support serious builders, open source projects, research, demos, and early product experiments. [Apply for more credits](/company#contact)
- **Ready for production?** Usage-based and annual plans are available when you are ready to scale. [View pricing](/pricing)

---

## Resources

The things developers actually need.

- **Docs** — [docs.videodb.io](https://docs.videodb.io) — reference, guides, recipes.
- **GitHub** — [github.com/video-db](https://github.com/video-db) — SDK source + OSS agents.
- **Quickstart** — From `npx skills add` to a working search in under five minutes. [docs.videodb.io/quickstart](https://docs.videodb.io/quickstart)
- **Community** — Discord for questions, showcases, roadmap. [discord.gg/videodb](https://discord.com/invite/py9P639jGz)
- **Status** — [status.videodb.io](https://status.videodb.io) — uptime, incidents, maintenance.
- **Showcase** — Live demos of every primitive. [labs.videodb.io/projects](https://labs.videodb.io/projects)

---

## Programs for builders

Two ways to plug in.

- **Run a VideoDB meetup in your city.** You bring the builders; we cover venue, swag, and budget. Workshop, hack night, or full meetup. [Apply to host](https://videodb.io/dev-advocate)
- **Ship in public with us.** Build at the frontier with infra support, direct access to the VideoDB team, launch support, and cash prizes. [Join the next cohort](https://forge.videodb.io/)

Events, past and upcoming — meetups, hackathons, workshops, community demos. [View past events](https://luma.com/videodb?period=past) · Get the monthly Dev Digest (SDK changelogs, reference agents, community demos, upcoming events).

---

## Closing

One install. One key. One quickstart. Free to start. Talk to us when you're running serious workloads.

CTAs: [Get an API key](https://docs.videodb.io) · [View on GitHub](https://github.com/video-db)

---

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /pricing.md -->
# Pricing — VideoDB

> Plans for solo developers to enterprise teams. Start free, scale on demand. Transparent unit rates for ingest, indexing, search, streaming, and generation.

Try for free. No credit card required. Use Pro without any limits.

CTAs: [Chat with our Pricing Assistant](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page) · [Talk to us](/company#contact)

---

## Developer Pricing

Transparent unit rates for ingest, indexing, search, streaming, and generation.

### Free — $0 / month

Best for testing APIs.

- $20 free credits to start
- Then pay as you go

CTA: [Sign up](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page)

### Pro — $20 / month

Best for indie devs and small teams.

- $20 monthly credit that rolls over
- No rate limits
- Unlimited credit top-ups
- Auto-recharge so you never run out of credit
- Priority support via Email and Slack Connect

CTA: [Subscribe](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page)

### Enterprise — Custom Pricing

Mission critical workloads.

- Custom plan tailored to your need
- Hybrid / on-premise deploy, custom models & fine-tuning
- 99.9% SLA, dedicated support team
- Expert consulting

CTA: [Contact Sales](/company#contact)

---

## What you can do with Pro credits

**Scenario 01 — Live Monitoring:** Connect 1 camera stream and run 2 continuous alert policies for the entire month.

**Scenario 02 — Archive Search:** Index 1,000 minutes of video archives and enable semantic search capabilities.

**Scenario 03 — RAG Application:** Build a Retrieval Augmented Generation bot for 50 videos with deep recall.

---

## Live rate card

### See — Ingest & Process (Signal Input)

| Item | Rate |
|------|------|
| Realtime | $0.084 / hour |
| File Uploads | $0.09 / GB |
| Transcoding SD (360p) | $0.0040 / min |
| Transcoding HD (720p/1080p) | $0.0090 / min |

### Understand — Indexing & Search

| Item | Rate |
|------|------|
| Transcription | $0.01 / min |
| Scene Processing | $0.003 / scene |
| Search Query | $1.50 / 1k queries |

Scene Processing covers segmentation and embedding creation. Token charges are billed separately as per the token rate.

### Remember — Storage

| Item | Rate |
|------|------|
| Media Storage | $0.03 / GB / mo |
| Index Storage | $0.0005 / min / mo |

### Act — Action & Generation

Video Editing (charged for edited segments only):

| Item | Rate |
|------|------|
| Inline Edit | $0.004 / min |
| Overlay Edit | $0.01 / min |
| Resize | $0.01 / min |

Generation:

| Item | Rate |
|------|------|
| Dubbing | $0.15 / min |
| Translation | $0.02 / min |
| Image Gen (Ultra) | $0.19 / img |
| Audio Gen | $0.12 / min |
| Video Gen | $0.5 / sec |

### Data Transfer — Delivery

| Item | Rate |
|------|------|
| Video Streaming | $0.07 / GB |
| Audio & Image Hosting | $0.06 / GB |
| Downloads (720p) | $0.03 / min |

### Cost of Prompts (Tokens) — Intelligence

LLM & VLM Models, per 1K token. Note: 1 frame ~ 1K token.

| Tier | Rate |
|------|------|
| Basic | $0.0016 |
| Advanced | $0.0065 |
| Ultra | $0.00875 |

---

## Keep costs predictable

Optimize performance without overspending by controlling usage, storage, and output limits.

- Use sampling policies to cap frames per second
- Split cheap monitoring indexes from deeper recall indexes
- Keep prompts short, output length capped
- Store indexes only when you need replay and recall

---

## Need a custom plan?

Tell us your streams, retention, and indexing depth. We will map it to a predictable plan.

CTAs: [Talk to us](/company#contact) · [Open Docs](https://docs.videodb.io)

---

<!-- source: /company.md -->
# Company — VideoDB

> Making video usable as data and memory for AI.

We are a team of engineers who have spent years working at the intersection of AI, video systems, and cloud infrastructure.

---

## About us

We came together because we saw the same problem: video is stuck in the past.

Early experiments with LLMs made one thing clear: agents are coming, but the infrastructure isn't. The tools we have today were built for humans watching screens, not for software that needs to perceive the world.

For AI to be truly useful in real-world environments, it needs more than text processing. It needs a real-time perception layer — the ability to see, understand, and act on visual data instantly.

So we're building something new: video infrastructure that treats video as a real-time data stream, not a static file on a server. Searchable. Programmable. Intelligent. This is the foundation AI needs to work with vision and audio the way it works with text. This is VideoDB.

---

## Partner with VideoDB

Build the next video AI workflows with us. VideoDB partners with agentic web companies, solution providers, and infrastructure teams bringing AI into video-heavy workflows. If your customers need searchable video, real-time alerts, clips, memory, or agent actions, we should talk.

- **Agentic web companies** — Add video, screen, browser, and meeting context to your agents. Use VideoDB as the media memory layer behind your product.
- **Solution providers** — Bring AI video workflows to customers in media, healthcare, security, education, defense, and industrial operations.
- **Infrastructure providers** — Offer VideoDB inside your cloud, region, marketplace, or managed environment for customers who need data control and scale.

Want to partner? Tell us what you are building and where VideoDB can help. [Talk to us](/company#contact) · partners@videodb.io

---

## Team — people working on this

Engineers, researchers, and operators with deep backgrounds in media, ML, and infrastructure.

We are a remote tribe — headquartered in San Francisco, building from India. You'll find us where nature, good vibes, music, and creativity thrive, often trading city smog for the quiet peaks of Dharamshala. We've been together for the past five years, solving complex challenges in video infrastructure for AI. We believe that if AI can see what we see, hear what we hear, and listen to us in real time, it can amplify human expression.

**Founder and CEO** — Ashu has spent decades at the crossroads of machine learning and distributed computing. He believes beauty emerges when complex systems are made simple to use. VideoDB has been a journey to find a tribe that loves innovation, grows together, and builds a foundation for technology and research that creates useful things that truly matter.

---

## Backed by — our investors and advisors

The people and firms who believe video should be machine-readable.

### Investors

- **Robin Vasan** — Veteran investor in enterprise cloud infrastructure and AI.
- **Tara Tan** — Early-stage deep-tech investor focused on bold product innovation.
- **Yohei Nakajima** — Seed investor specializing in AI-driven, agentic technologies.
- **Eric Chen** — Seed-stage tech investor and partner at OVO Fund.

### Advisors

- **Kranti K. Parisa** — Former Search Lead at Apple.
- **Saurabh Gupta** — Co-Founder, ZEUX Innovation & Loopify; 3 US design patents and Red Dot & iF awards.
- **Neeraj Gupta** — Angel investor in agentic AI and GTM expert across industries.

---

## Contact — the right way to reach us

| Reason | Email |
|---|---|
| Sales & enterprise | sales@videodb.io |
| Partnerships | partners@videodb.io |
| Press | press@videodb.io |
| Security | security@videodb.io |
| Careers | careers@videodb.io |
| Founders | founders@videodb.io |

---

## Closing

Video was built for people to watch. We're making it usable for machines to understand. If you're building something serious on top of video, we should talk.

CTAs: [Talk to us](mailto:founders@videodb.io) · [Build with VideoDB](/developers)

---

© 2026 VideoDB, Inc. · videodb.io

---

<!-- source: /agentic-perception.md -->
# Agentic Perception — VideoDB

> Build AI agents with visual perception. One SDK gives your agent eyes, ears, and memory across screen, mic, video, and live sessions. Native runtimes for Mac, Windows, Linux, and the web.

**Status:** Agentic Perception

---

## Hero

AI is moving out of the chatbox. Agents are creating content, running marketing, recording meetings, taking calls, and using the computer. The world they operate in is live, continuous, and perceived through vision and voice — not turns of text. VideoDB gives your agents realtime real-world context and memory: one SDK across screen, mic, files, and live streams, so your agent can see what just happened, recall what it watched, and act on what it heard.

Surfaces: Screen · Mic · Camera · Files · Live streams

CTAs: [Try the SDK](https://docs.videodb.io/pages/getting-started/welcome) · [View OSS agents](https://github.com/video-db)

---

## What builders ship

The next generation of software won't live in a chat window. It will watch your screen, work the web for you, and run inside containers that never sleep. Builders on VideoDB are already shipping all three.

| Surface | What it is | Reference |
|---|---|---|
| Desktop | Agents that watch the screen with you. Pair programmers, meeting copilots, and second brains that share your screen, never your data. | [/agentic-perception/desktop](/agentic-perception/desktop) |
| Web | Agents that work the open web for you. Long-running pipelines that research, create, and publish: faceless YouTube channels, daily marketing, video research briefs. | [/agentic-perception/web](/agentic-perception/web) |
| Sandbox | Agent runtimes with unlimited memory. Every container, every browser-use and computer-use agent gets eyes, ears, and persistent recall. Hand it a repo; get back a demo. | [/agentic-perception/sandboxes](/agentic-perception/sandboxes) |

---

## Capabilities — one SDK, all of media

Files, live streams, screen captures — all enter the same system.

- **One command.** `npx skills add video-db/skills` bootstraps every primitive into your agent runtime.
- **Files, RTSP, screen, mic.** One API across every source.
- **Compose understanding.** Custom indexes the way you compose endpoints.
- **Search returns a playable clip.** Not metadata. Not timestamps. A clip the agent can play.
- **Stream in, stream out.** Sub-second alert, act, respond.
- **Claude Code · OpenAI · Cursor · n8n · Zapier.** Drop into any agent that speaks tools.

---

## Two modes for agents

Realtime by default. Memory when you ask. Stream in, context out. Nothing is stored unless you say so. Flip one flag when a moment is worth keeping.

- **Mode 1 · Ephemeral (Default).** Realtime, no storage. Frames flow in, structured events flow out. Nothing touches disk. Best for live copilots, alerting, and anything sub-second.
- **Mode 2 · Memory (Optional).** Remember and search. Flip one flag and the moment becomes a searchable clip. Memory and search are opt-in: on for the moments you care about, off everywhere else.

---

## Build a perception box

A dedicated perception runtime for teams that need realtime throughput, predictable cost and load, and zero outbound calls to a model API. Sized to your fleet. Every frame, every inference, every retrieval runs inside the box. Use the bundled models, or bring your own open-weight model.

Highlights: Realtime processing · Zero outbound · One flat number · Bring your own model

- **Realtime, sub-second pipeline.** Ingest, perception, and event-out sized to your throughput from day one.
- **Built-in network monitor.** Verify isolation in one glance. A live view of every connection the runtime makes.
- **Bundled perception models.** Vision, speech, and embedding models pre-loaded and ready to use.
- **One capacity envelope.** Flat monthly cost. No per-token surprises, no traffic-driven spikes.

CTA: [Talk to us](/company#contact)

---

## No-code workflows

Every VideoDB primitive is exposed as a node on n8n and Zapier — same primitives, drag-and-drop. Index a feed, search for a moment, clip and deliver, all without writing code.

- **n8n.** Drag-and-drop video memory: capture a stream, index it, retrieve clips, post to Slack or your CMS, all in a visual flow with the VideoDB nodes. [View n8n nodes](https://github.com/video-db/n8n-workflows)
- **Zapier.** Triggers and actions: trigger on a new clip in memory, cut a highlight, push to Drive, message the team. Wire VideoDB into the 6,000+ apps Zapier already supports. [View Zapier app](https://zapier.com/apps/videodb/integrations)

---

## Try it yourself

Every agent on this page ships as an open-source repo or a runnable notebook.

- **call.md** — Meetings captured as markdown with playable clips for every decision. [GitHub repo](https://github.com/video-db/call.md)
- **Pair programmer** — An agent that watches your screen and YouTube tabs and brainstorms with full context. [GitHub repo](https://github.com/video-db/pair-programmer)
- **Research agents** — A report you can watch. An agent crawls the web and assembles a video brief. [Live notebook](https://github.com/video-db/agentic-streams)
- **Try my repo** — Hand it a repo. A Pi agent runs it, narrates it, and ships back a demo video. [GitHub repo](https://www.trymyrepo.com/)
- **Build desktop agents** — Install the native SDK on Mac, Windows, or Linux. Start streaming screen + mic in minutes. [Quickstart notebook](/developers#quickstart)

---

## Closing

Give your agents eyes and ears.

```bash
npx skills add video-db/skills
```

CTAs: [Start building](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page) · [View on GitHub](https://github.com/video-db)

---

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /agentic-perception/desktop.md -->
# Agents that watch the screen with you

> Desktop agents — background runtimes that continuously watch pixels and listen to audio, building memory of what you've done and a model of what you need next. VideoDB ships the OS bridge, ingest pipeline, and recall memory layer for them.

---

## Hero

The desktop is the highest-leverage surface AI has ever had access to. The next decade of productivity software gets built on top of it — starting with an agent that can see what's on your screen.

For forty years, every productivity tool has been blind. Word processors don't know what you're writing about. IDEs don't watch you debug. Slack has no idea what you just said on the call. The most expensive surface on your computer — the screen itself — has been a complete blind spot.

That's changing. The first generation of desktop agents is here, and they're nothing like the chatbots that came before them. They don't sit in a sidebar waiting to be summoned. They run in the background, continuously, watching pixels and listening to mic input, building a memory of what you've done and a model of what you need next.

If LLMs were the moment text got cheap, desktop agents are the moment *context* got cheap. And context is what work actually runs on.

---

## Why now

Three things had to happen in parallel for desktop agents to be possible:

- Models had to get fast and multimodal enough that watching a screen at 24fps wasn't a fantasy.
- OS-level capture APIs had to mature on every major platform.
- Someone had to build the runtime: the part that handles streams, indexes frames, manages memory, and exposes the whole stack as a single tool any agent loop can call.

VideoDB is that runtime. One install gives an agent a native bridge into the operating system, an ingest pipeline that runs at the speed of the OS, and a memory layer that lets it recall anything it watched as a searchable, replayable clip.

> "The desktop agent doesn't ask what you want. It already saw."

---

## Three things builders are shipping today

### 1. The pair programmer that actually pairs

Copilot in your editor is a good autocomplete. A pair programmer is something else: it watches the whole environment. It sees the architecture diagram opened in another window. It sees the YouTube deep-dive queued up about the auth flow being rewritten. It hears the question a colleague asked in the last call. When you ask *"how should I rebuild this?"* it answers from the same shared context a human would.

VideoDB ships a reference implementation at `video-db/pair-programmer`. It uses the screen capture stream as the primary input and threads a recall API into the assistant so any moment from the last hour, day, or week is one query away.

### 2. The meeting copilot without the bot

The standard meeting assistant joins your call as a third participant. Everyone watches a robot blink in the corner. That model is dying. The new one is local: capture the audio and screen-share on the device the meeting is already happening on, run perception locally, never invite a stranger to your call.

`call.md` is the open-source reference. Every meeting becomes a markdown document. Every decision has a playable clip attached. When someone asks *"what did we agree about pricing?"* the agent answers with the moment, not a paraphrase.

### 3. The second brain that finally works

People have tried to build "second brains" for two decades. Every attempt failed for the same reason: capture is too hard. Nobody wants to take notes, tag emails, transcribe meetings, save the right Slack threads. Desktop agents solve this without asking anyone to change a behavior. The capture happens whether you notice it or not. Cognitive load drops to zero.

Memory becomes a recall API. The agent watching your screen turns into the most reliable note-taker you've ever had — one that remembers everything and forgets only what you tell it to.

---

## What you actually get

Install the native SDK on Mac, Windows, or Linux. One package, three lines of code:

```python
# Stream screen + mic continuously
vdb = VideoDB()
async with vdb.desktop(screen=True, mic=True):
    async for event in vdb.stream():
        # transcripts, screen events, intents - typed
        agent.handle(event)
```

What comes back is not raw video. It's a stream of typed events: transcripts, screen changes, application focus, recognized intents. The agent subscribes to exactly the signals it cares about and ignores the rest. Frames flow through the process without ever touching disk — unless you decide a moment is worth remembering.

### Privacy, on by default

Every desktop deployment is ephemeral by default. Frames are processed in-memory and discarded. Persistence is opt-in, per-stream, and can be locked to your own cloud. SOC 2 and HIPAA-ready out of the box.

---

## The category that didn't exist three years ago

Desktop agents will be the most consequential product category of the next five years. Not because they're a flashier chatbot, but because they finally close the loop between what software *sees* a user doing and what it can *help* them do. That loop has been open since the GUI was invented. VideoDB is the cheapest, fastest way to close it.

The builders shipping on this stack today are building the next Slack, the next Notion, the next Figma. Not because they have better models — everyone has the same models. They have something nobody else has: a backend that actually understands what's happening on the screen.

---

## Build a desktop agent

One SDK. Mac, Windows, Linux. Stream screen and mic in three lines.

CTAs: [Open the quickstart](/developers#desktop-quickstart) · [View pair-programmer on GitHub](https://github.com/video-db/pair-programmer)

---

<!-- source: /agentic-perception/sandboxes.md -->
# Sandboxes with unlimited memory and perception

> Browser-use, computer-use, and code sandboxes share a blind spot: when the container shuts down, the memory is gone. VideoDB plugs in as a perception sidecar that gives every box eyes, ears, and persistent recall that outlives the session.

---

## Hero

The most underrated trend of 2026 is the rise of the agent sandbox. Browser-use, computer-use, code-sandboxes. The agent's working memory is no longer your laptop — it's a fresh container, spun up on demand, that lives for the duration of one task and then disappears. This is a fundamentally better model than running agents on your machine. It's also fundamentally *amnesiac*.

Every sandbox today has the same problem: the agent shuts down at the end of the task and everything it learned shuts down with it. The next run starts from zero. The screen events it saw, the moments it noticed, the patterns it picked up — all gone.

Plug VideoDB into the runtime and that stops being true. The container gets **eyes, ears, and a memory layer that outlives it.** Run `agent_42` today, run it again next month, and the agent remembers what it saw the last time.

---

## What "perception in the sandbox" actually means

Most agent sandboxes today work on a simple primitive: they expose a screen, a keyboard, a mouse, and maybe a shell. The agent does its thing inside the box and reports back. That works for narrow tasks. It falls apart the moment the agent has to *understand* what just happened on the screen.

VideoDB sits as a sidecar to the sandbox. It taps the screen capture stream, the stdout, the file changes — whatever the runtime exposes — and runs the same perception layer it runs on a desktop or a live video feed. The agent gets a recall API instead of having to "remember" via context-window stuffing.

> "The container is ephemeral. The memory isn't."

---

## The Try-my-repo agent

The cleanest demo of this whole pattern is something we've been calling **Try-my-repo**. You hand it a GitHub URL. A fresh sandbox spins up. It clones the repo. It runs the install. It runs the tests. It reads the README. A Pi-class agent narrates the whole session as a voiceover. The output is a demo video, under a minute, with every meaningful moment indexed and clippable.

What you get back is not a transcript or a PR comment. It's a watchable artifact. You see the agent actually trying the thing. You hear it explain what went wrong. You can scrub to the moment a test failed and see the screen at that exact frame.

This is what a useful agent *looks like*. Not a wall of text. A clip you can play.

---

## Three classes of sandbox we're powering today

### 1. Browser-use agents

Browser-use is the most popular sandbox category in 2026. The agent gets a browser. It navigates, fills forms, clicks buttons. The state of the art is still DOM-driven, but DOM is a lossy projection of what's actually on the screen. **Perception fixes that.** The agent sees the page the way a human sees it, including the parts that don't appear in the DOM at all.

### 2. Computer-use agents

Computer-use is the broader bet: give the agent a full operating system in a container. Anthropic and others have published reference implementations. The blocker for long-horizon tasks has been memory. The agent's context window can't hold a multi-hour session. With VideoDB plugged into the runtime, memory is no longer bounded by the model. The agent can run for hours and recall any moment from the session as a clip.

### 3. Code sandboxes

The classic code-execution sandbox (E2B, Modal, Daytona) gets a perception layer on top. The agent runs code, watches output, takes screenshots, builds an indexed history of what it tried. Failed runs become searchable. Successful runs become reproducible artifacts.

### Sandbox partnerships

VideoDB ships native integrations with the leading sandbox providers. One config flag in your runtime gives the box a perception sidecar. No glue code, no infra to manage.

---

## What the integration looks like

```python
# Attach VideoDB to any sandbox session
async with sandbox.session() as box:
    vdb.attach(box)             # perception sidecar

    box.run("npm install")
    box.run("npm test")
    box.read("README.md")

    vdb.narrate(voice="pi")     # voiceover the whole session
    demo = vdb.publish()        # → demo.mp4 with searchable index
```

---

## Why this matters more than it sounds like it does

The agent economy of the next five years is going to run on sandboxes. Not on your laptop, not in a chat window. In disposable containers that spin up by the millions. Every one of them needs three things: a runtime, a memory layer, and a way to show humans what happened inside.

VideoDB is all three. The container gets eyes and ears. The memory persists past the lifetime of the box. The human gets a watchable artifact at the end. That's the missing piece in every "agents will do X" story today.

---

## Plug VideoDB into your sandbox

Browser-use, computer-use, or your own runtime. One integration; eyes, ears, and memory for every box you spin up.

CTAs: [View integrations](/developers#sandbox-integrations) · [Try-my-repo on GitHub](https://github.com/video-db/try-my-repo)

---

<!-- source: /agentic-perception/web.md -->
# Agents that work the open web for you

> Web agents are long-running media pipelines, not chat sessions. They wake on a schedule, crawl what's new, generate something coherent, and publish it where the audience already is. VideoDB is the agentic-stream runtime: crawl, index, narrate, assemble, stream — with checkpoint memory at every step.

---

## Hero

A prompt is a request. An agent is a loop. The next generation of web software won't wait for you to ask. It will go out into the world, gather what's new, make something out of it, and ship it back. On a cadence.

YouTube is a four-hundred-billion-dollar company because somebody figured out that video is the highest-bandwidth medium humans have. Agents are about to figure out the same thing, but from the other side. They're going to *make* video, not consume it. And the surface they'll do it on is the open web.

We call this category **web agents**. They are long-running pipelines, not chat sessions. They wake up on a schedule, crawl what's new, research what's interesting, generate something coherent, and publish it where the audience already is. The output is media: a clip, a podcast, a daily brief, a research deep-dive you can watch.

If desktop agents are about closing the loop on your screen, web agents are about closing the loop on the public internet.

---

## The agentic stream

The hardest thing about web agents isn't the model. It's the loop. A good web agent is not one prompt — it's a hundred-step orchestration that includes crawling, summarizing, scripting, narration, assembly, and publishing. Most stacks fall apart by step three because they were never designed to keep state across a multi-hour run.

VideoDB is built around the concept of an **agentic stream**: a long-running media pipeline that an agent can drive end-to-end, with checkpoint memory at every step. The agent crawls. VideoDB indexes what was crawled. The agent narrates. VideoDB assembles. The agent publishes. VideoDB streams the result out as HLS.

> "The web becomes a substrate the agent grazes on, day after day, and turns into something the audience can watch."

---

## Four shapes builders are shipping

### 1. Faceless YouTube pipelines

The cleanest example of a web agent is the faceless YouTube channel. The agent picks a topic (or you do). It researches it across twenty sources. It writes a script. It generates a voice. It pulls B-roll from public domain archives, your stock library, or a model. It assembles. It publishes. The whole pipeline can run on a daily schedule, producing a video a day with no human in the loop.

This sounds dystopian until you realize what it actually unlocks: any niche too small to support a human creator now has a content stream. Long-tail education, hyper-local news, internal company knowledge — all suddenly viable.

### 2. Marketing agents that don't sleep

The campaign loop — brief, cutdown, distribute, measure — runs on a one-week cadence in most companies. Web agents collapse it to a one-hour cadence. The agent watches the source clip, generates channel-native variants (16:9, 9:16, 1:1), drops the brand mark, lays subtitles, queues to each surface. Every channel gets a fresh asset every day, custom-built for the audience there.

### 3. Micro-learning agents

The future of education is not a 50-minute lecture. It's a 90-second explainer, indexed, searchable, with a clip-level recall API. Web agents that read long-form content (lectures, podcasts, papers) and distill them into micro-lessons are the most under-loved category in this space. Once a company has indexed its own training material this way, internal knowledge transfer stops being a meeting.

### 4. Research agents with video outputs

The "deep research" pattern is mostly text today. That's about to flip. A research agent that produces a video brief (with clips as citations, not URLs) is dramatically more useful for the things humans actually want to know. *"Show me what changed in the EV battery market this month"* wants a watchable answer, not a 3,000-word essay.

---

## What the runtime gives you

```python
# A daily research-to-video pipeline
async def daily_brief(topic):
    sources = await vdb.crawl(topic, max=20)
    script  = await agent.draft(sources)
    timeline = vdb.timeline(script.segments)
    timeline.narrate(voice="editorial")
    timeline.overlay(brand="daily.png")
    return timeline.stream(format="hls")

# Schedule it
schedule.cron("0 6 * * *", daily_brief, "ev-batteries")
```

### State, not statelessness

Every step in an agentic stream commits to a memory the agent can recall later. Day 30 of the loop can see what Day 1 produced. The agent gets better at the topic the longer it runs.

---

## Why this is the next platform shift

The default loop of the consumer web for two decades has been: humans make content, machines distribute it. Web agents flip both halves. Machines make the content, machines distribute it, humans curate and consume. The economics are unrecognizable. The cost to produce a piece of media drops by three orders of magnitude. The number of channels that can be served goes up by six.

The companies that win will be the ones that figured out the agentic-stream stack early. Not the ones with the best prompt. The competitive moat is going to be the loop, the memory, and the publishing pipeline. VideoDB is exactly that stack.

---

## Ship a web agent

Long-running pipelines with memory built in. Crawl, generate, publish. On a cadence, not a prompt.

CTAs: [Open the quickstart](/developers#agentic-streams) · [View agentic-streams on GitHub](https://github.com/video-db/agentic-streams)

---

<!-- source: /live-camera-intelligence.md -->
# Real-time Monitoring — VideoDB

> VideoDB powers the next generation of live camera intelligence for healthcare, security, defense, and operations. Ingest thousands of feeds, build understanding layers in realtime, and ship the alerts and dashboards your team runs on.

**Live Camera Intelligence**

CTAs: [Talk to us](/company#contact) · [Book a solution call](/company#contact)

Live demos: Intruder Detection · Dashcam

---

## What you get

The infrastructure layer your monitoring product runs on. A single backend for live cameras. From raw feed to the alert on your operator's screen.

- **Ingest 1000s of feeds in realtime.** RTSP, ONVIF, RTMP, edge cameras, drones. One ingestion path. Your fleet stays online. We handle the scale.
- **Build understanding layers on live feeds.** Run your own perception, recognition, and event logic on every feed in realtime. No model lock-in, no custom infra to maintain.
- **Setup realtime alerts and data streams.** Build your own realtime dashboards, mission controls, and apps on top. Alerts fire in milliseconds, with playable evidence.

---

## Built by operators

Our team has spent decades building video intelligence. Deploy on your cloud. Save cost with edge deployments. Talk to us for demos and a free solution call. We've shipped this stack with the teams running it in production today.

CTA: [Book a free solution call](/company#contact)

- "VideoDB collapsed nine months of integration work into eight weeks. We kept our models; they ran the infra." — Vai, Computer Vision Lead, Voxel AI
- "Always-on, agent-callable cameras for the home. A backbone we couldn't have built ourselves at our stage." — Jon, withVale · Home security

---

## Industries transformed with VideoDB

Already running where milliseconds matter. From bedside cameras to drone feeds to front doors. The same platform under each.

| Vertical | Page | Summary |
|---|---|---|
| Healthcare · ICU | [/live-camera-intelligence/icu](/live-camera-intelligence/icu) | Healthcare monitoring with ICUs. Bedside cameras and vitals fused into one clinical signal. Sub-second alerts at the nurses' station. |
| Defense · Field ops | [/live-camera-intelligence/drone](/live-camera-intelligence/drone) | Drone feeds, queryable in flight. Field operators stream RTSP straight in. Edge classifies on-site, cloud indexes the full mission archive. |
| Home · Family safety | [/live-camera-intelligence/home-security](/live-camera-intelligence/home-security) | Home security, transformed. Every camera becomes an always-on agent that can recognise, recall, and respond before the doorbell finishes ringing. |

---

## Partner with VideoDB

Add AI video to your cloud or solution. Become a channel or industry partner if you already provide solutions today. Bring an AI-native video stack to your customers, with the cost savings of a single backend replacing the legacy plumbing.

CTA: [Book intro call](/company#contact)

**[ INDUSTRY PARTNER ]** — Sovereign · regulated. Power your platform with VideoDB. For sovereign clouds, datacenters, and infrastructure partners already serving regulated buyers.

- Deploy VideoDB inside your region. Data never leaves your perimeter.
- White-label the platform or co-sell with your existing GTM team.
- Add healthcare, defense, security, and monitoring workloads to your catalogue.
- Joint engineering by operators, not by a vendor.

**[ CHANNEL PARTNER ]** — Solution providers. Sell more with an AI-native stack. For systems integrators, ISVs, and solution providers who already serve healthcare, security, defense, or industrial customers.

- Replace bolt-on video plumbing with a single managed backend.
- Use VideoDB inside your own product. Your brand, your customer.
- Shrink onboarding from quarters to weeks with our solution engineers.
- Margin-friendly commercials sized for the deals you already win.

---

## Four-week POC

From first feed to fleet-grade alerting. A pilot loop scoped to ship value in a quarter, not a year. Same SDK, picked per site: Cloud · VPC · Edge.

- **Week 1 — First feeds live.** 10 cameras streaming. Indexing schema agreed. Solutions engineer paired in Slack.
- **Week 2 — Understanding layer live.** Your model wrapped as an index. Webhook + WebSocket routes wired into your dashboards.
- **Week 3 — Full site rollout.** 100+ feeds. Your model embedded as an index. Recall tested on the incident library.
- **Week 4 — Annual contract.** Volume tiers, SLA, customer-managed keys, VPC or edge selection.

---

## Try it yourself

Explore at your own pace. Docs, demos, and examples to help you get started.

- **Public spaces** — Multicam tracking of an individual. Follow a single person across overlapping camera angles. Re-identification across feeds, built in. → [Open notebook](/developers#oss)
- **Home · security** — Home security use case. A doorbell-style stream that recognises people, packages, and pets. It pages you only when it matters. → [Open notebook](/developers#oss)
- **Healthcare · ICU** — Patient monitoring. Bedside camera and vitals fused into a single clinical event stream. Playable evidence on every alert. → [Open notebook](/developers#oss)
- **Industrial · Floor ops** — Operations & floor monitoring. Plant-floor cameras that read activity, count throughput, and flag unsafe conditions in realtime. → [Open notebook](/developers#oss)

---

## Closing

Give every camera you run understanding, memory, and the ability to act. Talk to us about your fleet, your stack, and what the next quarter could look like.

CTA: [Talk to us](/company#contact)

---

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /live-camera-intelligence/drone.md -->
# Drone — Case study

> A field-operations team combined ArcGIS geospatial maps with realtime drone video on VideoDB. Every frame now arrives on the map, indexed and queryable, in the same second it was captured — turning recordings into in-flight signal.

---

## Hero

Drones see the ground. VideoDB makes that view decision-ready.

The team combined geospatial maps in ArcGIS with realtime drone video on VideoDB. Every frame now arrives on the map, indexed and queryable, in the same second it was captured.

CTAs: [Talk to us about your fleet](/company#contact)

---

## 01 — The problem

**Drone footage was a recording, not a signal.**

The operations team flew a growing fleet of drones across hundreds of square kilometres. The footage was useful — eventually. Files landed in storage after each sortie. Analysts opened them the next morning, scrubbed timelines, and wrote reports.

By then, the moment was gone. A vehicle had moved. A field had been harvested. A perimeter had been crossed. Decisions that needed to happen **in flight** were happening **in hindsight**.

The map in ArcGIS already showed where every drone was, in realtime. The video those drones were capturing did not.

---

## 02 — The build

**One picture. Position and perception, together.**

The team kept ArcGIS as the source of truth for geography. They added VideoDB as the source of truth for what the cameras were seeing. The two layers now live on the same screen.

Every drone streams RTSP into VideoDB. Each frame is timestamped and tagged with the drone's GPS pose from ArcGIS. The combined feed becomes a single object: a moving camera with known location, looking at known terrain, captioned by a model the team chose.

### Position-aware video

Every frame is tagged with GPS pose. Click a point on the map, get the clip that captured it.

### Realtime understanding

Detection, tracking, and counts run live on the feed. The team brought their own models.

### Coordinated decisions

Fleet-wide events route to one queue. The next drone can be retasked before the first one lands.

---

## 03 — The outcome

**A single screen where ground truth lives.**

The map went from a tracker to a live operations console. An analyst can pin a moving target on the map, see what each drone in range can offer, and replay the clip that proves the call.

The team stopped writing summary reports the next morning. They started making the call the same minute.

| Metric | Value | Detail |
|---|---|---|
| Latency | ~2 sec | From frame captured to event on map. Edge classified, cloud indexed. |
| Console | 1 | Position, video, and event queue together — ArcGIS + VideoDB. |
| Fleet | Full | Coordinated retasking across drones. One event queue. |

> "The map used to tell us where our drones were. Now it tells us what they are seeing, while they are still seeing it." — Operations lead, field deployment

---

## Build the same loop

Your drones are already in the air. Make every frame count. We will pair an SE with your team and ship a spatial-video pilot in weeks.

CTAs: [Talk to us](/company#contact) · [See the platform](/platform)

---

<!-- source: /live-camera-intelligence/home-security.md -->
# Home Security — Case study

> A home-security partner built a camera product that tags each household member, surfaces only what matters, sends a daily summary, and gives every family its own long-running memory of home.

---

## Hero

Cameras that know who lives here.

A home-security partner built a camera product that tags the family it serves. It pings on what matters, sends a clean daily summary, and gives every household its own long-running memory of home.

CTAs: [Talk to us](/company#contact)

---

## 01 — The problem

**A camera with no memory is just an alarm.**

Every household with cameras knows the routine. A motion ping at the front door. Open the app. Squint at a thumbnail. Was that your kid? The dog? A delivery? A stranger? You scroll back six clips and lose two minutes.

The team wanted something better. Not louder alerts. A camera that **knows the household**, remembers the day, and only interrupts when it matters.

---

## 02 — The build

**Tag the family. Summarise the day. Remember the year.**

Cameras stream into VideoDB. The household tags each family member once. From then on, every clip is recognised, sorted, and added to a memory the household owns. Daily and weekly summaries write themselves. Critical moments stand out, with one tap to play.

### Household

- **Aria** — Parent
- **Mateo** — Parent
- **Noor** — Child

### Daily digest — Tuesday · April 14

| Time | Moment | Source |
|---|---|---|
| 07:42 | Noor left for school | Front door · recognised |
| 13:18 | Package delivered | Front door · logged |
| 19:04 | Unknown visitor at door | Verified by Aria · safe |

---

## 03 — The outcome

**A camera with a memory. A household that feels safer.**

The cameras stopped competing for attention. They started behaving like a member of the household who pays attention so nobody else has to.

Long-term context kept getting better. The system learned the kid's school schedule, the dog walker's gait, the rhythm of a Saturday morning. Over time it learned what was normal, so it could be useful about what was not.

### Family-aware alerts

Pings on strangers, packages, and unknown patterns. Silence on the household it already knows.

### Memory of home

Every recognised moment becomes part of a long-running household record. Searchable. Owned by the family.

### Safer by default

Long-term behaviour patterns let the AI flag the unusual, not the routine. Safer because it is quieter.

> "The cameras finally stopped crying wolf. When they ping now, we look up." — Household pilot, consumer launch

---

## Build it on VideoDB

Your camera product deserves a memory. We will help you ship family-aware cameras with realtime intelligence in weeks, not quarters.

CTAs: [Talk to us](/company#contact) · [See the platform](/platform)

---

<!-- source: /live-camera-intelligence/icu.md -->
# ICU — Case study

> A healthcare partner deployed VideoDB across their ICU rooms. Bedside cameras now alert on falls, hygiene, respiration, and sedation in realtime — every alert paired with a playable clip and a clean record of what happened.

---

## Hero

A coworker for the night shift, watching every bed.

A healthcare partner deployed VideoDB across their ICU rooms. Bedside cameras now alert on falls, hygiene, respiration, and sedation, in realtime. Every alert comes with a playable clip and a clean record of what happened.

CTAs: [Talk to us](/company#contact)

---

## 01 — The shift

**Two nurses. Twenty beds. A long night.**

An ICU at 2 a.m. has a steady rhythm and an unsteady patient. The vitals monitor catches the body. The cameras catch the room. Until now, no one was watching both at the same time.

If a patient slipped out of bed at 2:14 a.m., the team learned about it on rounds. If a sedation level shifted, it took a vitals review the next morning to notice the posture change. Hand-hygiene compliance lived on a clipboard.

The team did not need more screens. They needed **another set of eyes** that never blinked.

---

## 02 — The build

**Realtime alerts. Connected to the whole patient record.**

Every bedside camera streams into VideoDB. A clinical model layers on top: posture, motion, hygiene, respiration rate. Vitals from existing trackers flow into the same patient context. The result is a single timeline per bed, where camera evidence and vitals tell one story.

### Alert types

- **Fall · Bed exit** — Patient left bed unassisted. Pages the nurse with a playable clip from 5 seconds before the event.
- **Hand hygiene** — Compliance, not on a clipboard. Counts hand-hygiene events at the dispenser. Daily report writes itself.
- **Respiration** — Camera-derived breath rate. Cross-checked against pulse-ox. Outliers flagged before the alarm fires.
- **Sedation · Position** — Posture and stillness, watched. Track time at each posture. Trigger turn reminders. Confirm sedation depth visually.

### Bed 12 — night handover

| Time | Event |
|---|---|
| 22:10 | Asleep · left lateral |
| 23:48 | Turn complete · nurse assist |
| 02:14 | Bed exit attempt (alert) |
| 02:16 | Nurse at bedside |
| 03:42 | Respiration steady · 16/min |

---

## 03 — The outcome

**Every claim, backed by a clip.**

The morning handover used to be a story told from memory. It is now a timeline a clinician can scrub. Each event, from a bed exit to a hand-hygiene moment, has a playable clip beside the note.

Hand-hygiene audits, posture-care compliance, and incident reviews stopped pulling time away from the floor. The system tells the story. The clinicians spend their time on the patient.

| Metric | Value | Detail |
|---|---|---|
| Latency | Sub-second | Alert latency at the bedside. From frame to pager. |
| Record | 1 | Camera + vitals + tracker in one place. Per patient. Per bed. |
| Evidence | Every clip | Playable evidence attached to every event. Audit-ready by default. |

> "On a 12-bed unit, we did not add another nurse. We added another set of eyes that never looked away." — Clinical lead, ICU deployment

---

## Bring this to your ICU

The bedside camera you already have can do this. Today. We will pair a clinical SE with your team and ship a unit-wide pilot in weeks.

CTAs: [Talk to us](/company#contact) · [See the platform](/platform)

---

<!-- source: /programmable-media.md -->
# Programmable Media — VideoDB

> One backend for every video moment, from archive to live to delivery. Make every video moment queryable, editable, and streamable by code, from archive search to clip factories, dubs, and recaps powered by the VideoDB Director Agent.

**Status:** Programmable Media · In production

---

## Hero

One backend for every video moment, from archive to live to delivery.

Make every video moment queryable, editable, and streamable by code, from archive search to clip factories, dubs, and recaps powered by the VideoDB Director Agent.

CTAs: [Book a demo](/company#contact) · [Try the SDK](https://docs.videodb.io/pages/getting-started/welcome)

---

## Real-world case studies

**From dark archives to engagement engines.**

OTT platforms, broadcasters, and brand teams are turning their catalogs into engagement, subscriptions, and reach with the Director SDK.

- **[OTT creating vertical clips to post on social media.](/programmable-media/case-study-ott)** — OTT · Streaming. A workflow built on VideoDB: the editorial team analyses 2,500+ files, generates platform-native vertical cuts (Shorts, Reels, TikTok), and ships a season's worth of social in a week. Stats: 2,500+ files analysed · Editorial search & edit co-pilot · 9:16 vertical cuts at scale.
- **[Brand & marketing agencies producing content at scale.](/programmable-media/case-study-brand-marketing)** — Brand · Marketing. Media-first brands use VideoDB to organise their assets, reuse them across campaigns, and automate stitching of custom assets, generating new content at scale without rebuilding pipelines per launch. Stats: asset library organise & reuse · auto-stitch custom cuts per channel · campaign throughput at scale.

Trusted by: CloudPhysician · Docket · Wisdocity · FutureSmart AI · Palette · Clario · Paradigm Life · Ezoic · Art of Living.

---

## Capabilities

**Search, edit, dub, stream. By code.**

The Director SDK is the agentic skill SDK for media editing. Cut, compose, dub, subtitle, reframe, brand.

1. **Connect your library.** Plug in your storage (S3, GCS, drives, existing archives). Cloud storage · Direct upload · Open Media.
2. **AI tags every moment automatically.** AI watches your footage like a producer — logging people, scenes, actions, and emotions so you can find any moment in seconds. Characters · Locations · Camera Angles · Soundbites · Reactions & emotions · On-screen text · Graphics · Music · Logos & brands · Crowd & ambience · + more.
3. **Search, discover, and select.** Use natural language to find moments across years of content using detailed visual, spoken, emotional and contextual understanding. Trailer cutdowns · B-roll locator · Auto chapters · Keyword search · Deep Search Mode · Face Search · Find Similar · Plot & subplot · Chat-based Interface.
4. **Spin up clips, reels, and edits.** Turn media into clips, reels, dubbed cuts, and rough edits powered by AI — export to Premiere/Resolve, or straight to your publishing stack. Intro/outro bumpers · Logos · Watermarks · Transitions · Custom subtitles · Smart vertical (object-tracking) · Effects.

---

## Where teams use it

**Pilots that used to take three months run in three weeks.**

OTT, broadcasters, archives, and brand teams shipping with Director SDK + custom indexes.

- **OTT · Streaming — Catalog search + AI clip factories.** Premium catalogs indexed. Editors search by scene description. Director composes clip pipelines in Python.
- **Broadcasters — Live event enrichment.** Match-day recaps, social cuts, and multi-language dubs ready before the production team breaks for dinner.
- **Archives · Heritage — Decades of footage, searchable.** Faith, education, and nonprofit archives bring legacy media online, with provenance.
- **AI editing · Creator tools — Editor copilots inside Premiere and Resolve.** Catalog search and AI clipping land where editors already work. No NLE replacement project.

---

## Want to start right away?

**Bring AI to your media library, your way. Choose the path that fits your team.**

Explore with Director, give teams a chat interface over your archives, or partner with our experts to build the full workflow.

- **Director: Open Source.** The open-source agentic framework adopted by hundreds of media houses to build their own AI workflows on top of their catalog. [Explore on GitHub](https://github.com/video-db/Director)
- **Self-Hosted Chat.** For marketing and branding teams. Drop in our self-hosted chat interface and start working on your archives in plain language. No setup, no pipeline build. [Get access](/company#contact)
- **Work With Our Experts.** Want our team to realize your AI vision for your content? We scope, ship, and stand up the production workflow with you. [Book a demo](/company#contact)

---

## Deployment

**One SDK. Managed or in your cloud.**

Start managed, move to your own cloud later. No code changes.

- **[FULLY MANAGED] Run on VideoDB Cloud.** Fastest path. We manage the compute, storage, and scale. You build with the SDK. Best for Media Teams, Agentic Startups, ML Product Teams. CTA: [Build with VideoDB](https://console.videodb.io/auth?utm_source=videodb.io&utm_medium=cta&utm_campaign=landing_page)
- **[BRING YOUR OWN CLOUD] Run inside your AWS, GCP, or Azure.** Sovereign · regulated. Run it in your cloud. Keep videos in your storage. Use the same SDK. Best for Healthcare, Fintech, Defense, Telco, Regulated EU Buyers. CTA: [Talk to us](/company#contact)

---

## Try it yourself

**Real notebooks for media workflows. Click any card to open.**

Working examples we've published: clip factories, multi-language dubs, recaps, and catalog search. Run them in your browser before you ever pick up the phone.

- **[Vertical clip factory for OTT.](/developers#oss)** Auto-surface the most shareable moments from a long-form title and generate 9:16 vertical cuts ready for Shorts, Reels, and TikTok.
- **[Multi-language dubs at scene level.](/developers#oss)** Translate, dub, and re-sync scene-by-scene. Ship dozens of language tracks for a single title in hours, not weeks.
- **[Live event recap & highlights.](/developers#oss)** Index events as they happen. Director composes recap reels and social cuts before the broadcast ends.
- **[Semantic catalog search for editors.](/developers#oss)** Search a premium catalog by scene description, talent, or brand. Every hit a playable clip in milliseconds.

---

## Closing

Want our experts to realize your AI vision for your content?

CTAs: [Book a demo](/company#contact) · [Try the SDK](https://docs.videodb.io/pages/getting-started/welcome)

---

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /programmable-media/case-study-ott.md -->
# OTT Platform — Case study

> A leading OTT platform built a search engine over its own catalog. 2,500+ files indexed by faces, plot, scene, music, dialogue, and on-screen action. A Director SDK editing pipeline on top compiles, captions, reformats, and ships clips the moment a search returns a hit.

---

## Hero

A leading OTT platform built a search engine over its own catalog.

2,500+ files indexed by faces, plot, scene, music, dialogue, and on-screen action. A Director SDK editing pipeline on top compiles, captions, reformats, and ships clips the moment a search returns a hit.

- **2,500+** files indexed
- **11** custom indexes live
- **12x** editorial throughput

Status: In production

---

## The story

**The catalog wasn't the moat. The indexes on top of it were.**

The platform had a decade of premium content. An enviable library, on paper. In practice, it was just files. The editorial team knew which scenes made fans cry, which lines became memes, which faces audiences kept coming back for — but none of that lived in a system. It sat in heads and spreadsheets.

> "We didn't build a clip tool. We built a search engine that understands our show."

Their leadership saw clearly that the next year wouldn't be won by who had the most content. **It would be won by who indexed it best.** AI made every catalog operationally identical; only the layer of context on top differentiated. So the editorial team got to work on what they alone could do: designing indexes that captured how *they* understood their own content.

---

## The indexes they built

**Eleven indexes. Each one was the team's opinion, encoded.**

| # | Index | Sample query |
|---|---|---|
| 01 | Identity — character & actor faces | "every close-up of the lead, season 2" |
| 02 | Narrative — plot & subplot beats | "betrayal-arc scenes across the season" |
| 03 | Dialogue — spoken-word + byte search | "the line about second chances" |
| 04 | Location — locations & backgrounds | "all rooftop scenes after dark" |
| 05 | Music — music cues & needle drops | "every needle-drop in the finale" |
| 06 | Emotion — emotional beats | "the tearjerker moments, season 3" |
| 07 | Action — action, violence, gunshots | "chase scenes longer than 90 seconds" |
| 08 | Cast graph — all actors in a title | "scenes with leads A and B together" |
| 09 | Brand — on-screen brands & props | "all branded-prop appearances" |
| 10 | Scene — exact-scene retrieval | "the scene right before credits" |
| 11 | Virality — shareability signal | "likeliest-to-go-viral moments" |

A new index can be added in a week and retrofitted across the catalog automatically.

---

## The editing pipeline

**Search returns hits. The pipeline ships the reel.**

The intelligence isn't just in finding the moment. It's in everything that runs the second a search returns: compile, caption, smart-vertical, brand, publish. One Director SDK pipeline, configured per output.

1. **Compile the moment.** Search hits become a sequenced cut. Director picks the right in/out points, stitches scene-to-scene, balances pacing.
2. **Caption + subtitles.** Dynamic, word-by-word captions in the show's voice. Localised subtitle tracks for every language the platform serves.
3. **Smart vertical reframe.** Track the subject across the frame. 9:16 cuts that keep the right face in centre, never the back of someone's head.
4. **Brand & publish.** Brand kit, end-card, deep-link to free trial, platform-native export. Reels, Shorts, TikTok, all from one source pipeline.

Some teams have automated this end-to-end. Faceless reel channels run 24/7 on the same indexes, with no editor in the loop. The search engine picks the moments, the pipeline ships the cuts, the channel grows in its sleep.

---

## By the numbers

What an indexed catalog plus pipeline actually change:

- **11** custom indexes live. Every dimension the editorial team cares about, queryable in plain language.
- **2,500+** files retro-indexed. Full catalog brought online behind every new index, automatically.
- **12x** editorial throughput. A season's social + recap output ships in a week, on the same pipeline.
- **24/7** faceless channels run. Full automation for sub-brands. Search picks, pipeline publishes, no human in the loop.

---

## Quote

> "The indexes are the product. Every clip, every recap, every recommendation we ship sits on top of opinions our team had about our own content. VideoDB just made those opinions queryable, and the pipeline does the rest." — Editorial lead, leading OTT platform (name on request)

---

## Why leaders choose VideoDB

**AI is only as powerful as the indexes you give it.**

Generative AI flattened the output layer. What still differentiates is the context the model gets to reason over. For a media business, that context is the indexes you've built on your own content. The richer your indexes, the better every downstream AI workflow performs.

VideoDB is the framework that lets a team build that layer continuously: new index this month, retrofitted across the back catalog by Friday, powering production workflows by Monday. That's the AI advantage smart leaders pick — a compounding intelligence layer on top of content only they own.

---

## Closing

Give your editorial team an AI co-worker. Not another tool to learn.

CTAs: [Talk to us](/company#contact) · [Back to Media Solutions](/programmable-media)

---

<!-- source: /programmable-media/case-study-brand-marketing.md -->
# Brand & Marketing — Case study

> Brand and marketing teams turned a footage library into a queryable brand intelligence layer. Every shot indexed by product, talent, mood, lighting, location, brand cue, and music. A Director SDK pipeline compiles, captions, smart-reformats, and publishes channel-ready cuts the moment a search returns a hit.

---

## Hero

Brand & marketing teams turned a footage library into a queryable brand intelligence layer.

Every shot indexed by product, talent, mood, lighting, location, brand cue, and music. A Director SDK pipeline on top compiles, captions, smart-reformats, and publishes channel-ready cuts the moment a search returns a hit.

- **9** brand indexes live
- **8x** asset reuse per launch
- **~60%** shoot days avoided

Status: In production

---

## Trusted by media-first brands

Art of Living · Palette · Ezoic · FutureSmart · Wisdocity · Paradigm Life

---

## The story

**A library full of unused footage. An asset only when it answers questions.**

Every brand team has the same closet: petabytes of footage, every new campaign answered with "let's shoot more." Not laziness. **The library doesn't answer questions.** A marketing lead can't ask "every shot of the product on a kitchen counter, warm light, no people," so booking a new shoot is faster than finding the right one.

> "We didn't need a better file cabinet. We needed our library to know what the brand stands for."

The brand leaders we work with figured this out a year ago. The competitive edge in the AI era is the intelligence layer over a brand's own footage. So they designed nine indexes that encode how they see their content: product, mood, lighting, talent, brand cue, music, scene, location, dialogue. Then they let VideoDB run them across the whole library.

---

## The indexes they built

**Nine indexes. Each one the brand's opinion, encoded.**

| # | Index | Sample query |
|---|---|---|
| 01 | Product — product-in-frame | "hero product, neutral bg, 3-4s" |
| 02 | Mood — mood & emotional tone | "calm, aspirational, golden-hour" |
| 03 | Lighting — lighting & time of day | "matching golden-hour shots" |
| 04 | Talent — talent & people | "every shot of ambassador X" |
| 05 | Brand cue — on-brand & off-brand | "on-brand only, exclude rejects" |
| 06 | Music — music tempo & mood | "upbeat 90+ BPM, brand-licensed" |
| 07 | Scene — scene type & setting | "kitchen-counter B-roll, daytime" |
| 08 | Location — location & geography | "shots filmed in APAC region" |
| 09 | Dialogue — dialogue & voiceover | "testimonials mentioning 'every day'" |

A new index can be added in a week.

---

## The content engine

**Once AI knows your content, the work becomes templates.**

Structures the brand defines once, run by VideoDB across the library forever.

- **Faceless YouTube channels.** Sub-channels publishing daily. No on-camera talent, no editor in the loop.
- **Demo videos at SKU scale.** One template per product. Infinite variations by colour, market, and channel.
- **GenAI as a teammate.** Sora, Veo, Runway, ElevenLabs in the same pipeline as your indexed library.
- **Dubs and captions at scale.** Scene-level dubbing in dozens of languages. Brand-styled captions per market.

A content engine that scales without a production team. Templates run. The engine compounds. AI does the busywork the brand never wanted to own.

---

## By the numbers

What a queryable library plus pipeline change:

- **9** brand indexes live. Product, mood, lighting, talent, brand cue, music, scene, location, dialogue.
- **8x** asset reuse per launch. Same hero shot powers eight channel cuts instead of sitting unused.
- **~60%** fewer shoot days. Re-use replaces re-shoot for most campaign needs.
- **24/7** faceless channels run. Fully automated brand sub-channels publishing daily, no human in the loop.

---

## Quote

> "The library used to be a cost centre we kept paying for. Now it's the smartest member of the marketing team, because the indexes are ours, the pipeline runs itself, and both get sharper every quarter." — Marketing lead, media-first brand (name on request)

---

## Why leaders choose VideoDB

**AI is only as good as the indexes you feed it.**

Every brand can generate a cut now. What still differentiates is the context AI gets to reason over: the indexes you've built on your own footage. A richer on-brand index means cleaner AI cuts; a richer mood index means sharper performance creative. Each new index makes every existing AI workflow smarter on the brand's data, instantly.

VideoDB is what gives brand teams the framework to build that intelligence layer continuously. That's the AI edge brand leaders pick — a growing set of opinions, encoded as indexes, that AI gets to use every time the brand ships.

---

## Closing

Give your marketing team an AI co-worker. Not another platform to manage.

CTAs: [Talk to us](/company#contact) · [Back to Media Solutions](/programmable-media)

---

<!-- source: /training-video-data.md -->
# Training Video Data — VideoDB

> The training-data layer for world models, VLAs, and physical AI. We partner with some of the largest video data providers and frontier labs to take petabyte-scale raw footage and stand up a queryable, training-grade dataset. Custom labeling models, human-in-the-loop verification, provenance per clip.

**Status:** Training Video Data

---

## Hero stats

- **1M+ videos** — Indexed across partner programs
- **PB-scale** — Tuned for training pipelines
- **<200ms** — Search across the full corpus

CTAs: [Book a consult](/company#contact) · [See the lab](https://labs.videodb.io/)

---

## The bottleneck

World models are hungry for video. Most teams are still stuck building pipelines. Before training can begin, raw footage has to be cleaned, clipped, indexed, labeled, and delivered. That work slows teams down long before the model ever sees the data.

1. **Scale.** "Our pipeline cracks every time the corpus doubles." Multi-million-hour archives outrun ad-hoc scripts. Even the biggest labs end up shipping their own curators just to keep the ingest moving.
2. **Specificity.** "Off-the-shelf tags don't speak our taxonomy." Every team wants a different slice: camera motion, contact-rich manipulation, edge cases, locomotion gait. Generic annotators give you generic labels.
3. **Provenance.** "Every clip needs a paper trail." Source, license, capture context, consent: non-negotiable for any defensible training run. Most pipelines bolt this on later. We start with it.
4. **Reuse.** "Today's prepped samples die in S3." Painstaking sample-prep work disappears into a bucket and never gets queried again. The next training run re-does most of it.

---

## Case study — Turning 100,000+ hours of archived footage into training-ready clips

**Case study · Video data provider**

A large video data provider had a massive archive, but the metadata was only useful at the video level. A model lab didn't want full videos. They needed precise 6- to 10-second clips for training, pulled from hundreds of thousands of hours of footage.

The archive already had the raw material. The problem was retrieval. VideoDB processed the footage into scene-level understanding. Existing video-level tags became a starting point, then each scene was indexed with richer context, custom labels, and searchable metadata.

The provider could now search across the full archive, find the exact moments a model lab needed, and extract clips instantly. The clips weren't limited to a fixed duration. Teams could retrieve a 6-second moment, a 10-second sequence, or a longer training sample depending on the use case.

> "We didn't just unlock old footage. We turned a dark archive into a product model labs can search."

What was once a dark archive became a searchable data product. Model labs got the specific video samples they needed for training. The provider got a repeatable way to turn old footage into new revenue. Every new batch added more searchable memory to the archive.

**The opportunity.** Most media archives are sitting on the data model labs want. The issue is that the data is trapped inside long videos, coarse tags, and storage systems built for playback. VideoDB turns those archives into scene-level, searchable, clip-ready datasets.

---

## Case study — A query interface over a multi-petabyte training corpus

**Case study · Searchable training catalog**

### 01 · Aggregation queries — find how many of a specific clip you have

"How many clips do we have with people and a dog, outdoors, no NSFW?" Count and slice the corpus before you plan the training run. The question every dataset planner asks first.

### 02 · Deep search engine — tag filters and natural language, together

Compose structured filters (location, safety, audio class, visual class) with free-form scene descriptions. "Outdoor + safe + violin playing + sunset" returns the exact moments, not just the videos that contain them.

### 03 · Flexible clip length — re-clip without re-encoding

VideoDB doesn't generate a new mp4 for every clip. The training team can sweep clip lengths (2s, 8s, 16s, episode-level) without re-encoding the corpus. A genuine superpower when you're tuning context windows.

### 04 · Modify samples in pipeline — redact, enhance, resize, transcode in one pass

Remove PII (faces, plates, on-screen text), redact restricted content, enhance low-light, resize to the target resolution, transcode to your training format, all in the same pipeline that retrieved the clip. No round-trip to a separate processing job.

---

## For robotics & VLA teams

**Real-world video in. Validated robot data out.**

VideoDB turns robot streams, sim renders, and camera feeds into searchable context for training and evaluation. Use one layer to inspect rollouts, find edge cases, compare real and synthetic data, and export the exact clips your models need.

- **Real-time perception ingest.** Connect RTSP feeds, robot cameras, desktop streams, and sim renders. Index fresh video as it arrives.
- **VLA and world-model validation.** Wrap model outputs as indexes. Score rollouts, catch regressions, and surface edge cases.
- **Sim2real bridge.** Search real and synthetic episodes through one layer. Export reusable slices for Isaac Sim, Newton, and MuJoCo.

---

## Our approach

**We embed. Your data prep gets fast. Your samples stay queryable forever.**

A research-grade partnership, not a vendor relationship. We've built this twice. We know the failure modes.

1. **Phase 01 · Audit.** We sit with your team for a week. Map your taxonomy, your storage layout, your eval needs, your gaps. Leave with a concrete pipeline brief.
2. **Phase 02 · Build.** Custom labeling models. Indexes wired to your taxonomy. Human-in-the-loop for the long tail. Provenance and license trail attached to every clip. Immutable.
3. **Phase 03 · Hand off.** A searchable, versioned dataset on your infra, reusable across training runs. The same indexes scale to every new batch you ingest. No more raw-bucket dead weight.

---

## Under the hood — capabilities

From raw footage to training-ready data. The pipeline the modeling team would otherwise build by hand: standardised, reproducible, audited.

- **Petabyte ingest.** Files, datasets, RTSP captures. Throughput tuned for corpus-scale pipelines.
- **Quality scoring & dedup.** Per-clip quality, near-duplicate detection. Train on what's worth training on.
- **Scene & event segmentation.** Reproducible scene and event boundaries you can slice the corpus against.
- **Custom labeling models.** Bring your taxonomy. Your labeling model wraps cleanly as an index.
- **Provenance trail.** Source, license, capture context attached to every clip. Immutable.
- **Lab-grade reproducibility.** Versioned slices, run logs, deterministic exports. Auditable training runs.

---

## Built in our own lab

**Validated on real video workloads.**

VideoDB comes from our own work in multimodal retrieval, evaluation, and video data preparation. When we work with you, we bring patterns already tested on large archives, model datasets, and production workflows.

- **Research note — Evaluate VLMs on your own video data.** A practical playbook for benchmarking vision-language models against your corpus. [Read the post →](https://labs.videodb.io/research)
- **Inside the lab — What we're building now.** Open notes on retrieval, eval design, sample efficiency, and video-language alignment. [Visit the lab →](https://labs.videodb.io/)

---

## Closing

Bring your corpus. Ship a training-grade dataset in days. Some of the largest video data providers run on this pipeline.

CTAs: [Book a consultation](/company#contact) · [See the platform](/platform)

---

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /world-model-data.md -->
# World Model Training Data — VideoDB

> Training data infrastructure for the physical AI era. World models, robotics, and autonomy don't need another upload tool. They need structured video at scale, with provenance.

**Status:** World Model Data · Program in market · Partner-first

---

## Hero stats

- **PB-scale** — Multi-million-hour corpora across cloud and partner sources
- **Lab-grade** — Reproducibility, deduplication, quality scoring per clip
- **Provenance** — Source, license, capture context — every clip carries the trail

CTAs: [Talk to us](/company#contact) · [See the platform](/platform)

---

## The problem

Getting from raw footage to training-ready data shouldn't take a quarter. World model and physical AI pipelines need scale, structure, and provenance — and most teams build it by hand.

1. **Scale problem.** Multi-million-hour corpora crack internal pipelines.
2. **Structure problem.** Models need scenes, events, quality tiers. Upload tools don't do that.
3. **Provenance problem.** Source, license, capture context — every clip needs a paper trail.

---

## Built for

Teams training models on the physical world. Partnership-first. Data infrastructure partner, not a competing model lab.

- **World model labs** — Curated, structured video at the scale and quality model training requires.
- **Robotics & autonomy** — Filter for the events, scenes, and edge cases that matter to the policy.
- **Simulation** — Ground simulation against real-world video — searchable, structured.
- **Video data providers** — Productize raw footage as queryable, licensable datasets.
- **Research consortia** — Multi-party datasets with consistent structure, access controls, provenance.
- **Internal data platforms** — Replace bespoke labeling and curation with one platform.

---

## Capabilities — from raw footage to training-ready data

The pipeline the modeling team would otherwise build by hand — standardised, reproducible, audited.

- **Petabyte ingest.** Files, datasets, RTSP captures — throughput tuned for corpus-scale pipelines.
- **Quality scoring.** Per-clip quality, dedup, near-duplicate detection. Train on what's worth training on.
- **Scene segmentation.** Reproducible scene + event boundaries you can slice the corpus against.
- **Event labeling.** Bring your taxonomy. Indexes as code — your labeling model wraps cleanly.
- **Provenance trail.** Source, license, capture context attached to every clip — immutable.
- **Lab-grade reproducibility.** Versioned slices, run logs, deterministic exports. Auditable training runs.

---

## The shortest path

Bring your corpus. Get a queryable, training-ready dataset. The data infra you would otherwise build — configured for your scenes, events, and provenance schema.

```bash
videodb dataset create --schema robotics.yml --source s3://your-corpus/
```

---

## Two tracks for the world-model wedge

### Research track

Co-build the training pipeline for a frontier world model.

For lab teams · physical AI · robotics · autonomy.

We embed an engineer in your team. Your model wraps as an index. Reproducible slices, audit logs, deterministic exports. [Talk to us →](/company#contact)

### Partner track

Productize your corpus as a queryable, licensable dataset.

For data providers · sovereign clouds · research consortia.

VideoDB sits on your hosting as the structured-video layer. Sovereign cloud partnership in market. [Read the brief →](/company#partners)

---

## Closing

Stop building data infrastructure. Start training models. Partner-first. Lab-grade reproducibility. Provenance per clip.

CTAs: [Talk to us](/company#contact) · [See the platform](/platform)

---

© 2026 VideoDB, Inc. · videodb.io · hello@videodb.io

---

<!-- source: /platform/personalized-streaming.md -->
# Personalized streaming

> A reference technical solution for composing per-user HLS feeds from indexed moments. Retrieve clips by query, compose them on a chunk-level timeline, apply branding and subtitles as overlay tracks, publish as an HLS manifest. No re-encode step.

---

## Hero

Path: Platform → Solutions → Personalized streaming

Flow: data (video) → transform → stream

A reference technical solution for composing per-user HLS feeds from indexed moments. Retrieve clips by query, compose them on a chunk-level timeline, apply branding and subtitles as overlay tracks, publish as an HLS manifest. No re-encode step.

CTAs: [Try the SDK](/developers) · [Talk to architecture](/company#contact)

---

## Architecture

**The composition pipeline.**

Request resolves against the indexes already built on your collection. Selected moments are composed on a Timeline; overlays apply at the manifest layer; output is an HLS m3u8 URL with per-segment caching.

Stages: REQUEST (source video + "Create highlights of this match" prompt) → MEMORY (scan the source, 2 moments detected) → TRANSFORM (Create clips, Add logo, Add subtitles, Change aspect ratio) → NEW STREAM (hls://highlights.m3u8, focal).

---

## The three stages

**Retrieve. Compose. Stream.** Each stage is a method on the SDK. No infrastructure to operate. No re-encoding pipelines to run.

### 01 — Retrieve from memory

Query the indexes you've already built. Filters, semantic recall, structured constraints — all in one call. Returns ranked moments with playable clip URLs.

```python
# pull moments
moments = vdb.retrieve("goals by player_42")
```

### 02 — Compose the timeline

Merge the moments. Lay overlays on top. Add multilingual subtitles. Reframe for the target aspect ratio. All chunk-level operations. No re-encode.

```python
# compose
tl = vdb.timeline(moments)
tl.overlay(logo="brand.png")
tl.subtitle(lang="en")
```

### 03 — Stream the result

Output is an HLS stream URL. Embeddable in any player, cacheable at the CDN edge, addressable per-user. Generated fresh, on demand.

```python
# publish
stream = tl.stream(format="hls")
# hls://cdn/u/42/highlights.m3u8
```

---

## Full implementation

A complete personalized highlight reel in under 20 lines.

```python
from videodb import connect

vdb = connect(api_key="<your-key>")

# 1. Retrieve. Query the indexes you already built.
#    Filters narrow scope; the index does the ranking.
moments = vdb.retrieve(
    query="goals scored by player_42",
    collection="season_2026",
    filters={"team": "home", "crowd_peak": True},
    limit=20,
)

# 2. Compose. Chunk-level edits, no re-encode.
timeline = vdb.timeline(moments)
timeline.overlay(image="brand.png", position="top-right", opacity=0.85)
timeline.subtitle(language="en", style="caption")
timeline.reframe(aspect="9:16")  # vertical for mobile

# 3. Stream. Fresh HLS URL, cacheable per-user.
stream = timeline.stream(format="hls", ttl="7d")

print(stream.url)
# → https://cdn.videodb.io/u/42/highlights.m3u8
```

---

## Data flow

**What actually happens between request and stream.** The whole pipeline executes inside VideoDB — no round-trip to your services, no re-encoding step.

| Stage | Detail | Note |
|---|---|---|
| REQUEST | user_42 + context | |
| INDEX SCAN | scenes, speech, embeddings, tags | ~90 ms |
| TIMELINE | cut, merge, overlay, caption | chunk-level |
| MANIFEST | HLS playlist + overlay layer | 0 re-encode |
| CDN EDGE | segment cache, per-user TTL | multi-region |
| PLAYER | .m3u8 embedded | |

End-to-end budget: under 1 second for first segment.

---

## What to know before you ship

- **Latency — Sub-second to first segment.** Retrieve is ~120 ms p95. Compose is chunk-pointer math. No re-encode. First HLS segment is served while later segments resolve in the background.
- **Caching — Per-user TTL at the edge.** Each personalized manifest gets a unique URL; CDN segments shared across users are deduplicated. Per-user TTL keeps invalidation cheap.
- **Overlays — Burned at the manifest layer.** Brand marks, subtitles, and reframes are applied as overlay tracks on the HLS manifest. Same source bits; new visual surface.
- **Cost — Usage-based, not minute-based.** Pay for retrievals + unique segments served. No charge for compose-time CPU. The format makes the cuts free.

---

## Use cases

What teams ship with this solution.

- **Sports highlight reels.** Per-fan highlights composed from live games and seasons of archive. Crowd-peak audio detection, player tracking, scoreboard overlay rendering at the manifest layer.
- **Episodic recap for OTT.** "Previously on" cuts generated per viewer based on the last episode watched and which character arcs they're following. Composed inline; output is a per-user HLS manifest.
- **Dynamic ad insertion.** Slot ads at scene boundaries, not on a fixed clock. Brand-safety and audience filters apply at compose time so the same source produces different manifests per viewer.
- **Agent-generated explainers.** An agent pulls source clips from a knowledge base, composes them with synchronized subtitles, and streams the result back to the user. The agent owns the query; VideoDB owns the render.

---

## Closing

Start streaming personalized video in minutes. Same SDK as the rest of the platform. Same six primitives underneath.

CTAs: [Start building](/developers) · [See Realtime alert](/platform/realtime-alert)

---

<!-- source: /platform/realtime-alert.md -->
# Realtime alert

> A reference technical solution for sub-second event detection on live feeds. Attach an RTSP source, declare an event in natural language or as a composed index, subscribe to alerts. Every payload carries the matching clip, frame, and confidence score.

---

## Hero

Path: Platform → Solutions → Realtime alert

Flow: live feed → understanding → alert

A reference technical solution for sub-second event detection on live feeds. Attach an RTSP source, declare an event in natural language or as a composed index, subscribe to alerts. Every payload carries the matching clip, frame, and confidence score.

CTAs: [Try the SDK](/developers) · [Talk to architecture](/company#contact)

---

## Architecture

**The detection pipeline.**

RTSP frames are normalized on ingest, sampled at a configurable rate, and run through the index stack (including BYO classifiers). A match builds an alert envelope (event, confidence, clip URL, frame, source tags) and dispatches it via your subscriber.

Stages: CAPTURE (RTSP, cam-04 aisle, 30 fps, h.264, live) → PERCEPTION (indexes · custom event · BYO classifiers, on normalized frames sampled on ingest) → EVENT LOG (e.g. `fall_in_aisle`, conf 0.94, ~110 ms; clip URL, frame, source tags) → POST (subscriber webhook to /ops/page + clip.m3u8).

---

## The three stages

**Attach. Define. Subscribe.** Three SDK calls. No frame pipelines to operate. No model serving to manage.

### 01 — Attach a live source

RTSP, ONVIF, RTMP, screen + audio capture, WebRTC. One ingestion path; VideoDB handles the codec normalization and clock alignment.

```python
# attach
feed = vdb.stream.attach(
    rtsp="rtsp://cam-04"
)
```

### 02 — Define the event

In natural language or as a composition of indexes. Confidence threshold, cooldown, scope. The platform compiles it into a continuous index pass.

```python
# define
feed.define_event(
    name="fall",
    prompt="a person falls"
)
```

### 03 — Subscribe to alerts

Webhook, WebSocket, or pager. Every alert carries the matching clip URL, the frame timestamp, and the source context. Auditable by default.

```python
# subscribe
feed.on_event(
    webhook="https://ops/page"
)
```

---

## Full implementation

A complete fall-detection pipeline across a 1,000-camera fleet.

```python
from videodb import connect

vdb = connect(api_key="<your-key>")

# 1. Attach the fleet. Same call shape per camera.
#    VideoDB normalizes codecs and clocks on ingestion.
for cam in fleet:
    feed = vdb.stream.attach(
        rtsp=cam.url,
        tags={"site": cam.site, "zone": cam.zone},
    )

    # 2. Define the event. Natural language or composed indexes.
    feed.define_event(
        name="fall_in_aisle",
        prompt="a person falls or slips on the floor",
        custom_index="action_classifier_v3",  # BYO model
        confidence=0.85,
        cooldown="30s",  # debounce repeats
    )

    # 3. Subscribe. Every alert ships a clip + context.
    feed.on_event(
        name="fall_in_aisle",
        webhook="https://ops.example.com/page",
        include=["clip_url", "frame_url", "site_tags"],
    )

# Each alert delivered as: {event, conf, ts, clip_url, frame_url, tags}
```

---

## Data flow

**What runs continuously, per camera, per frame.** The entire pipeline runs inside VideoDB. Your application only sees the alerts.

| Stage | Detail | Note |
|---|---|---|
| RTSP IN | cam-04, 30 fps, h.264 + aac | |
| EXTRACT | keyframes + audio chunks | @ 5 Hz default |
| INDEX RUN | scene, object, custom (BYO) | ~180 ms p95 |
| MATCH | confidence threshold, cooldown | debounced |
| ALERT BUILD | + clip.m3u8, + frame.jpg, + tags | |
| DELIVER | webhook, WebSocket, pager | |

End-to-end alert latency: ~280 ms typical.

---

## What to know before you ship

- **Latency — Sub-second, end to end.** Extract + index + match + deliver typically lands under 300 ms. Configurable sample rate (default 5 Hz) trades latency vs. cost.
- **Cooldown — Built-in debounce.** A defined event won't refire for its cooldown window. Prevents alert storms when the underlying condition persists across many frames.
- **Evidence — Every alert is auditable.** Each alert ships with the matching clip URL, a still frame, the confidence score, and your source tags. Forwardable to a SIEM or case-management system.
- **Fleet — Scales to 1,000+ feeds.** Same SDK shape per camera; the platform handles fan-out, GPU scheduling, and back-pressure. Per-camera or fleet-wide event policies.

---

## Use cases

What teams ship with this solution.

- **Industrial safety.** Fall detection in warehouses, PPE compliance on factory floors, hazard alerts on construction sites. Custom-trained classifiers plug in as indexes; the platform runs them on every frame.
- **Surveillance and security.** Loitering, intrusion, perimeter breaches across fleets of cameras. Each alert lands as a webhook with the matching clip URL and the source frame, ready for triage in a SOC.
- **Healthcare operations.** Patient-fall detection in hospital rooms with per-room confidence thresholds and cooldown windows. Alert payloads carry the clip, frame, and ward context for the on-call nurse.
- **Broadcast moderation.** Detect unsafe content, sponsor-logo violations, and brand-safety flags across live broadcast feeds. The standards desk gets a clip plus the frame; the platform handles cooldowns per policy.

---

## Closing

Wire your camera fleet to alerts in an afternoon. Same SDK as the rest of the platform. Same six primitives underneath.

CTAs: [Start building](/developers) · [See Personalized streaming](/platform/personalized-streaming)

---

<!-- source: /research.md -->
# Research — VideoDB

> Papers, evaluations, and technical reports from the VideoDB team on video retrieval, persistent visual memory, and multimodal models.

---

## Technical reports

### [Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence](https://videodb.io/research/search-over-the-visual-world)
Nagaonkar, Garg, Raj, Choithani, Trivedi · VideoDB Technical Report · July 2026
A formal model of search over continuously growing visual corpora, plus a complete-system comparison across 9,834 queries on four public benchmarks (MSVD, YouCook2, VATEX, MSR-VTT) where a pipeline of general-purpose components beats a commercial video-native engine on macro Recall@1/@3/@10 (73.09/83.39/91.20 vs 65.75/77.13/89.10). [PDF](https://labs.videodb.io/papers/search-over-the-visual-world.pdf)

## Peer-reviewed & preprints

- [Do Thought Streams Matter? Evaluating Reasoning in Gemini Vision-Language Models for Video Scene Understanding](https://arxiv.org/abs/2604.11177). arXiv, April 2026
- [Benchmarking Vision-Language Models on Optical Character Recognition in Dynamic Video Environments](https://arxiv.org/abs/2502.06445). arXiv, February 2025 · [code + dataset](https://github.com/video-db/ocr-benchmark)

## Research notes

- [What Video Retrieval Benchmarks Get Wrong About Ground Truth](https://videodb.io/blogs/video-benchmark-ground-truth-ambiguity). July 2026
- [JEPA: From Language Models to World Models](https://videodb.io/blogs/jepa-from-language-models-to-world-models). July 2026
- [How to Evaluate Multimodal VLMs for Your Video Use Case](https://videodb.io/blogs/how-to-evaluate-multimodal-vlms-for-your-video-use-case). May 2026
- [Strong VLMs Can Still Fail on Downstream Vision Tasks](https://videodb.io/blogs/claude-chessboard-spatial-reasoning). May 2026

---

<!-- source: /research/search-over-the-visual-world.md -->
# Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

> VideoDB Technical Report · July 2026 · Sankalp Nagaonkar, Rohit Garg, Ankit Raj, Ashish Choithani, Ashutosh Trivedi

[Read the PDF](https://labs.videodb.io/papers/search-over-the-visual-world.pdf)

---

## Abstract

Most video-retrieval systems assume a bounded corpus and return ranked files or timestamps. Agents operating over cameras, screens, streams, and archives face a different systems problem: observations arrive continuously; different models interpret them at different temporal granularities; useful context must be selected without replaying the complete visual record; and a result must remain connected to inspectable source evidence. We argue that search over such a corpus is an infrastructure problem that cannot be reduced to ranking video files. We develop a conceptual and formal model of search over the visual world built on analyzer-defined scenes, persistent understanding artifacts, visual memory as coexisting scene spaces over shared source time, and capability-declared indexes; we draw a strict distinction between memory (everything retained), context (what is selected for a task), and evidence (the source intervals that ground it). The VideoDB data format (VDB) realizes this conceptual model in production, and a typed search surface exposes it: planned retrieval, a bounded stateful investigation mode, direct semantic, structured, and aggregate access, and grounded synthesis. We contrast this model-agnostic infrastructure (segmentation, sampling, model choice, embeddings, and ranking all exposed as system decisions, with live streams as first-class sources) with video-native foundation models offered as fixed APIs. In a complete-system semantic-retrieval comparison against a commercial video-native retrieval engine spanning 9,800+ natural-language queries over four public datasets, a pipeline of general-purpose components, none trained end-to-end for video retrieval, achieves higher macro-averaged Recall@1/@3/@10 (73.09/83.39/91.20 versus 65.75/77.13/89.10), while the baseline is higher at Recall@50 (96.42 versus 96.07). The results suggest that, today, retrieval quality over the visual world is governed more by system design than by video-specific pretraining, and that visual-memory infrastructure can deliver it while keeping source-grounded, playable evidence a first-class system object.

## Headline results (Table 9, complete-system comparison, 9,834 queries)

| Dataset | System | R@1 | R@3 | R@10 | R@50 |
|---|---|---|---|---|---|
| MSVD | VideoDB | 70.10 | 80.73 | 89.17 | 96.39 |
| MSVD | TwelveLabs | 67.86 | 78.22 | 89.48 | 97.10 |
| YouCook2 | VideoDB | 65.96 | 80.98 | 93.48 | 97.87 |
| YouCook2 | TwelveLabs | 47.07 | 65.56 | 84.18 | 96.41 |
| VATEX | VideoDB | 83.46 | 90.28 | 95.24 | 97.77 |
| VATEX | TwelveLabs | 85.43 | 92.40 | 97.30 | 99.43 |
| MSR-VTT | VideoDB | 72.82 | 81.55 | 86.89 | 92.23 |
| MSR-VTT | TwelveLabs | 62.62 | 72.33 | 85.44 | 92.72 |
| **Macro-average** | **VideoDB** | **73.09** | **83.39** | **91.20** | 96.07 |
| **Macro-average** | **TwelveLabs** | 65.75 | 77.13 | 89.10 | **96.42** |

Honest reading: the general-purpose pipeline leads at the early ranks that matter for agents (R@1/@3/@10); the video-native baseline is slightly ahead at R@50 and stronger on VATEX.

## Cite

```bibtex
@techreport{videodb2026searchvisualworld,
  title  = {Search over the Visual World: Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence},
  author = {Nagaonkar, Sankalp and Garg, Rohit and Raj, Ankit and Choithani, Ashish and Trivedi, Ashutosh},
  institution = {VideoDB},
  year   = {2026},
  month  = {July},
  url    = {https://videodb.io/research/search-over-the-visual-world}
}
```

---

<!-- source: /blogs.md -->
# Blog

> Welcome to everything that's happening at VideoDB — from product updates and partnerships to engineering deep dives and tutorials.

---

## Latest posts

- [Search over the Visual World (technical report)](/blogs/search-over-the-visual-world.md) — Research
- [NFL Game Analysis: Cutting VLM Hallucinations by 80%](/blogs/nfl-game-analysis-vlm-hallucinations.md) — Engineering
- [Conference Slide Extraction: Search What Was On the Screen](/blogs/conference-slide-extraction.md) — Tutorials
- [Claude + VideoDB Skills Edited Our Launch Video. Then It Watched Its Own Cut.](/blogs/claude-edited-our-launch-video.md) — Product
- [Give Your AI Agents Eyes: A Builder's Guide to Real-Time Visual Perception](/blogs/give-your-ai-agents-eyes.md) — Product
- [Your AI Agent Can't Watch YouTube. Here's the Fix.](/blogs/ai-agent-watch-youtube.md) — Tutorials
- [RTSP + AI: Turn Any Camera Stream into Agent-Readable Events](/blogs/rtsp-ai-analysis.md) — Tutorials
- [Video RAG: The Definitive Guide (2026)](/blogs/video-rag.md) — Engineering
- [Twelve Labs Alternatives in 2026: An Honest Guide](/blogs/twelve-labs-alternatives.md) — Product
- [TinyFish × VideoDB: The Internet, Finally Visible to Your Agents](/blogs/tinyfish.md) — Partnerships
- [VideoDB × TwelveLabs: Search Any Moment Across Your Video Library](/blogs/twelvelabs.md) — Partnerships
- [VideoDB × LlamaIndex: Plug Video Into Your RAG Pipeline](/blogs/llama-index.md) — Partnerships

Categories: All · Product · Partnerships · Engineering · Tutorials · Company

---

<!-- source: /blogs/search-over-the-visual-world.md -->
# Search over the Visual World

> Persistent Visual Memory, Layered Indexes, and Source-Grounded Evidence

Sankalp Nagaonkar · Rohit Garg · Ankit Raj · Ashish Choithani · Ashutosh Trivedi  
VideoDB · {sankalp, rohit, ankit, ashish, ashu}@videodb.io  
Technical Report · July 2026

Category: Research

---

**The canonical document is the PDF**, published and revised at
<https://labs.videodb.io/papers/search-over-the-visual-world.pdf>.
Read it there rather than relying on any summary of it, including this page.

Search over the visual world cannot be reduced to ranking video files. The report develops the
infrastructure it does require — analyzer-defined scenes, persistent visual memory,
capability-declared indexes, and evidence that stays playable at the original source — and
evaluates it against a commercial video-native retrieval engine over 9,834 natural-language
queries drawn from four public datasets.

## What the report covers

1. **Introduction** — why agents over cameras, screens, and archives face a different retrieval problem.
2. **The Shape of the Problem** — where classical retrieval assumptions fail, and the layer missing between models and media.
3. **A Conceptual Model** — sources and source time, analyzers and scenes, understanding artifacts, visual memory, indexes, and the memory / context / evidence distinction.
4. **The VideoDB Data Format (VDB)** — a logical format binding identity, source time, artifacts, provenance, and capability-declared indexes, in seven principles.
5. **Evidence as a Stream** — any temporal selection realized as playable media on demand.
6. **Search over Visual Memory** — one typed search surface: planned retrieval, stateful investigation, direct semantic / structured / aggregate contracts, and grounded synthesis; indexed versus resolved scenes.
7. **Closed-Loop Visual Search** — query-time refinement and representation refinement.
8. **Model or System** — video-native model APIs contrasted with model-agnostic visual data infrastructure, as documented interfaces.
9. **Empirical Study** — a shared-corpus retrieval comparison, plus analyzer-choice and scene-construction studies.
10. **Future and a Research Agenda** — persistent but not yet adaptive memory, native temporal representations, and governance.
11. **Conclusion**

Appendices carry the exact dataset-specific prompts, the detailed reranker study, and the
operational measurement setup.

## Links

- Paper (PDF, canonical): <https://labs.videodb.io/papers/search-over-the-visual-world.pdf>
- Benchmark configurations and reproduction instructions: <https://github.com/video-db/search-over-the-visual-world>
- Open-source Deep Search implementation of the stateful retrieval loop: <https://github.com/video-db/deepsearch>

---

<!-- source: /blogs/nfl-game-analysis-vlm-hallucinations.md -->
# NFL Game Analysis: Cutting VLM Hallucinations by 80%

> Three approaches to event-dense sports footage, measured on the same game. Play-by-play segmentation cut hallucinations from 68.1% to 11.4% and cost up to 70% less than 1 fps into Gemini.

Category: Engineering
Published: 2026-08-25

---

Vision language models shine in controlled benchmarks, then stumble on real-world, event-dense footage such as an NFL game. VideoDB bridges that gap by letting developers slice video at the right semantic boundaries, combine external stats, and run multi-tier visual/LLM pipelines that cut hallucinations by >80% while costing up to 70% less than a naive "1 fps into Gemini" workflow.

Source footage: https://www.youtube.com/watch?v=pA_xAsb5hbA

Four evaluation axes:

| Evaluation metric | What it measures |
| --- | --- |
| Hallucination | Frequency of incorrect or irrelevant information produced by the VLM. |
| Temporal Context | How accurately the VLM maintains correct chronological relationships within the video. |
| Performance on Granular Queries | The VLM's effectiveness in accurately responding to detailed and specific queries. |
| VideoDB Involvement | The extent to which VideoDB's capabilities were leveraged to enhance VLM performance. |

## 1. The naive Gemini approach

Complete NFL game footage sent directly to Gemini.

| Evaluation metric | Observation | Notes |
| --- | --- | --- |
| Hallucination | 68.1% | Frequent irrelevant predictions. |
| Temporal Context | Bloated | Model often lost critical event continuity. |
| Performance on Granular Queries | Moderate | Struggled significantly. |
| VideoDB Involvement | Low | |

Known VLM limitations:

- **Finite context windows.** Even a 1M-token window can't hold one NFL quarter at 30 fps.
- **Image-tile token explosion.** Every 1080p frame splits into ~4-9 tiles (~1-4k tokens) before the model sees it.
- **Weak event reasoning.** VLMs reason per-frame, not per-play, missing temporal causality ("Was the QB still behind the line when he released?").
- **Cost scales linearly** with frames.

## 2. Uniform-length chunks (possible with VideoDB)

Fixed 2s / 5s / 10s clips via VideoDB's scene index API.

```python
import videodb

conn = videodb.connect(api_key="YOUR_API_KEY")
collection = conn.get_collection()

video = collection.upload(url="https://www.youtube.com/watch?v=pA_xAsb5hbA")

# Analyze fixed five-second windows with eight representative frames per window
uniform_understanding = video.understand(
    segmentation={"type": "time", "seconds": 5},
    analyzers=[
        {
            "type": "vlm",
            "name": "uniform_play_analysis",
            "sampling": {"strategy": "uniform", "frame_count": 8},
            "config": {
                "prompt": "Summarize the football action in this five-second segment.",
                "schema": {"summary": "string"},
            },
        }
    ],
)
uniform_understanding.wait_until_complete()

uniform_analyzer = uniform_understanding.get_analyzer("uniform_play_analysis")
uniform_output = uniform_analyzer.get_output()

# Make the summaries searchable
uniform_index = video.index(
    name="uniform_play_analysis",
    source=uniform_analyzer,
    use_for=["semantic", "query"],
    fields={"semantic": ["summary"]},
)
uniform_index.wait_until_complete()

print(f"Understanding ID: {uniform_understanding.id}")
print(f"Index ID: {uniform_index.index_id} ({uniform_index.status})")
print(uniform_output["scenes"][0])
```

| Evaluation metric | Observation |
| --- | --- |
| Hallucination | 74.2% |
| Temporal Context | Insufficient |
| Performance on Granular Queries | Moderate |
| VideoDB Involvement | Moderate |

Uniform chunking scored *worse* than sending the whole game. Arbitrary boundaries split plays (a QB throw lands across two clips), so the model can't see the full action and invents the missing half. Too long means overload, too short means no context — neither extreme works.

## 3. Play-by-play segmentation (advanced pipeline with VideoDB)

Detailed statistical reports for major sports games are public, and include exact start/end timestamps for each play. The problem: those timestamps are on the official *game clock*, not video runtime.

### Aligning game-time with video-time

The on-screen scoreboard is the bridge. It displays scores, quarter, down and yardage, and the game clock.

- **OCR-based timestamp extraction** from the scoreboard throughout the video.
- **Frame sampling optimization**: 1 fps for OCR, cutting compute without losing timestamp accuracy.
- **Timestamp mapping**: OCR results correlate official game time to video runtime, enabling per-play segmentation.

```python
# Analyze one representative frame for every second of video
scoreboard_understanding = video.understand(
    segmentation={"type": "time", "seconds": 1},
    analyzers=[
        {
            "type": "vlm",
            "name": "scoreboard",
            "sampling": {"strategy": "uniform", "frame_count": 1},
            "config": {
                "prompt": (
                    "Read the scorebar at the bottom of the frame. Extract both "
                    "team names and scores, the quarter number, and the game clock."
                ),
                "schema": {
                    "team_1_name": "string",
                    "team_1_score": "integer",
                    "team_2_name": "string",
                    "team_2_score": "integer",
                    "quarter_number": "integer",
                    "game_clock": "string",
                },
            },
        }
    ],
)
scoreboard_understanding.wait_until_complete()

scoreboard_output = scoreboard_understanding.get_analyzer("scoreboard").get_output()

# Map each video-runtime second to its structured scoreboard reading
scene_ocr_results = {
    float(scene["start"]): scene["data"]
    for scene in scoreboard_output["scenes"]
}

for video_time, scoreboard in list(scene_ocr_results.items())[:5]:
    print(video_time, scoreboard)
```

### Integrating play-by-play segmentation with VideoDB

```python
# Step 1: Use the stats PDF to filter all play timestamps (game clock) where a
# catch occurred into `catch_play_scenes` as a list of (start, end) for plays with catches

# Step 2: Map game clocks to video timestamps using OCR outputs

# Analyze short windows once; they will be joined to official play ranges below
catch_understanding = video.understand(
    segmentation={"type": "time", "seconds": 5},
    analyzers=[
        {
            "type": "vlm",
            "name": "catch_analysis",
            "sampling": {"strategy": "uniform", "frame_count": 8},
            "config": {
                "prompt": (
                    "This segment is part of a play containing a catch. Extract the "
                    "catch type, player position, and whether it is an interception."
                ),
                "schema": {
                    "catch_type": "string",
                    "player_position": "string",
                    "interception": "boolean",
                },
            },
        }
    ],
)
catch_understanding.wait_until_complete()
catch_output = catch_understanding.get_analyzer("catch_analysis").get_output()

def overlaps(scene, start, end):
    return float(scene["start"]) < end and start < float(scene["end"])


catch_details = []
for start_time, end_time in catch_play_scenes:
    evidence = [
        scene["data"]
        for scene in catch_output["scenes"]
        if overlaps(scene, start_time, end_time)
    ]
    if not evidence:
        continue

    catch_details.append(
        {
            "start": start_time,
            "end": end_time,
            "catch_type": ", ".join(
                dict.fromkeys(item["catch_type"] for item in evidence if item["catch_type"])
            ) or "none",
            "player_position": ", ".join(
                dict.fromkeys(
                    item["player_position"]
                    for item in evidence
                    if item["player_position"]
                )
            ) or "none",
            "interception": any(item["interception"] for item in evidence),
        }
    )

catch_index = video.index(
    name="catch_plays",
    source=catch_details,
    use_for=["semantic", "query"],
    fields={
        "semantic": ["catch_type", "player_position"],
        "filter": ["interception"],
    },
)
catch_index.wait_until_complete()
```

| Evaluation metric | Observation |
| --- | --- |
| Hallucination | 11.4% |
| Temporal Context | Perfect |
| Performance on Granular Queries | High |
| VideoDB Involvement | High |

## Approach comparison

| Evaluation metric | Naive whole-video | Uniform chunks | Play-by-play |
| --- | --- | --- | --- |
| Hallucination | 68.1% | 74.2% | **11.4%** |
| Temporal Context | Poor | Insufficient | **Perfect** |
| Granular Queries | Moderate | Moderate | **High** |
| VideoDB Use | Low | Moderate | **High** |

## Key takeaways

1. **Define key sports concepts** — catch (yes/no), running play (yes/no), scoring event (yes/no).
2. **Check availability of statistical data.** Available: use it to isolate plays. Not available: use the VLM directly for visual extraction.
3. **Extract relevant plays using statistical data** via the VideoDB timeline; record timestamps and metadata.
4. **Run visual analysis with VideoDB indexing** — pass extracted scenes to the VLM for detail (catch type "overhead", position "near sidelines").
5. **Structure the output data clearly:**

```json
[
  {
    "play_start_time": 12,
    "play_end_time": 52,
    "details": {
      "catch": true,
      "type": "overhead",
      "position": "near sidelines",
      "interception": false,
      "running_play": true
    }
  }
]
```

6. **Add a query and reasoning engine (small LLM).** Feed structured data plus the user query into the VideoDB search interface for accurate play-by-play results.

## Pricing: VideoDB vs. Gemini at 1 fps

| 60-min NFL game | Frames analysed | VideoDB (Balanced tier) | Gemini 1.5 Pro* |
| --- | --- | --- | --- |
| 1 fps, 1080p | 3,600 | **$2.00** index + ~$0.35 tokens | $1.1 - $7.4 |
| 5 fps | 18,000 | **$10.00** index | $5.6 - $37.0 |
| 30 fps | 108,000 | **$12.00** index | $33 - $220 |

*Prices use Google's published rate card: $0.10/M input tokens, $0.40/M output; HD frames tokenize into 1,024-4,128 tokens each.*

As frame rate or resolution rises, VideoDB's flat visual-index pricing stays predictable while pure-Gemini costs explode.

## Why choose VideoDB

- **Event-aligned indexing** — cut by play, scene, or any custom timeline, not crude 1s slices.
- **Hybrid reasoning pipelines** — blend stats, embeddings, and VLMs to slice hallucinations to ~11%.
- **Serverless scale** — petabytes or a single clip, zero idle cost.
- **Developer-first API** — Python, JS, REST.
- **Transparent pricing** — pay once for storage + index; pick Entry / Advanced / SOTA LLM pricing per query.

## FAQ

**Why do VLMs hallucinate on sports footage?** Context windows can't hold a game at broadcast frame rates, every 1080p frame explodes into thousands of tokens, and VLMs reason per-frame rather than per-play. Naive whole-video hallucinated on 68.1% of events.

**Do smaller video chunks reduce hallucinations?** No — uniform 2s/5s/10s chunks scored worse (74.2%). Fixed-length cuts split plays across boundaries. The problem is where you cut, not how small.

**What is play-by-play segmentation?** Segmenting at real semantic boundaries instead of on a timer: official play start/end timestamps from the public game summary, mapped onto video runtime by OCR-ing the on-screen game clock at 1 fps. Hallucinations dropped to 11.4%.

**How do you align the official game clock with video timestamps?** Sample 1 fps, run OCR with a structured prompt returning scores, quarter, and game clock as JSON, and build a lookup table from game clock to video runtime.

**Is VideoDB cheaper than sending frames straight to Gemini?** Comparable at 1 fps. The gap opens as frame rate rises: at 30 fps for a 60-minute game, $12.00 of indexing against $33-$220 of Gemini tokens.

**Does this only work for American football?** No. It generalizes to any domain with an authoritative event log and an on-screen clock: cricket, basketball, soccer, esports, broadcast production.

---

Docs: https://docs.videodb.io/pages/understand/indexing-pipelines/create-an-index · Retrieval architecture: https://videodb.io/blogs/video-rag · Model selection: https://videodb.io/blogs/how-to-evaluate-multimodal-vlms-for-your-video-use-case · Questions: engg@videodb.io

---

<!-- source: /blogs/conference-slide-extraction.md -->
# Conference Slide Extraction: Search What Was On the Screen

> Combine spoken-word search with visual scene indexing to pull slide content out of any conference talk, retrieved by what the speaker was saying at the time.

Category: Tutorials
Published: 2026-08-25

---

When you try to recall a specific part of a talk, usually only a few keywords come to mind — and often what caught your attention was on the *slide*, not in the speech. This tutorial builds a pipeline that stores talks in VideoDB and returns the on-screen slide content, in text form, from a spoken-word query.

Runnable notebook: https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/examples/conference_slide_scraper.ipynb

Example question: what was on the screen when the speaker discussed the "hard and fast rule" in https://www.youtube.com/watch?v=IEe-5VOv0Js

## The approach

Index the video on two modalities, then join them on time:

1. **Spoken content** — transcribe and index the speech, making it keyword-searchable.
2. **Visual content** — shot-based scene extraction, each scene described by a prompt that reads slide text.

At query time, keyword-search the transcript, take the returned time ranges, and keep the scenes that overlap them. The transcript locates the moment; the scene index supplies what was on screen.

## Setup

```bash
!pip install videodb
```

Get an API key from https://console.videodb.io (free for the first 50 uploads, no credit card required).

### Step 1: Connect to VideoDB

```python
import videodb

# Set your API key
api_key = "your_api_key"

# Connect to VideoDB
conn = videodb.connect(api_key=api_key)
coll = conn.get_collection()
```

### Step 2: Upload the video

```python
# Upload a video by URL
video = coll.upload(url="https://www.youtube.com/watch?v=IEe-5VOv0Js")
```

### Step 3: Understand and index both modalities

Run spoken-word and slide analyzers together. Conference videos need a lower shot threshold to capture subtle slide changes, while one representative frame per shot is enough to read a static slide.

```python
slide_prompt = (
    "Extract all text on the presentation slide. "
    "Return None if no slide is visible."
)

understanding = video.understand(
    segmentation={"type": "shot", "threshold": 10},
    analyzers=[
        {"type": "spoken_words", "name": "transcript"},
        {
            "type": "vlm",
            "name": "slides",
            "sampling": {"strategy": "uniform", "frame_count": 1},
            "config": {
                "prompt": slide_prompt,
                "schema": {"description": "string"},
            },
        },
    ],
)
understanding.wait_until_complete()

transcript_analyzer = understanding.get_analyzer("transcript")
slides_analyzer = understanding.get_analyzer("slides")
transcript_output = transcript_analyzer.get_output()
slides_output = slides_analyzer.get_output()
```

Build a sentence-level transcript index so exact phrase matches remain tied to tight time ranges. Index the slide artifact separately for direct visual retrieval.

```python
import re

sentence_records = []
sentence_words = []

for scene in transcript_output["scenes"]:
    for word in scene["data"]["words"]:
        sentence_words.append(word)
        if re.search(r'''[.!?]+["'”’)\]]*$''', word["text"]):
            sentence_records.append(
                {
                    "start": sentence_words[0]["start"],
                    "end": sentence_words[-1]["end"],
                    "text": " ".join(item["text"] for item in sentence_words),
                }
            )
            sentence_words = []

if sentence_words:
    sentence_records.append(
        {
            "start": sentence_words[0]["start"],
            "end": sentence_words[-1]["end"],
            "text": " ".join(item["text"] for item in sentence_words),
        }
    )

transcript_index = video.index(
    name="conference_transcript_sentences",
    source=sentence_records,
    use_for=["query"],
    fields={"filter": ["text"]},
)
slides_index = video.index(
    name="conference_slides",
    source=slides_analyzer,
    use_for=["semantic", "query"],
    fields={"semantic": ["description"], "filter": ["description"]},
)

transcript_index.wait_until_complete()
slides_index.wait_until_complete()
print(transcript_index.status, slides_index.status)
```

### Step 4: Search pipeline implementation

Query the spoken index, extract time ranges, then filter the timed slide artifact by overlap.

```python
scene_index = [
    {
        "start": float(scene["start"]),
        "end": float(scene["end"]),
        "description": scene["data"]["description"],
    }
    for scene in slides_output["scenes"]
]
scene_index.sort(key=lambda scene: (scene["start"], scene["end"]))


def filter_overlapping_scenes(time_ranges, scenes):
    def overlaps(scene, range_start, range_end):
        return scene["start"] < range_end and range_start < scene["end"]

    filtered_scenes = []
    for start, end in time_ranges:
        filtered_scenes.extend(scene for scene in scenes if overlaps(scene, start, end))

    # Remove duplicates while preserving order
    seen = set()
    return [
        scene
        for scene in filtered_scenes
        if not (
            (scene["start"], scene["end"]) in seen
            or seen.add((scene["start"], scene["end"]))
        )
    ]
```

```python
def search_pipeline(query, video):
    transcript_result = video.query(
        index_id=transcript_index.index_id,
        filter=[{"field": "text", "op": "contains", "value": query}],
        limit=100,
        sort=[("start", "asc")],
    )
    time_ranges = [
        (shot.start, shot.end) for shot in transcript_result.get_shots()
    ]

    final_result = filter_overlapping_scenes(time_ranges, scene_index)

    result_text = "\n\n".join(
        result_entry["description"]
        for result_entry in final_result
        if result_entry.get("description", "").lower().strip() != "none"
    )
    result_timeline = [
        (result_entry.get("start"), result_entry.get("end"))
        for result_entry in final_result
    ]

    return result_text, result_timeline
```

### Step 5: Viewing the search results

```python
from videodb import play_stream

query = "hard and fast rule"

result_text, result_timeline = search_pipeline(query, video)

stream_link = video.generate_stream(result_timeline)
play_stream(stream_link)

print(result_text)
```

It returns scenes where the spoken words match the query, plus the content of any slides visible in those scenes, plus a playable stream of only the matching moments.

## Results

Query: **"hard and fast rule"** — the slide reads:

```text
IT'S ALL IN THE DETAILS

- Prefer American English for naming
- Avoid payment-industry jargon
- Timestamp fields should use <verbed_at>
- Amount properties should also provide a currency
- API resources with IDs are top-level
- New API resources should be retrieved and listed one way
- API resource mutations should be reflected in API responses
- Use nested structures for future extensibility
- Prefer enums to booleans for new properties
- Use a type field for polymorphic objects
- Use verbs for properties with side effects
- Use top-level namespaces for product APIs
- Evaluate new features in the Dashboard before building an API
- Use simple, unambiguous language
- Always paginate unbounded lists
- Iterate on designs with beta users with the feature behind a gate
```

Query: **"stripe api review"** — the slide is an API review checklist (title, gavel block for pinging PMs, and a change summary section).

Query: **"Friction Log"** — the slide is a set of internal Terminal dogfooding instructions, including how to order hardware and set up the iOS/Android SDK environment.

## Conclusion

The technique adapts to any case where visual information needs to be retrieved from audio content:

- Finding product demonstrations in long-form video content
- Identifying key moments in educational videos
- Searching for specific visual elements in recorded meetings or presentations

## FAQ

**How do you search for what was on a slide, not just what was said?** Understand and index both channels. Query the transcript to locate the moment, then pull the slide artifact records that overlap those time ranges.

**Why keyword search on the transcript instead of semantic search?** The transcript only locates the moment, it doesn't answer the question. Speakers say the phrase you half-remember, so exact matching gives tight, high-precision ranges.

**What scene extraction threshold works for conference talks?** Lower than the default. Slide decks change gradually, so the default shot detection skips transitions between similar slides. Shot-based extraction with a threshold of 10 captured every slide change.

**How do I keep indexing costs down while tuning the prompt?** Use a short representative test video, inspect `slides_analyzer.get_output()`, and only run the finalized analyzer configuration over the complete talk.

**Can I get a playable clip of the results, not just text?** Yes. Pass the returned time ranges to `generate_stream()` for a single stream of only the matching moments.

**Does this work on live talks and streams?** Yes, with streaming ingestion — the same two-channel pattern runs continuously and the index is queryable while the talk is happening. See https://videodb.io/blogs/rtsp-ai-analysis

---

Understanding Artifacts: https://docs.videodb.io/pages/understand/indexing-pipelines/understanding-artifacts · Create an Index: https://docs.videodb.io/pages/understand/indexing-pipelines/create-an-index · Search and Retrieval: https://docs.videodb.io/pages/understand/search-and-retrieval/natural-language-query · Architecture: https://videodb.io/blogs/video-rag · Discord: https://discord.gg/py9P639jGz

---

<!-- source: /blogs/claude-edited-our-launch-video.md -->
# Claude + VideoDB skills edited our launch video. Then it watched its own cut.

> Raw takes, a hand-drawn deck, and a research paper went in. A subtitled, mastered, post-ready launch video came out. Every edit was decided by understanding.

Category: Product
Published: 2026-08-09

---

Last week we shipped the launch video for our research paper, [*Search over the Visual World*](https://labs.videodb.io/papers/search-over-the-visual-world.pdf). It runs 2 minutes 14 seconds across twelve narrative beats: talking-head takes intercut with a hand-drawn deck and a scroll through the paper itself, a generated music bed, and brand-styled captions burned in. Plus three vertical variants for social.

<!-- TODO: embed final cut (YouTube/CDN) -->

![A frame from the final cut: the research paper scrolling past its benchmark table, with the founder overlay](https://videodb.io/assets/images/Blog-post-preview/claude-edit-paper-scroll.webp)
*A frame from the final cut. The paper scroll is anchored on the benchmark table the agent found on its own.*

No human opened an editor. The editing was done by Claude in a chat session, using VideoDB's [agent skills](https://github.com/video-db/skills) as its eyes and hands. What follows is the receipts, real screenshots from the session. This is the clearest demonstration yet of what we mean when we say *the edit is a query*.

## What went in

Multiple talking-head takes (the usual founder-fumbling-lines footage, recorded on Loom), an animated HTML deck (12 slides, one per beat), the paper PDF, and a brief that fits in one breath: follow this script kit, overlay my takes, show the paper scrolling for credibility, generate a light music bed.

One scoping instruction is worth noting: *"for my videos you can just use the transcript index, there's nothing much visually to do"*. Take-picking is an audio problem, and the human knew it. The agent agreed and indexed the takes transcript-first.

## The first cut: 2:35, with judgment calls a junior editor would miss

The agent created a collection, then uploaded and transcribed the takes in the background. Meanwhile it rendered the deck headless and built a paper-scroll clip. On its own it found the exact results table in the PDF worth scrolling past: it went looking for the benchmark numbers and anchored the scroll on them.

Then it pulled timed transcripts, computed **word-boundary cut points for each beat**, and composited everything server-side on a [VideoDB timeline](https://videodb.io/blogs/infrastructure-that-sees-and-edits). The hook take runs full-frame, the paper cameo sits under the intro, and the deck carries beats 2 to 12 with the founder in a corner bubble. Slides advance on sentence starts, with the ambient bed at 12% volume.

![Session screenshot: the agent delivers the full 2:35 draft with a streaming link and its own notes](https://videodb.io/assets/images/Blog-post-preview/claude-edit-draft-delivery.webp)
*The draft delivery, in the session: 2:35, streamed for review before any MP4 download.*

Two details from the draft delivery that tell you this isn't template assembly:

- *"Math beat uses pass 1 — **pass 2 said '93,000'**."* Two takes covered the same beat. In one of them the founder said the wrong number. The agent read both transcripts, caught the discrepancy against the deck, and cut in the take where the number was right.
- *"Your recordings are HDR (that washed-out look) — all takes were **tone-mapped to SDR** before editing."* Nobody asked. It noticed.

It also flagged its own imperfections unprompted, each with a proposed one-line fix: the hook framing covering a title line, the bubble clipping a chart label.

## Directing an AI editor sounds exactly like directing a human one

The feedback round, verbatim:

> *"move the video overlay to the left bottom it won't mess with the diagrams on the right. The start of second slide starts randomly. The end closing slide and last 3 needs to choose correct take and correct starting points. The Hook should be only my video and no background… reverify the start and end of each cut in timeline and choose the exact start perfectly. Not bad for the first draft"*

No timecodes. No EDL. Director's notes. The agent triaged them the way a good editor would. Four of five were mechanical: bubble to bottom-left, hook full-frame, the slide-2 reveal anchored to a specific spoken line instead of popping mid-beat, and a **silence-aware re-audit of every cut-in and cut-out** so no segment catches the tail syllable of the previous sentence.

The fifth was the closing-take choice. It pulled the candidate takes from the transcripts and put the creative call back where it belongs, with the human.

![Session screenshot: the note 'bw slide 5 and 6 there's a lot of mess' and the agent's frame checks that followed](https://videodb.io/assets/images/Blog-post-preview/claude-edit-directors-note.webp)
*Another round, verbatim: "bw slide 5 and 6 there's a lot of mess." The agent went and looked.*

![Contact sheet of frames the agent sampled from its own render to check slide timing](https://videodb.io/assets/images/Blog-post-preview/claude-edit-frame-sampling.webp)
*How it looked: contact sheets of its own render, sampled to check the slide transitions frame by frame.*

## The part no other editing stack can do

Then came the instruction that makes this a different category of thing:

> *"whenever you have the final video, reupload and do visual index and audio index and verify that everything is perfect from editing and consumption point of view, if not try to replace the messedup portions. Do at least 1 pass of quality"*

![Session screenshot: the reupload-and-verify instruction, the agent's QA contact sheets, and the final cut delivered QA-passed on both indexes](https://videodb.io/assets/images/Blog-post-preview/claude-edit-final-cut-qa.webp)
*The instruction and the delivery: 40 commands later, the final cut ran 2:48, QA-passed on both indexes.*

So the agent rendered, then **uploaded its own cut back into VideoDB, indexed it for spoken words and visuals, and reviewed its own work.** And the pass earned its keep. It found that the draft's cut points were systematically early. Spoken-word timestamps were running about 1.3 to 2.6 seconds ahead of the file timeline, varying per file. That is exactly why some beats started mid-breath.

The fix: re-align every boundary against the real audio using silence fingerprints, re-cut in each file's own timeline, gate-check all seven talking segments with local ASR (first and last word verified), and re-stretch the deck piecewise so all twelve slide changes land on the corrected schedule.

QA on the shipped file. The transcript audit: every beat starts clean, every sentence completes. The visual pass: 12 of 14 windows with zero defects. The other two showed the corner bubble covering the paper's corner, which is inherent to any overlay. Final: **2:48, QA-passed on both indexes.**

Sit with the loop: *ingest → understand → decide → render → re-ingest → verify.* Every other programmatic editing tool stops at "render". The output ships un-watched, because assembly APIs have no eyes. Here, the same perception layer that picked the takes **watched the finished cut and fixed it.**

And because this was real dogfooding, the timestamp-offset finding went into a gaps doc for our product team, with repros. The edit session filed its own bug report.

![Session screenshot: commissioning the gaps doc for the team alongside the vertical output](https://videodb.io/assets/images/Blog-post-preview/claude-edit-gaps-doc.webp)
*Dogfooding for real: the gaps doc commissioned mid-session, with findings and repros straight to the product team.*

## Then the notes you'd give any editor

*"okay now 1.25x the whole video"* → 2:14, tempo-shifted with pitch preserved, dead-air verified gone. It also added an unprompted editor's caution: 1.25x on an already-brisk hook can read rushed. A variable-speed alternative was offered.

*"the audio of the section where I am on full screen the intro is low. can you fix it?"* → measured, not guessed. The hook came from a different recording session and was **7.7 dB quieter** than everything else. It boosted just that window (all beats now within 0.9 dB), then mastered the whole mix from −30 LUFS to **−16 LUFS**. That is the loudness X and LinkedIn feeds expect.

## Captions that became a brand asset

The subtitle request produced the detail we'd show any media team. Instead of shipping raw ASR, the agent rebuilt the caption text as a **corrected canonical script**. So the captions say *Qwen* (not "Quinn"), *TwelveLabs* (not "12 labs"), and *GPT* (not "GPD").

The text was chunked into 1 to 2 word pops (200 of them), butt-joined inside sentences so a word is always on screen. The styling is brand: bold uppercase, white with black outline, numbers in VideoDB orange, 70 ms fade. It was delivered as a custom caption track on the timeline, streamed for review *before* any MP4 was rendered.

![QA frames from the shipped cut with brand captions burned in, the keyword highlighted in VideoDB orange](https://videodb.io/assets/images/Blog-post-preview/claude-edit-final-frames.webp)
*QA frames from the shipped cut, with brand captions burned in and keywords in VideoDB orange.*

And then one more instruction: *"also store this subtitle style for future as videodb brand."* The style is now project memory: the template, the corrections list, and the review-first pipeline. Any future session reproduces it on any new video by asking for **"brand subtitles."** The edit session didn't just produce a video. It produced a reusable house style.

## Three verticals, art-directed in plain language

*"from timeline of the successful one, extract my video and audio but the deck build it specific to 9:16… you can also create a split variant where my video on top and the deck at the bottom (like social media has many such videos) do you get it?"* It got it. Variant A was deck-native 9:16, with a blurred deck texture, cinematic. Variant B was the creator-style split.

Then one more note: *"I want my video down and on 1/4th of the screen and give paper more"*. That produced Variant C: face as a full-width band in exactly the bottom quarter, captions in the seam, and the paper cameo **re-rendered as a portrait crop** so the results tables are actually readable in a phone feed instead of a letterboxed sliver. The deck was re-laid for vertical, not squeezed.

![A vertical variant streamed for review: the deck re-laid for 9:16 with the founder in a portrait band](https://videodb.io/assets/images/Blog-post-preview/claude-edit-vertical-variant.webp)
*A vertical variant streamed for review, with the deck re-laid for 9:16, not squeezed.*

<!-- TODO: embed Variant C vertical (YouTube/CDN) -->

## Why this works (and why "AI editing" tools can't)

Timeline-assembly APIs are hands without eyes: you compute every cut yourself and express it in JSON. Consumer AI editors have eyes but no API, so a human drives. Rough-cut copilots pick takes, but inside human NLE workflows.

The loop above requires something structurally different: **an index that decides, a timeline that renders, and the same index again to verify, because the output is just video.**

Understanding-driven editing. The edit is a query. The QA is a query too.

## Do this with your own footage

This wasn't a bespoke pipeline. It was Claude with VideoDB's agent skills (`npx skills add video-db/skills`). Any agent that can call tools can run the loop: upload takes, index, search for moments, compose, render, re-ingest, verify.

Start with the [quickstart](https://docs.videodb.io/pages/getting-started/quickstart), or read the [builder's guide to giving agents eyes](https://videodb.io/blogs/give-your-ai-agents-eyes). This post is what the "Act" verb looks like at full stretch.

## FAQ

**Can AI actually edit raw footage into a finished video?**
Yes, with understanding-driven infrastructure. The agent transcribes and indexes the takes. It selects the best take per script beat by querying transcripts, including catching factual slips between takes. Then it composes the timeline programmatically, renders, and re-indexes its own output to verify the edit. The session above shipped a QA-passed, mastered, subtitled 2:14 final plus three vertical variants.

**How does the agent pick the "best take"?**
By querying the indexes, not scrubbing footage. Transcripts show which take covers which beat cleanly. When two takes disagree (one said the wrong number), the transcript comparison catches it. Where the choice is genuinely creative, the agent surfaces candidates and asks.

**Can it check its own work?**
Yes, that's the differentiating loop. The rendered video is re-uploaded, indexed on both channels, and audited against the intended structure. Misaligned cuts get re-cut. Assembly-only APIs cannot do this, because they never perceive their output.

**Does it handle subtitles, loudness, speed changes, and vertical formats?**
All demonstrated in this one session: corrected-script captions in a stored brand style, per-beat loudness matching and −16 LUFS mastering, pitch-preserved 1.25x, and three art-directed 9:16 variants.

**Is this video generation?**
No. Every frame is real footage: takes, slides, paper pages. VideoDB is video understanding and editing infrastructure. It doesn't synthesize avatars or synthetic scenes. (The only generated element was the background music bed, which was requested.)

---

Run the loop on your own footage: start with the [quickstart](https://docs.videodb.io/pages/getting-started/quickstart) (free tier), or [talk to us](https://videodb.io/company#contact) about what an editing agent could do on your footage.

---

<!-- source: /blogs/give-your-ai-agents-eyes.md -->
# Give Your AI Agents Eyes: A Builder's Guide to Real-Time Visual Perception

> AI agents can browse, code, and search, but they're blind. How to give your agent real-time visual perception over YouTube videos, live cameras, and screens, in a few lines of Python.

Category: Product
Published: 2026-08-09

---

Your agent can write code, book flights, search the entire web, and operate a browser like a caffeinated intern. Point it at a security camera, a YouTube video, or the screen it's supposedly automating, and it's helpless. It can act on the world, but it cannot watch the world.

**The short version:** AI agents get vision by plugging into video infrastructure. That layer ingests any video source: files, YouTube URLs, live RTSP cameras, screen capture. It converts them continuously into indexed, searchable context, and exposes search, events, and alerts through an API. That's VideoDB, and it gives your agent eyes plus the memory to act on what it sees, in a few lines of Python or TypeScript.

## Agents got hands before they got eyes

Exa and Parallel gave agents web search and research. Firecrawl gave them clean page context. Browserbase and browser-use gave them hands. TinyFish runs fleets of web agents. Every one of those unlocked a product category. And every one operates on text and DOM.

Meanwhile, most of what actually happens in the world never touches a DOM. Meetings happen on camera. Work happens on screens. Operations happen in front of cameras.

The largest knowledge source on the internet is a video platform, and your agent can't watch it. That gap is not a model problem. GPT-5, Claude, and Gemini can all describe a frame beautifully. It's an infrastructure problem.

## Why screenshots don't add up to sight

The screenshot-in-prompt hack collapses for five reasons:

1. Moments aren't events. Meaning in video lives in time.
2. Nothing persists. Perception without memory is a party trick.
3. Tokens explode. You re-buy the same understanding on every question.
4. Polling misses things. Real-time perception has to be push, not pull.
5. You can't search what you never indexed.

Sight, for an agent, is a pipeline: **ingest → understand → remember → retrieve → act.**

## Build 1: An agent that actually watches YouTube

```python
import videodb

conn = videodb.connect()
coll = conn.get_collection()

video = coll.upload(url="https://www.youtube.com/watch?v=LPZh9BOjkQs")

understanding = video.understand(analyzers=[
    {"type": "spoken_words", "name": "transcript"},
    {"type": "vlm", "name": "scene", "config": {"prompt": "Describe the visual content."}}
])
understanding.wait_until_complete()

transcript_index = video.index(name="transcript", source=understanding.get_analyzer("transcript"))
scene_index = video.index(name="scene", source=understanding.get_analyzer("scene"))
transcript_index.wait_until_complete()
scene_index.wait_until_complete()

results = video.semantic_search(
    query="an explanation supported by diagrams and on-screen text",
    index_ids=[transcript_index.index_id, scene_index.index_id],
    top_k=5
)
for shot in results.get_shots():
    print(f"{shot.start:.1f}s-{shot.end:.1f}s")

evidence_url = results.compile()  # playable evidence reel
```

The visual channel is indexed too (the VLM prompt is your extraction schema), search runs across both modalities, every hit is timestamped, and `compile()` returns a watchable clip. Your agent doesn't summarize a video. It cites it.

## Build 2: An agent with a live camera feed

Connect an RTSP stream, run continuous understanding in rolling windows, and index it live. Then define events in plain language, like "Detect when a person appears in the monitored area." Alerts arrive over WebSocket or webhook with a label, confidence, explanation, and a playable `stream_url` of the moment.

The feed is continuously understood whether or not anyone is asking questions. That is the difference between an agent that *can look* and an agent that *is watching*. Full code: https://videodb.io/blogs/rtsp-ai-analysis

## Build 3: Agents that remember what happened on screen

VideoDB's capture SDK (`pip install "videodb[capture]"`) turns a desktop or browser session into the same kind of stream: captured, continuously understood, indexed, searchable. An agent can ask questions about its own visual history. Call it episodic memory, in the concrete sense.

## One layer, every source

| Your agent | Its eyes | What it can now do |
|---|---|---|
| Research / web agent | YouTube + any video URL | Watch, extract, cite with timestamped clips |
| Computer-use agent | Screen capture | Recall and replay anything it ever saw |
| Browser agent | Session recordings, web video | Verify visually, debug from replays |
| Ops / monitoring agent | RTSP cameras, drones | Watch continuously, act on plain-language events |
| Meeting copilot | Calls and meetings | Search what was shown, not just said |
| Media agent | Archives and libraries | Find any moment, compile new cuts programmatically |

## "Can't I just send frames to GPT-5 or Gemini?"

For a single image, absolutely. Raw VLM calls are the right tool for one-shot perception of individual moments.

You outgrow them the day your agent needs continuous sources, search over hours of footage, push events, memory that outlives the session, or costs that don't scale linearly with watching. VideoDB isn't a competitor to the models. It's the layer that feeds them.

## Plug it into the agent you already have

- Agent Skills: `npx skills add video-db/skills` for Claude, Cursor, and other agents (https://github.com/video-db/skills)
- Frameworks: LlamaIndex retriever, LangChain, REST
- No-code: n8n and Zapier

## FAQ

**How do AI agents see video?** Through a perception layer: video infrastructure that ingests sources, runs speech and vision analyzers continuously, indexes the results, and exposes semantic search plus real-time events via API.

**Can my agent watch a YouTube video and answer questions about it?** Yes. Ingest the URL, index speech and visuals, search semantically, and get timestamped matches and evidence clips. About 20 lines of Python.

**Can agents process live streams like RTSP cameras in real time?** Yes. RTSP Connect, rolling-window understanding, live indexes, plain-language events, and WebSocket/webhook alerts.

**What's the difference between VideoDB and a vision model API?** The model perceives frames. VideoDB is the model-agnostic infrastructure around it: ingestion, segmentation, indexing, memory, retrieval, events, playback.

**Does this work with Claude, Cursor, LangChain, or my agent framework?** Yes. Agent Skills, a LlamaIndex retriever, Python/Node SDKs, and n8n/Zapier.

**What about privacy and deployment?** Zero Data Retention options, SOC 2 Type II, SSO, managed cloud or BYOC (AWS/GCP/Azure), VPC and edge.

---

The next wave of AI will understand the visual world. Give your agents eyes: https://docs.videodb.io/pages/getting-started/quickstart

---

<!-- source: /blogs/ai-agent-watch-youtube.md -->
# Your AI Agent Can't Watch YouTube. Here's the Fix.

> Give your agent YouTube videos as searchable, timestamped context in about 20 lines of Python, with speech and visuals indexed and evidence clips compiled.

Category: Tutorials
Published: 2026-08-09

---

Give an agent a paper, a blog post, or an entire website and it's brilliant. Give it the URL of a 40-minute conference talk and it shrugs. The fix: an AI agent "watches" a YouTube video by ingesting the URL into video infrastructure that analyzes both what's said and what's shown. That infrastructure indexes the results with timestamps and exposes semantic search. The agent gets back not just answers, but the exact moments that prove them, playable as a clip.

## The stack everyone builds first (and regrets)

yt-dlp to fetch, Whisper to transcribe, ffmpeg to sample frames, a VLM to caption them, a vector DB to store both, plus glue code to keep timestamps aligned. That is five dependencies, two of which break when YouTube changes something. And the output is still just text *about* the video.

The failure isn't any single tool. It's that you're rebuilding video infrastructure to answer one question.

## The 20-line version

```python
import videodb

conn = videodb.connect()
coll = conn.get_collection()

# 1. Ingest straight from the URL
video = coll.upload(url="https://www.youtube.com/watch?v=LPZh9BOjkQs")

# 2. Understand both channels: speech AND visuals
understanding = video.understand(analyzers=[
    {"type": "spoken_words", "name": "transcript"},
    {"type": "vlm", "name": "scene",
     "config": {"prompt": "Describe the visual content, including any text, charts, or product shots on screen."}}
])
understanding.wait_until_complete()

# 3. Index
t_index = video.index(name="transcript", source=understanding.get_analyzer("transcript"))
s_index = video.index(name="scene", source=understanding.get_analyzer("scene"))
t_index.wait_until_complete()
s_index.wait_until_complete()

# 4. Ask
results = video.semantic_search(
    query="pricing or cost claims",
    index_ids=[t_index.index_id, s_index.index_id],
    top_k=5
)
for shot in results.get_shots():
    print(f"{shot.start:.1f}s-{shot.end:.1f}s")

clip_url = results.compile()  # playable reel of just the matching moments
```

The VLM prompt is yours, so it works as a schema for what to extract. Search runs across speech and visuals together, so "the part where they show the dashboard" works even if nobody said "dashboard." And `compile()` turns search results into a watchable clip. Your agent cites video the way a good analyst cites sources.

## Turning it into an agent tool

Wrap the flow as a `watch(url, question)` tool that returns timestamped moments plus an evidence clip. Or skip the wiring entirely: `npx skills add video-db/skills` gives Claude this capability as tools, and the same primitives exist as n8n and Zapier nodes.

Indexing is one-time per video, so every later question is just a search call. For batches, loop uploads into one collection and aggregate: a research agent that does in minutes what an analyst does in a week, with receipts.

## FAQ

**Can an AI agent summarize a YouTube video through an API?** Yes. Ingest the URL, run speech + visual analyzers, index, then search or summarize grounded in the indexed content, with timestamps and compiled clips.

**Does this only work for YouTube?** No. The same path takes direct file URLs and local files, plus live RTSP streams and screen recordings. (Process content you have the rights to work with.)

**What does it cost to index a video?** Indexing is one-time per video, and searches are cheap API calls. Free tier at console.videodb.io. Rates at https://videodb.io/pricing

**Is this just transcript search?** No. The visual channel is analyzed and indexed too. For transcript-only budget workloads, `video.index_spoken_words()` + `video.search()` is a leaner path.

---

Full guide: https://videodb.io/blogs/give-your-ai-agents-eyes · Quickstart: https://docs.videodb.io/pages/getting-started/quickstart

---

<!-- source: /blogs/rtsp-ai-analysis.md -->
# RTSP + AI: Turn Any Camera Stream into Agent-Readable Events

> Turn any RTSP camera stream into AI-readable events: continuous understanding, plain-language alerts, and searchable history. No CV pipeline to build.

Category: Tutorials
Published: 2026-08-09

---

RTSP stream AI analysis means connecting a live camera feed to infrastructure that continuously understands it. You get rolling-window analysis of what's happening, a live index you can search, and events you define in plain language that push alerts the moment they occur. No frame-grabbing loops, no training a classifier per event, no pipeline babysitting.

## The DIY wall

The do-it-yourself version is OpenCV frame grabs + YOLO or a VLM call per frame + your own event logic. Three problems never go away. Frames aren't events, because events live in time windows. Streams misbehave, and keeping a pipeline alive 24/7 is an SRE job. And there's no history unless you also build storage and indexing.

## Connect a stream, define an event, get alerts

```python
import asyncio
import videodb

RTSP_URL = "rtsp://samples.rts.videodb.io:8554/intruder"  # public test stream

async def main():
    conn = videodb.connect()
    coll = conn.get_collection()

    # 1. Ingest the live stream (store=True keeps searchable history)
    rtstream = coll.connect_rtstream(url=RTSP_URL, name="Dock Cam", store=True)

    # 2. Continuous understanding in rolling windows
    understanding = rtstream.understand(
        segmentation={"type": "time", "window": "5s"},
        analyzers=[{
            "type": "vlm",
            "name": "scene",
            "sampling": {"frame_count": 2},
            "config": {"prompt": "Describe the people, vehicles, movement, and activity."}
        }],
        store=True
    )

    # 3. A live index - queryable while the stream runs
    index = rtstream.index(name="dock-scenes",
                           source=understanding.outputs["scene"],
                           use_for=["semantic"])

    # 4. Define the event in plain language
    event_id = conn.create_event(
        event_prompt="Detect when a person appears in the monitored area.",
        label="person_appears"
    )

    # 5. Get pushed alerts (webhook and/or websocket)
    websocket = conn.connect_websocket(coll.id)
    await websocket.connect()
    alert_id = index.create_alert(event_id=event_id,
                                  callback_url="https://your-app.example.com/hooks/camera",
                                  ws_connection_id=websocket.connection_id)

    async for message in websocket.receive():
        if message.get("channel") == "alert":
            alert = message["data"]
            print(alert["label"], alert["confidence"])
            print(alert["explanation"])
            print(alert["stream_url"])
            break

asyncio.run(main())
```

The event definition is a sentence, not a trained model. "Detect water pooling on the floor." "Detect a vehicle blocking the fire lane." Changing what you monitor for is editing a prompt.

Every alert arrives with an explanation and a playable clip of the exact moment. With `store=True`, questions work backwards too: search "delivery truck at the dock" against last week.

## Production notes

Stop resources deliberately (`index.stop()`, `understanding.stop()`, `rtstream.stop()`). Rolling windows and frame sampling are your cost dial. Fleets scale to thousands of concurrent feeds. Sensitive footage runs in your own cloud (AWS/GCP/Azure), VPC, or edge, with Zero Data Retention options and SOC 2 Type II.

## FAQ

**What is RTSP, and will my camera work?** RTSP is the standard streaming protocol spoken by virtually every IP camera, NVR, and drone gateway. If your device exposes an RTSP URL, it connects. No vendor SDK needed.

**Do I need to train a model?** No. Analyzers are prompt-driven VLMs. Events are natural language. Custom CV models are supported when needed.

**How fast are alerts?** Understanding runs in rolling windows you configure (e.g. 5s). Alerts push over WebSocket/webhook as events are detected, so awareness is seconds-scale.

**Can I search past footage?** Yes. With storage enabled, stream history is indexed and searchable, and results return playable clips.

---

Runnable demos (intrusion, flood, baby monitor): https://github.com/video-db/videodb-cookbook · Fleet scale: https://videodb.io/live-camera-intelligence

---

<!-- source: /blogs/video-rag.md -->
# Video RAG: The Definitive Guide (2026)

> What video RAG is, why it breaks text-RAG assumptions, the four architectures that work, and how to build one, including RAG over live streams.

Category: Engineering
Published: 2026-08-09

---

Video RAG (retrieval-augmented generation for video) lets an LLM answer questions using video as its knowledge source. The video is analyzed and indexed ahead of time, relevant moments are retrieved at question time, and the model generates an answer grounded in those moments. Ideally that answer carries timestamps and playable clips as citations.

Text RAG retrieves paragraphs. Video RAG retrieves *time ranges*. That one difference drives every architectural decision.

## Why video RAG is not text RAG with extra steps

1. **The retrieval unit is a time range, not a chunk.** Chunking becomes segmentation: by time window, scene, or speech turns.
2. **Meaning lives in two channels that must stay aligned.** Transcript-only is blind. Frames-only is deaf. Both need agreeing timestamps.
3. **Cost is asymmetric.** VLM-analyzing an hour of video is expensive. Decide what understanding you pay for at index time so retrieval stays cheap forever.
4. **Citations should be playable.** The gold standard is a timestamped, watchable clip of the exact evidence.

## The four architectures

1. **Transcript-only RAG.** Cheap, and right for talking heads. Fails when meaning is on screen.
2. **Frame sampling + captions.** Better, but it fails on events that unfold over time, and alignment glue grows with scale.
3. **Native multimodal embeddings** (e.g. Twelve Labs Marengo). Strong for produced-media search. Embeddings aren't prompt-steerable, and you still build the surrounding pipeline.
4. **Hybrid indexes over both channels** (our recommendation for agents). Run speech analyzers AND prompt-steered VLM analyzers over time-segmented video. Index each channel, retrieve across both, and compile evidence clips. Index-time prompts act as an extraction schema.

## Reference implementation (architecture 4)

```python
import videodb

conn = videodb.connect()
coll = conn.get_collection()
video = coll.upload(url="https://www.youtube.com/watch?v=LPZh9BOjkQs")

understanding = video.understand(analyzers=[
    {"type": "spoken_words", "name": "transcript"},
    {"type": "vlm", "name": "scene",
     "config": {"prompt": "Describe the visual content: people, actions, on-screen text, charts."}}
])
understanding.wait_until_complete()

t_index = video.index(name="transcript", source=understanding.get_analyzer("transcript"))
s_index = video.index(name="scene", source=understanding.get_analyzer("scene"))
t_index.wait_until_complete()
s_index.wait_until_complete()

results = video.semantic_search(
    query="an explanation supported by diagrams and on-screen text",
    index_ids=[t_index.index_id, s_index.index_id],
    top_k=5
)
context = [(s.start, s.end) for s in results.get_shots()]  # feed your LLM
evidence = results.compile()                               # playable citation reel
```

Already on LlamaIndex? The VideoDB retriever (`llama-index-retrievers-videodb`) drops video in as another retriever: https://videodb.io/blogs/llama-index

Production notes. Keep `top_k` small, because video context is expensive to verify, so precision beats recall. Store shot timestamps in traces so users can dispute answers by watching evidence. Treat the index-time VLM prompt as versioned config.

## The frontier: RAG over live video

The harder case: the video never ends. Understanding must run continuously and the index must be queryable while it grows. Connect an RTSP stream, run rolling-window analyzers, and index live.

Live RAG also inverts question direction: pull (you ask) plus push (standing plain-language events that alert your agent the moment the answer becomes yes). Full loop: https://videodb.io/blogs/rtsp-ai-analysis

## What the research says

*Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension* (arXiv 2411.13093) and *VideoRAG: RAG over Video Corpus* (arXiv 2501.05874) converge on the same conclusion: retrieval over aligned multimodal signals beats both raw long-context stuffing and transcript-only pipelines on long-video tasks. That is the hybrid-index bet.

## Evaluating a video RAG system

Retrieval hit rate (gold set of question/timestamp pairs), temporal grounding accuracy (how tightly ranges bracket the true moment), answer faithfulness (does the answer match what the clip shows, which you audit by watching the evidence).

## FAQ

**What is video RAG?** RAG where the knowledge source is video: indexed ahead of time, time ranges retrieved per question, answers grounded with timestamps and playable clips.

**Video RAG vs multimodal RAG?** A subset. Video adds time as the retrieval dimension and two channels needing aligned timestamps.

**Do I need a vector database?** DIY architectures, yes. Managed hybrid indexes ship indexing, storage, and retrieval in the infrastructure.

**Can RAG work on live video?** Yes. Continuous rolling-window understanding feeds a live, growing index, plus standing events that push.

**Why not just long-context models?** Expensive and slow at library scale, impossible on never-ending streams, and no reusable index or playable citations. Index once, ask forever.

---

<!-- source: /blogs/twelve-labs-alternatives.md -->
# Twelve Labs Alternatives in 2026: An Honest Guide

> Comparing Twelve Labs alternatives for video AI in 2026 (VideoDB, Mixpeek, Google, AWS, Azure, NVIDIA), organized by what you're actually building.

Category: Product
Published: 2026-08-09

---

A disclosure before anything else: this guide is written by VideoDB, and we're on the list. We also integrate with Twelve Labs (https://videodb.io/blogs/twelvelabs), so this isn't a takedown. It's the guide we wish existed. The pages ranking for "Twelve Labs alternatives" today are auto-generated aggregator lists, and one even confuses Twelve Labs with ElevenLabs, the voice company.

**The short answer:** the right alternative depends on the job. Real-time streams and agent-native video infrastructure: VideoDB. Self-hosted multimodal indexing: Mixpeek. Hyperscaler-native label extraction: Google Video Intelligence, AWS Rekognition, or Azure Video Indexer. Self-managed GPU stack: NVIDIA's video analytics blueprints. Video-native embeddings behind a single model API: Twelve Labs itself.

## What Twelve Labs is

Video understanding foundation models exposed through an API: Marengo (embeddings/search) and Pegasus (video-to-text). Indexing and search run as a managed path. How video is segmented, embedded, and ranked is internal to the model and fixed at training time, and the API does not expose per-dataset prompts or artifact schemas.

Its documented customer base is concentrated in media archives, sports, and advertising. It ships integrations with the major clouds.

## Why people look for alternatives

They need live streams, not just uploaded files. They need infrastructure (storage, playback, clips, editing, eventing), not just inference. They're building agents and want agent skills and tool-call-native workflows and real-time alerts. They want self-hosting or their own cloud. Or procurement constrains them to a hyperscaler.

## The alternatives, by job to be done

- **VideoDB** (us). Complete video infrastructure with a multimodal agentic pipeline, covering the whole arc: **ingestion** (files, YouTube URLs, RTSP live streams, screen capture) → **understanding** (multimodal agentic pipeline with continuous rolling-window analysis and indexing over speech and visuals) → **retrieval** (semantic search with playable timestamped results, plus plain-language events pushed over WebSocket/webhook alerts) → **delivery** (a streaming engine that, based on that understanding, can modify the video itself by clipping, sequencing, stacking, overlaying, and streaming the result: https://videodb.io/blogs/infrastructure-that-sees-and-edits). Case in point: an agent edited VideoDB's launch video end-to-end, then re-ingested its own cut to verify it, at https://videodb.io/blogs/claude-edited-our-launch-video. Agent skills, managed or BYOC (AWS/GCP/Azure), ZDR, SOC 2 Type II. It's model-agnostic, so Twelve Labs models can run in the loop. Best for agents that need eyes, live camera intelligence, video RAG with evidence.
- **Mixpeek**. Self-hosted multimodal indexing, and you operate the stack.
- **Google Cloud Video Intelligence + Gemini**. Hyperscaler-scale labels plus prompt-based analysis. Agent workflows are your glue code.
- **AWS Rekognition (Video)**. AWS-native detection and streaming events.
- **Azure AI Video Indexer**. Enterprise media indexing in the Microsoft world.
- **NVIDIA video analytics blueprints (Metropolis/VSS)**. Reference architectures on your own GPUs, and you operate everything.
- **Cloudglue**. Lightweight video-to-LLM context extraction on files.

## The architectural difference, in one table

VideoDB's research report, "Search over the Visual World" (https://labs.videodb.io/papers/search-over-the-visual-world.pdf), documents the contrast between the two architectures on this list: a video-native model API and visual data infrastructure (Table 4 in the paper). Read it for what it is: a comparison of documented interfaces, not an independent assessment of model quality. And the two architectures compose rather than exclude each other. Condensed:

| Dimension | Twelve Labs (video-native model API) | VideoDB (visual data infrastructure) |
|---|---|---|
| Primary artifact | A trained model behind an API | A logical data format (VDB) and a pipeline over many models |
| Understanding produced by | One proprietary model family (Marengo embeddings, Pegasus generation) | An open portfolio: ASR, detectors, OCR, VLMs of any tier, and domain models, including video-native models |
| Live streams | Managed search documented over uploaded, indexed assets. Real-time workflows documented through the VideoDB integration | First-class sources. Scenes append in real time. Evidence from moments ago |
| Evidence output | Timestamped segments with playback metadata for indexed sources. Cross-source composition left to the application | Query-resolved scenes. Programmable composition (sequence, stack, overlay) into one playable evidence stream, archived or live |
| Visibility of understanding | Embeddings and internal state opaque | Artifacts readable and exportable. Users index their own way |
| New access path over old media | Re-index through the model | Derive a new index from stored artifacts, with no re-analysis |
| Improvement path | Train and migrate to the next model version | Recompose: swap an analyzer or index. Provenance keeps old and new comparable |

Condensed from Table 4 of the report. Interfaces as documented July 2026.

## The measured comparison

The same report runs a complete-system retrieval comparison between the two systems on a shared corpus: 885 videos, 47.6 hours, and 9,834 natural-language queries drawn from MSVD, YouCook2, VATEX, and MSR-VTT.

VideoDB ran collection-level semantic retrieval followed by a fixed BGE Gemma reranker. Twelve Labs ran its managed Marengo 3.0 search path, one index per dataset. Recall@k is the percentage of queries for which a relevant item appears in the first k results, and the macro-average is the unweighted mean of the four datasets.

Disclosure: VideoDB ran this benchmark and published the paper. It is a self-reported comparison, not an independent evaluation. Both sources are public: the [paper](https://labs.videodb.io/papers/search-over-the-visual-world.pdf) and the [benchmark code](https://github.com/video-db/search-over-the-visual-world).

| Dataset | System | R@1 | R@3 | R@10 | R@50 |
|---|---|---|---|---|---|
| MSVD | VideoDB | 70.10 | 80.73 | 89.17 | 96.39 |
| MSVD | Twelve Labs | 67.86 | 78.22 | 89.48 | 97.10 |
| YouCook2 | VideoDB | 65.96 | 80.98 | 93.48 | 97.87 |
| YouCook2 | Twelve Labs | 47.07 | 65.56 | 84.18 | 96.41 |
| VATEX | VideoDB | 83.46 | 90.28 | 95.24 | 97.77 |
| VATEX | Twelve Labs | 85.43 | 92.40 | 97.30 | 99.43 |
| MSR-VTT | VideoDB | 72.82 | 81.55 | 86.89 | 92.23 |
| MSR-VTT | Twelve Labs | 62.62 | 72.33 | 85.44 | 92.72 |
| Macro-average | VideoDB | 73.09 | 83.39 | 91.20 | 96.07 |
| Macro-average | Twelve Labs | 65.75 | 77.13 | 89.10 | 96.42 |

Table 9 of the report. Percentages.

On the macro-average VideoDB is higher through the first ten results. That is 73.09 against 65.75 at R@1, 83.39 against 77.13 at R@3, and 91.20 against 89.10 at R@10. The margin narrows as the cutoff grows, and Twelve Labs is higher at R@50 (96.42 against 96.07).

The ordering is not uniform across datasets. YouCook2 is the widest separation, 18.89 points at R@1 in VideoDB's favor. Twelve Labs is higher at every reported cutoff on VATEX. MSVD and MSR-VTT split by cutoff: VideoDB higher early, Twelve Labs higher deeper in the result list.

The report's own reading is that this is not a universal ordering of the two systems. Retrieval quality depends on how the representation and the ranking path fit the query distribution. And YouCook2's margin travels with a prompt, schema, index, and reranker all shaped around procedural cooking language.

## Pick by need

| If you need... | Look at |
|---|---|
| Agents with real-time eyes (streams, screens, alerts) | VideoDB |
| Video-native embeddings over a produced-media library, managed end to end | Twelve Labs |
| Self-hosted control of the whole index | Mixpeek (or NVIDIA blueprints) |
| Label extraction inside your existing cloud contract | Google / AWS / Azure |
| Video RAG with timestamped, playable evidence | VideoDB |
| Media-archive search with enterprise integrations | Twelve Labs (or VideoDB x TwelveLabs together) |

## The combination nobody mentions

"VideoDB vs Twelve Labs" is sometimes the wrong frame. We run them together: Twelve Labs models for understanding inside VideoDB's ingest → index → retrieve → act loop.

From the infrastructure's point of view, a video-native model is simply one more analyzer in the portfolio. This isn't hypothetical, because both companies document the pairing. Already invested in Marengo embeddings but need live streams, events, or clips around them? That's the path: https://videodb.io/blogs/twelvelabs

## FAQ

**What is the best Twelve Labs alternative?** It depends on the job. VideoDB for real-time and agent-native infrastructure, Mixpeek for self-hosted indexing, Google/AWS/Azure for cloud-native labels, NVIDIA blueprints for self-managed GPU stacks.

**Is Twelve Labs the same as ElevenLabs?** No. ElevenLabs is voice AI, and Twelve Labs is video understanding.

**Can I use VideoDB and Twelve Labs together?** Yes. VideoDB is model-agnostic infrastructure, and Twelve Labs models can run as the understanding layer inside VideoDB pipelines.

---

<!-- source: /blogs/video-benchmark-ground-truth-ambiguity.md -->
# What video retrieval benchmarks get wrong about ground truth

> Manual review of MSRVTT, MSVD, VATEX, DiDeMo, and QVHighlights shows cases where the benchmark ground truth is too narrow, shifted, or ambiguous. Valid retrieved clips end up scored as misses.

Category: Engineering
Published: 2026-07-24

---

We use public video retrieval datasets to benchmark search quality. The setup looks simple: take a text query, retrieve video moments, and compare the result against the dataset's ground-truth clip.

The benchmark score depends on one assumption.

The dataset ground truth has to be right.

That assumption does not always hold. During manual review, we found problem cases across MSRVTT, MSVD, VATEX, DiDeMo, and QVHighlights. In those cases the dataset annotation looked wrong, too narrow, or too hard to defend as the only valid answer.

Some retrieved clips matched the query visually, but they sat outside the accepted ground-truth window. In a benchmark report, those results look like retrieval misses. When you watch the clips, some of them look like dataset misses.

For example:

- A query like `a car is in a wreck` can match several crash moments, not just one timestamp.
- A query like `a video game is played` is broad enough to describe many gameplay clips.
- A query like `A woman is doing a big hole in a pumpkin.` can match the carving process before the annotated window.
- A query like `Most beautiful resort i have ever seen` depends on subjective context, not a single visual fact.

This matters because a retrieval score can mix two things. One is how well the system finds the right moment. The other is how well the dataset represents the possible right moments.

The second part is easy to miss when we only look at aggregate metrics.

We saw five annotation patterns:

- Narrow windows that capture only part of an event.
- Shifted timestamps where matching evidence appears nearby.
- Missing alternate positives when more than one clip is valid.
- Broad captions that allow several plausible clips.
- Subjective captions that depend on context or judgment.

## Methodology

We reviewed the query, the ground-truth clip, and the top retrieved shots side by side.

A case was included only when the retrieved clips had visible evidence matching the query. The useful question in those cases was whether the dataset annotation was complete and correct.

We left out examples that looked better explained by search quality, indexing, clip localization, or retrieved-shot descriptions with limited visual detail.

This was a qualitative review. The goal was to understand where one accepted timestamp can hide reasonable visual alternatives.

## Annotation patterns

The examples below use the same vocabulary throughout the note.

| Dataset issue | What it looks like |
| --- | --- |
| Narrow GT window | The accepted clip captures only a small part of the event. |
| Shifted GT timestamp | Relevant evidence appears before or after the accepted window. |
| Missing alternate positives | Other clips satisfy the query, but the benchmark accepts one answer. |
| Broad query with one target | The query supports multiple valid clips, while scoring expects one. |
| Subjective caption | The query depends on judgment, such as beauty or context. |

## Clip evidence

Each example shows the accepted GT clip first, followed by the top 5 retrieved shots.

The clips make the dataset issue easier to inspect.

### MSRVTT: `a car is in a wreck`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Ff4867ee8-f7e0-4df0-9bf6-ddbc6f7144a2.m3u8) on `video9800`
- Many returned segments clearly show car crashes or wrecks, but they fall outside or partially outside the annotated GT window. The GT appears narrower than the set of semantically matching segments.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `0.0-5.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Ff65a45a4-e726-4177-ad83-38b22347f158.m3u8) | A red rally car speeds along a winding paved road, loses control, and slides off the road into a grassy embankment. |
| 2 | `20.0-25.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fb17c82e5-b386-4178-a140-3ccc63d3df6f.m3u8) | A yellow-and-blue rally car skids off a tree-lined road and flips multiple times. |
| 3 | `30.0-40.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fafc3d55e-216b-44d4-a729-5bbb98f53246.m3u8) | A rally car loses control, slides off the road, and kicks up tire smoke. |
| 4 | `45.0-80.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F8b9e857d-fcde-4675-b15b-fe137f1ddf24.m3u8) | A yellow rally car drifts through a turn, skids off course, and crashes into roadside protection. |
| 5 | `115.0-120.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F3c13732c-cad3-4153-a2eb-1bbfed242d89.m3u8) | A rally car overturns on a bend and lands upside down at the edge of the road. |

### MSRVTT: `a video game is played`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F19ff0ce9-50c9-4925-ae57-57e548579ebf.m3u8) on `video8027`
- Many returned scenes explicitly show video games being played across multiple timestamps. The query is generic, so one accepted GT window captures only part of the visual concept.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `525.225-530.23s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F7b4fb5e0-cf6b-4998-832c-6f3daf9b3bab.m3u8) | A game screen displays a round-end message and visible gameplay interface. |
| 2 | `540.24-545.245s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F94226199-b8c6-4de7-8714-62d5f932fc41.m3u8) | A game character fires a projectile across a dark arena with scores and time on screen. |
| 3 | `550.25-555.255s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Ffcaa97b2-deb4-44ce-9349-0a2a9f4f138c.m3u8) | A game character attacks enemies in a grid-patterned arena. |
| 4 | `560.26-565.265s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F31bada25-05d8-444a-8091-8e1f979aa232.m3u8) | Arcade-style gameplay continues with a player character moving and firing. |
| 5 | `595.295-600.3s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F911df91a-6c65-4f00-9977-e2fc74473d84.m3u8) | A player character navigates a game screen with enemies and interface elements. |

### MSVD: `a delicious japanese dish`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F920ed36e-8bd1-4cd4-87dd-6ea429f37dbf.m3u8) on `-wa0umYJVGg_100_115`
- The candidate moments show several different Japanese-food preparation scenes, but the benchmark accepts only a narrow annotated answer for a broad and subjective caption.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `13.012-17.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F4b0b48fd-a735-43f5-96c5-eb1d2d232915.m3u8) | Onigiri is arranged with bento sides. |
| 2 | `5.005-13.012s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Ff5ec290a-132c-4989-92e0-951207669cab.m3u8) | A battered piece of food, likely tempura, is turned with chopsticks while frying in bubbling oil. |
| 3 | `0.0-5.004s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F2dafe092-1234-49dc-9cd3-f9a1569f76dc.m3u8) | Hands shape white rice around cooked salmon to make a filled onigiri. |
| 4 | `30.0-40.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F4d674f80-ae39-4db8-ac19-cbc90d67aaf9.m3u8) | A cutlet is rolled and pressed into a sesame or crumb coating. |
| 5 | `45.0-50.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F60206ebd-5d80-4c96-bb26-36bff54eef2c.m3u8) | The cutlet is flipped and pressed through breadcrumbs. |

### VATEX: `A woman is doing a big hole in a pumpkin.`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fadd1fdb1-36a1-4e72-b02e-2e94a83198f1.m3u8) on `efvOYBo03XM_000351_000361`
- Multiple returned clips are clear semantic matches for making a big hole in a pumpkin, but they occur outside the annotated GT window. The action appears to span more than the accepted timestamp.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `60.06-65.065s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fcc85e938-4705-41d0-b6dd-4006ae10abf3.m3u8) | A young woman stands beside a large pumpkin and begins working on it. |
| 2 | `65.065-70.07s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F250d279e-a671-4051-b746-040a1ce0c043.m3u8) | A woman is focused on carving a large orange pumpkin at a table. |
| 3 | `75.075-100.1s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F9056db89-218d-466f-8f98-265f0da41eeb.m3u8) | The same pumpkin-carving activity continues in a longer window. |
| 4 | `100.1-115.115s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fda599c64-274e-4159-8f2b-995f72c6c43d.m3u8) | The person continues actively carving the pumpkin on the table. |
| 5 | `115.115-125.125s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F715af6de-d8e5-42d4-91b7-ab551d60522b.m3u8) | A carving tool is used to saw into the large pumpkin. |

### DiDeMo: `head passes in front of camera`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F5d86463d-4de0-4afb-9f4f-db39f4d10f4d.m3u8) on `53301297@N00_5826898997_0a951bea4f`
- Returned clips show heads or faces moving into the immediate foreground at other timestamps. The query is broad, and the accepted GT captures one instance of a repeated visual event.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `40.04-45.045s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F42b21a70-7598-451b-b269-4da01210c341.m3u8) | The camera moves suddenly from people in the street toward the foreground. |
| 2 | `54.054-57.057s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fecffe748-ffb4-4dd2-a86a-610656867ee3.m3u8) | A person in a Santa hat moves their face closer to the camera. |
| 3 | `0.0-5.005s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fe23bbeda-8a8b-4e77-a7f3-5b5f573a13b8.m3u8) | A man leans close until his face fills the frame. |
| 4 | `48.048-51.051s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F1bfa7c06-0220-4bde-9258-d1bd2b473678.m3u8) | A man approaches from a doorway and leans into the camera. |
| 5 | `18.018-21.021s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fc1f14ce6-d658-4cfd-b332-38d3e95987fb.m3u8) | A close face leaves the frame as the camera tilts upward. |

### DiDeMo: `castle comes into view`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F73361a6a-a487-430a-8631-12eb73c83585.m3u8) on `51167579@N06_6829893951_f10a25a3c8`
- Several returned windows show the same semantic target, a tower or castle-like structure entering view, but the benchmark accepts a narrow GT span.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `69.0-77.2s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F27cdff2a-7664-4783-8a57-55826d7b577d.m3u8) | A camera pan reveals a tall, dark church or tower structure. |
| 2 | `5.005-10.01s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fd2997dca-1e8e-4f60-845d-ef124f573ea2.m3u8) | The camera moves along a path toward distant stone architecture. |
| 3 | `12.012-20.02s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Ff7b1067e-767b-4cd9-8c4a-4b4f7062f759.m3u8) | A stone tower remains visible while the viewpoint moves along the path. |
| 4 | `39.039-42.042s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F68d6d9d4-60e9-4bd0-8921-e69a45afcf58.m3u8) | A slow zoom brings the stone church tower closer in a churchyard view. |
| 5 | `20.02-25.025s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F926b96ba-2f74-453b-9168-2a47dd9f08d9.m3u8) | The viewpoint reveals more of the stone tower behind trees. |

### DiDeMo: `light starts blue`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fe5356584-20ff-4c6c-98e1-9099f4f68c52.m3u8) on `26292851@N04_4223864218_8a531d1c08`
- The accepted GT is subtle, while top-ranked clips contain clearer blue-light events outside the GT. The query is visually under-specified for a single narrow target.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `30.03-35.035s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F033b929b-0e80-434c-adeb-2f41f7385430.m3u8) | Blue lights appear and move across the airport tarmac view. |
| 2 | `40.04-45.045s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F3c931453-11cc-4ef7-83f3-85cf1320d64d.m3u8) | A line of blue runway lights appears from an aircraft-window view. |
| 3 | `70.07-75.075s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fe6ed13f4-3a9a-4fbd-8660-a041c7df531c.m3u8) | Bright ground lights pass through the frame during aircraft movement. |
| 4 | `0.0-5.005s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F8c7a5c06-1c73-4d48-8833-88a6091cf131.m3u8) | A silhouetted rower moves through glowing blue cave water. |
| 5 | `5.005-15.015s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F1d5ad418-cd0e-4d49-b90b-22d8e5ac46a0.m3u8) | A performer stands on an outdoor stage under blue light and smoke. |

### QVHighlights: `A girl speaking from her car`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F18695f65-3d3b-48c9-968d-4e543274b85b.m3u8) on `Zhx9Ki9bUkE_360.0_510.0`
- The GT is only about two seconds, while the video contains many car-speaking moments that satisfy the query.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `45.045-60.06s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F8dfbf5db-df32-4e83-93a4-441e0dc9abf8.m3u8) | A young woman sits in the driver's seat of a car and speaks to camera. |
| 2 | `105.105-120.12s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F6744c302-7c04-43c6-80cd-3c6981884149.m3u8) | A young woman speaks inside a car at night. |
| 3 | `150.15-165.165s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F0c8067cb-5918-4d76-9028-3903457dff26.m3u8) | A blonde woman with glasses speaks from the driver's seat. |
| 4 | `450.15-465.165s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fdfc61426-442f-4eec-9482-b1b9846beacc.m3u8) | A conversational car scene shows a young woman and an older man. |
| 5 | `120.12-135.135s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fc027962c-181a-4f26-8d31-90239513bb6e.m3u8) | A young woman sits in a car backseat and speaks directly to the camera. |

### QVHighlights: `Most beautiful resort i have ever seen`

- GT: [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fc281dd24-503d-4d04-bd2e-c3bf30e3cc1f.m3u8) on `lyGaTk4MLVM_60.0_210.0`
- The caption is subjective, and many resort scenes can satisfy it outside the small annotated GT windows.

| Rank | Window | Clip | Why it matters |
| --- | --- | --- | --- |
| 1 | `255.0-270.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F463cba80-f7c0-48a9-a33e-85927aa21ded.m3u8) | A panoramic beachfront resort view transitions into a hotel-room view. |
| 2 | `285.0-315.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F1e807ccd-9902-4b77-bc4c-245c987720fa.m3u8) | A guided room tour shows a luxury room and its outdoor surroundings. |
| 3 | `750.0-765.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2F52e3a137-ea21-4d87-a69a-0094c2b9e35b.m3u8) | The camera moves through a lush tropical resort toward a spa entrance. |
| 4 | `795.0-810.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Faec7e70a-0654-4eab-a076-9b28d3fff4e1.m3u8) | A relaxing tropical resort or spa scene appears. |
| 5 | `840.0-855.0s` | [clip](https://player.videodb.io/watch?v=https%3A%2F%2Fd1zudc7ewmc6ey.cloudfront.net%2Fv1%2Fa1a1e458-3e40-4ffb-894f-174086a979c6.m3u8) | The segment shows a resort or spa vacation setting. |

## Findings

The strongest examples followed a few repeatable annotation patterns.

Many annotations were wrong because they were too narrow for repeated or extended actions. Car wrecks, gameplay, pumpkin carving, and car-speaking moments can occur across many windows in the same or similar videos.

A single GT span can be brittle for those captions.

Generic captions create multiple plausible answers. Queries like `a video game is played`, `head passes in front of camera`, and `a delicious japanese dish` describe broad visual concepts.

When the benchmark accepts one clip, semantically valid results can still land outside the target window. That points back to the dataset as well as retrieval.

Subjective or context-heavy captions are harder to anchor visually. `Most beautiful resort i have ever seen` depends on context and judgment.

A clip can match the text while missing the dataset's selected moment. For those queries, the label is not a complete representation of the caption.

## Discussion

These examples make the benchmarks more useful when read carefully.

For retrieval tasks, a caption can describe an event, a repeated action, a broad visual category, or a subjective impression.

When that caption is paired with one accepted timestamp, the benchmark becomes sensitive to annotation coverage. If the timestamp is wrong or incomplete, the score can hide a valid retrieval.

That is worth keeping in mind when reading scores. Some misses are retrieval misses. Some are places where the dataset is wrong, even when the dataset is considered a standard benchmark.

Before evaluating a system against a benchmark, the dataset itself needs review. That can be manual review, LLM-assisted review, or both.

In video retrieval, one ground-truth window is not always the only correct answer.

## Citation

```
Samuel Alexander, "What video retrieval benchmarks taught us about ground truth",
VideoDB Labs, July 2026.
```

```bibtex
@article{alexander2026whatVideoRetrieval,
  author = {Samuel Alexander},
  title = {What video retrieval benchmarks taught us about ground truth},
  journal = {VideoDB Labs},
  year = {2026},
  note = {https://videodb.io/blogs/video-benchmark-ground-truth-ambiguity},
}
```

*First published on [VideoDB Labs](https://labs.videodb.io/engineering/field-notes/video-benchmark-ground-truth-ambiguity), July 24, 2026. This is the canonical home.*

---

<!-- source: /blogs/jepa-from-language-models-to-world-models.md -->
# JEPA: From Language Models to World Models

> Why Yann LeCun's JEPA bet makes the hidden state the training target instead of the next token, and what that shift means for vision-language models, robots, and long-horizon planning.

Category: Engineering
Published: 2026-07-07

---

A language model can tell you what usually comes next. A world model should tell you what happens if you act. That is the core of Yann LeCun's JEPA bet.

## Why next-token prediction got us this far

Before getting into JEPA, it is worth being fair to LLMs. Next-token prediction is brutally scalable: every document, code file, transcript, and forum thread becomes labeled training data. It also works because language is already compressed human experience. Text contains physics, social behavior, software conventions, recipes, plans, contracts, emotions, and arguments.

That is why LLMs became useful general interfaces. The mismatch appears when we ask them to become agents. Token prediction can still produce useful hidden states, but the training signal only checks whether the next visible token was likely.

It does not directly check whether the model represented the physical state, the future consequences, or the variables needed for control. For an agent that has to perceive, predict, act, and correct itself over time, that gap can become the bottleneck.

## Token prediction can be locally right and globally weak

A language model can generate the right next token while still having a weak internal trajectory for what comes next. That sounds subtle, but it matters. A token is a local target. A world state is a stronger constraint.

Say the model produces the right word while its hidden state drifts into a bad part of representation space. The immediate output may look fine. The system is still poorly positioned for planning, consistency, or future action.

The Semantic Tube Prediction paper is useful here. It separates two things that are easy to collapse: producing the right next token, and maintaining a useful hidden-state trajectory. The paper argues that a model can land in the correct token region while drifting away from the representation path that would support later predictions [1]. One of its core assumptions is:

> "The trajectory of x≤t is locally linear almost everywhere." [1]

The narrow point is that the next-token loss can reward the right next word without guaranteeing that the hidden trajectory remains useful for later prediction.

*A model can land in the right token region while its hidden state drifts away from a coherent future. Adapted from Semantic Tube Prediction [1].*

This is the first intuition behind JEPA: the target should not only be the next visible thing. The target should shape the internal state that makes future prediction useful.

If the goal is a better autocomplete system, predicting tokens is natural. If the goal is an autonomous system that can handle changing environments, future consequences, and multi-step plans, that is not enough. The hidden state itself probably needs to become a first-class training target.

![A model landing in the right token region while its hidden-state trajectory drifts](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/right-token-wrong-trajectory.webp)

*A model can land in the right token region while its hidden state drifts away from a coherent future. Adapted from Semantic Tube Prediction [1]. ([animated version](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/right-token-wrong-trajectory.mp4))*

## JEPA changes the object of prediction

A JEPA-style model predicts the representation of another view, missing part, or future state. It does not need to reproduce every pixel or emit every word. It learns an encoder and a predictor so that one latent representation can predict another. A simple version looks like this:

```
z_c = E(x_context)
z_t = E_target(x_target)
L_JEPA = D(P(z_c), stopgrad(z_t)) + R(E)
```

Here, E encodes the context and E_target encodes the target view or target state. P predicts the target representation, D measures distance in latent space, and R(E) keeps the representation from collapsing. That anti-collapse term is not a detail. It is the difference between learning a useful latent space and mapping everything to the same vector.

A naming note before going further: I use JEPA for the broad joint-embedding predictive idea. LeJEPA is a specific JEPA variant from a separate paper [3]. Its main difference is the anti-collapse mechanism: it adds SIGReg, a regularizer that pushes embeddings toward an isotropic Gaussian.

I use LeJEPA only when discussing that specific regularization and the theory built around it. It is not a synonym for every JEPA model.

This is the core shift:

| Architecture | Main prediction target | What it encourages |
| --- | --- | --- |
| LLM | next token | linguistic continuation |
| VLM | next token conditioned on visual context and prompt | visual-language alignment through generated answers |
| JEPA | latent target representation | representations shaped for prediction |

The important part is not that JEPA uses embeddings. Everyone uses embeddings. The important part is that JEPA trains the embedding space to be predictive.

V-JEPA 2.1 defines JEPA as a framework that learns representations "by making predictions in a learned latent space, rather than directly in the observation (input) space" [4]. That sentence is the architectural pivot.

*JEPA trains a predictor to match target embeddings rather than reconstructing pixels. Adapted from V-JEPA 2.1 [4]. Robot-arm photo source: Shixart1985 / Wikimedia Commons, CC BY 2.0 [12].*

For LLMs and most VLMs, the training check is still token-level: did the model assign enough probability to the next answer token? JEPA changes the check. Given a context view, can the model predict the target embedding for another view, masked region, or future state?

That does not solve agency by itself. But it puts pressure on the representation to carry information that survives across views and time. For a robot or long-running visual agent, that can include object position, pose, contact, reachable surfaces, and whether the scene is moving closer to the goal.

![A JEPA training loop comparing context and target embeddings](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/jepa-training-loop.webp)

*JEPA trains a predictor to match target embeddings rather than reconstructing pixels. Adapted from V-JEPA 2.1 [4]. ([animated version](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/jepa-training-loop.mp4))*

## What changes for vision-language models

This is where the JEPA argument becomes more concrete. Most VLMs turn perception into text. The model sees an image or video, takes a question or instruction, and produces tokens. That is useful, but it puts language generation in the loop even when the system mainly needs an updated state.

A vision-language JEPA changes the default object being predicted. Instead of mapping every visual question into an autoregressive text sequence, it can predict the semantic target embedding directly. VL-JEPA makes this argument clearly. The paper says:

> "Instead of autoregressively generating tokens as in classical VLMs, VL-JEPA predicts continuous embeddings of the target texts." [8]

That matters because the world is underdetermined at the surface. Ask a model what happens if a switch is flipped down. "The lamp turns off," "the room goes dark," and "the light shuts off" can all be correct.

In token space, those answers are different strings. In semantic embedding space, they should be nearby states. VL-JEPA uses the same example: "the lamp is turned off" and "room will go dark" are separate in raw token space, but ideally close in embedding space [8].

The implication is bigger than parameter efficiency. A VLM can become a continuous perception system. It can watch, update a semantic state, compare that state to a query or goal, and decode text only when language is actually needed.

VL-JEPA reports roughly 50% fewer trainable parameters in a controlled comparison. It also reports about 2.85x fewer decoding operations with selective decoding, while output quality stays similar [8]. The exact numbers will change with models and tasks. But the direction is important: language becomes an interface to a visual state, not the only form the state can take.

That changes what VLMs are good for. Instead of treating every frame as more context for a token generator, a JEPA-style VLM can maintain a compact, queryable representation of what is happening.

For video systems, that is a major shift. A model watching a long stream should not need to narrate everything. It should keep track of the meaningful state changes and speak when the task requires it.

## A world model is more than an embedding

Calling JEPA a world-model approach raises the bar. In his 2022 position paper, LeCun framed common sense as "a collection of models of the world" that tell an agent what is likely, plausible, and impossible [11].

That standard shifts attention from whether a model has embeddings to what those embeddings preserve. A representation can be useful for retrieval or classification and still drop the state variables needed for control. For agents, the latent space has to preserve the variables that matter for prediction and action.

A useful world model needs at least four things: a compact state representation, a way to predict future states, a way to condition those predictions on actions, and a planner that can search over possible futures. The strongest theoretical idea here is linear identifiability. If z is the true latent state of the world, a good learned representation h(z) should recover it up to a simple transformation:

```
h(z) = Qz
```

Here Q can be a rotation or another simple linear transform. The exact coordinates may change, but the geometry should remain usable. This matters because planning is geometry.

Suppose the real world has a smooth path from state A to state B. If the learned latent space twists that path into something warped, a planner working in latent space will make confident mistakes. It may optimize the wrong distance. It may choose a straight line in embedding space that becomes a bad trajectory in the real world.

A theory paper on LeJEPA makes this precise: under Gaussian latent variables, OU-style transitions, alignment, and Gaussian regularization, the learned representation can recover the true latent state up to rotation [2].

That is a narrow theorem about a specific setup, not a universal claim about all JEPAs. But it gives the right test: a world model is not an embedding that "looks semantic." It is an embedding whose geometry preserves the world well enough for prediction and planning.

This is the cleanest way to state the embedding-space objection: JEPA's biggest advantage is also its biggest risk, because everything important happens in latent space. If the latent space is faithful, JEPA gives agents a compact substrate for prediction. If it is collapsed, over-compressed, or geometrically wrong, the system can fail while looking mathematically elegant.

## The anti-collapse problem is the whole game

A naive latent-prediction model has an easy way to win: make every embedding the same. If E(x) = c for every input, then predicting the target embedding is trivial. The loss can look good while the representation is useless.

This is why JEPA methods care so much about stop-gradients, target encoders, exponential moving averages, whitening, variance constraints, contrastive losses, or explicit distributional regularizers. They are what prevent the model from cheating.

The LeJEPA paper's answer is SIGReg: push the distribution of embeddings toward an isotropic Gaussian [3]. The intuition is that a good latent space should be spread out, statistically well-conditioned, and difficult to collapse. The paper emphasizes that LeJEPA combines prediction with a regularizer that makes the embedding distribution behave well, reducing reliance on a bag of heuristics [3].

This also explains why "everything happens in embedding space" should not be dismissed as hand-waving. Embedding space is not magic. It has to be engineered, constrained, and tested against whether it preserves the right variables.

JEPA is powerful only if the learned space is shaped by objectives that make prediction, planning, and control possible. A bad latent space is worse than a bad image. At least a bad image can be inspected. A bad latent space may fail silently.

## Video makes the argument obvious

Language hides this problem because text is already an abstraction. Video makes it obvious. A video is not just a list of frames. It is a stream of state changes: objects move, hands interact with objects, cameras shift, and actions create consequences.

A pixel generator can learn to produce plausible future frames. That is useful, but it forces the model to spend capacity on surface detail.

LeCun's dashcam example is perfect. A generative video model may waste resources predicting the random motion of leaves beside the road. Those leaves occupy many pixels, even though they are mostly irrelevant to driving [9]. A JEPA-style video model asks a different question: how should the representation move?

```
ẑ_{t+1} = z_t + Δz_t
```

That is the right kind of abstraction for agents. The model does not need to render every texture to understand that an object moved left, a hand approached a cup, or a door is now more open.

V-JEPA 2.1 pushes in this direction by predicting masked or future visual representations, making features spatially dense, and preserving temporal consistency [4]. The paper's target is not just semantic recognition. It tries to make the latent state "spatially structured, semantically coherent, and temporally consistent" [4].

That phrase matters. A robot does not only need to classify a scene. It needs to know where things are, how they move, which surfaces matter, and which changes persist across time.

If a model sees a hand move toward a cup, the most important prediction is not the exact next pixel color of the hand. It is the evolving state: distance to cup, grasp possibility, object pose, likely contact, future occlusion, and maybe the intention implied by the motion. Those are latent variables, and this is where JEPA starts to look like a bridge from perception to action.

## What changes for vision-language-action models

VLMs connect perception to language. VLAs connect perception, language, and action. A VLA can read an instruction, look at a workspace, and output motor commands or action tokens. That is already a large step beyond captioning.

The question is whether the action system has a reusable model of consequences. It may instead be learning a direct mapping from observation and instruction to action.

JEPA points to a different middle layer for VLAs. The stack becomes: perceive the scene, encode the current latent state, predict how candidate actions change that state, then choose actions that move the world toward a goal. Language still matters because it specifies goals, constraints, and explanations. But the control loop needs a state space where consequences can be predicted before the robot acts.

This is the difference between a VLA that reacts and a VLA that can plan. If the instruction is "put the cup in the drawer," a direct VLA policy may learn useful behaviors from demonstrations.

A JEPA-style world model should also represent intermediate states: the cup is visible, the gripper is aligned, the cup is graspable, the drawer is open, the cup is above the drawer, the cup is released. Those are not just words. They are latent states the system should recognize, predict, and reach.

The implication for multimodal AI is practical. Bigger context windows help a VLM remember more frames. They do not by themselves give the model action-conditioned dynamics.

More demonstrations help a VLA imitate more behaviors. They do not by themselves give the system a compact space for counterfactual search. JEPA is interesting because it tries to make the hidden state itself predictive enough to support planning.

## Action is the line between representation and agency

A model that predicts what happens next from observation alone is still missing a key question for agency: what changes if the system chooses an action? For a world model, the core equation is:

```
z_t = E(o_t)
ẑ_{t+1} = F(z_t, a_t)
```

The model observes o_t, encodes it into latent state z_t, then predicts the next latent state after action a_t. That changes the question from "what comes next?" to "what happens if I do this?" This is the line LeCun keeps drawing around real agency: "I do not understand how you can even think of building an agentic system without the ability to predict the consequences of its actions" [10].

LeWorldModel is a concrete example of this direction. It learns from pixels, predicts next latent states conditioned on actions, and plans by rolling out candidate actions in latent space [5]. The paper states the JEPA instinct clearly:

> "Instead of attempting to model every aspect of the environment, JEPA focuses on capturing the most relevant features needed to predict future states." [5]

The planner can optimize toward a goal embedding:

```
a*_{1:H} = arg min_{a_{1:H}} D(ẑ_{t+H}, z_goal)
```

*[Video illustration of action-conditioned world model]*

This is still early. It is bounded, short-horizon, and evaluated in controlled settings. But architecturally it is the right shape.

The model does not need to generate a full video of the future. It needs to predict the state variables that make action selection possible.

This is why JEPA might replace language models as the core of some future AI systems. Language is still useful for instructions, tools, and explanation. But an agent that lives in the world needs a predictive state engine.

![The same current state leading to different predicted next scenes depending on the action](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/action-conditioned-world-model.webp)

*The same current state leads to different predicted next scenes depending on the action. Adapted from LeWorldModel [5]. ([animated version](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/action-conditioned-world-model.mp4))*

## Long-horizon autonomy is a decomposition problem

If a task requires many steps, a flat planner has two problems. Prediction error compounds with every rollout step, and the search space grows with the horizon. The Hierarchical World Models (HWM) paper [6] says single-level planning fails in two regimes: non-greedy tasks and long-horizon tasks where "prediction errors compound over autoregressive rollouts and the action search space grows exponentially with horizon."

Hierarchical latent world models attack this directly. The high-level model plans over coarse future states or macro-actions. Its first predicted future state becomes a subgoal. The low-level model then plans primitive actions to reach that subgoal. The system replans repeatedly as new observations arrive.

```
z_subgoal = F_high(z_t, l_t)
a*_{1:h} = arg min_{a_{1:h}} D(F_low(z_t, a_{1:h}), z_subgoal)
```

The high-level model handles long-horizon direction while the low-level model handles short-horizon precision. LeCun's version is simple: low levels make "short-term prediction with a lot of details," while longer-term prediction has to throw away detail so it does not diverge from reality [10].

*[Video illustration of hierarchical latent planning]*

This is one of the strongest arguments for JEPA-style world models. In the HWM paper, hierarchy improves planning across latent world-model backbones and task suites. One reported robot result is especially sharp: under the evaluated setup, flat VJEPA2-AC gets 0% success on Franka pick-and-place, while the hierarchical version reaches 70% for the cup task [6].

The failure is not only representation quality. Manual subgoals can rescue flat planners, which means the bottleneck is often decomposition.

This is also where language-only planning starts to feel brittle. A chain-of-thought can describe substeps, and a task list can store a plan. But an embodied agent needs subgoals in the same space where it predicts consequences.

LeCun puts the point bluntly: "your cat can do hierarchical planning, and your cat does not know language" [10]. The point is not that cats are the benchmark, but that hierarchical planning does not have to be expressed first as language. A useful agent should be able to represent "the cup is in a graspable pose" not merely as a sentence, but as a latent state it can recognize, predict, and reach.

![High-level latent subgoals decomposing one long rollout into short plans](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/hierarchical-latent-planning.webp)

*Hierarchy turns one long brittle rollout into repeated short plans toward latent subgoals. Adapted from Hierarchical Planning with Latent World Models [6]. ([animated version](https://videodb.io/assets/blog-media/jepa-from-language-models-to-world-models/hierarchical-latent-planning.mp4))*

## The risk: elegant latent spaces can lie

This is the part that should make us cautious. JEPA only works if the latent space preserves the right structure. There are many ways to fail:

- Collapse: everything maps to the same point.
- Over-compression: useful details disappear.
- Wrong variables: the representation captures texture instead of state.
- Warped geometry: latent distances do not match controllable changes.
- Off-manifold planning: the planner searches states the model never learned.
- Weak action coverage: the model cannot predict actions outside its data.
- Missing hierarchy: short-horizon predictions do not compose into long-horizon behavior.

This is also where human-like intelligence comparisons should be handled carefully. Humans do not merely compress the world into minimal vectors. We preserve messy, useful detail. We keep context, remember exceptions, and carry "inefficient" structure because it helps us adapt.

The *From Tokens to Thoughts* paper makes a related point about LLM embeddings: they "broadly align with human category boundaries, yet fall short on fine-grained semantic distinctions" [7]. A JEPA-style system that compresses too aggressively may be efficient and still miss what matters.

The core question becomes: which constraints make embeddings preserve the world variables needed for action? That is where the field is heading. Better objectives, better anti-collapse methods, better temporal structure, action-conditioned prediction, hierarchy, memory, and meta-control.

## The likely future stack

I do not think the future is a giant JEPA that simply replaces every language model use case. A more plausible architecture is simpler: language remains the interaction layer, while a JEPA-style world model holds and predicts state.

At a high level:

- Encoders map observations and instructions into latent state.
- A JEPA-style world model predicts how that state changes.
- A planner searches over future states and selects a path.
- Decoders turn the result into human-readable language or action-readable commands.

The important point is where language sits in the system. Language remains how humans instruct, inspect, and coordinate with machines. It is also how the system explains itself back to us.

If the latent space becomes faithful enough, future agents may use language as the interface over a predictive world model. The loop is simple: encode observations into state, predict how the state changes, plan over that state, and decode the result into words or actions.

## References

1. *Semantic Tube Prediction: Beating LLM Data Efficiency with JEPA*. https://arxiv.org/pdf/2602.22617
2. *When Does LeJEPA Learn a World Model?* https://arxiv.org/pdf/2605.26379
3. *LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics*. https://arxiv.org/pdf/2511.08544
4. *V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning*. https://arxiv.org/pdf/2603.14482
5. *LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels*. https://arxiv.org/pdf/2603.19312
6. *Hierarchical Planning with Latent World Models*. https://arxiv.org/pdf/2604.03208
7. *From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning*. https://arxiv.org/pdf/2505.17117
8. *VL-JEPA: Joint Embedding Predictive Architecture for Vision-language*. https://arxiv.org/pdf/2512.10942
9. Welch Labs, JEPA interview / explainer with Yann LeCun, part 1. https://www.youtube.com/watch?v=kYkIdXwW2AE
10. Welch Labs, JEPA interview / explainer with Yann LeCun, part 2. https://youtu.be/v_jDvpEGTIg
11. Yann LeCun, *A Path Towards Autonomous Machine Intelligence*, 2022. https://openreview.net/forum?id=BZ5a1r-kVsf
12. Shixart1985, *Robotic arm...*, Wikimedia Commons. https://commons.wikimedia.org/wiki/File:Robotic_arm_at_work_lifting_a_box_during_a_technology_exhibition.jpg

*First published on [VideoDB Labs](https://labs.videodb.io/research/jepa-from-language-models-to-world-models), July 7, 2026. This is the canonical home.*

---

<!-- source: /blogs/how-to-evaluate-multimodal-vlms-for-your-video-use-case.md -->
# How to Evaluate Multimodal VLMs for Your Video Use Case

> A practical workflow for evaluating video VLM setups with VideoDB and Langfuse, from task definition and dataset design to tracing, scoring, and deployment decisions.

Category: Engineering
Published: 2026-05-15

---

This blog explains how we evaluate VLMs for real video use cases and how to build a repeatable workflow around VideoDB and Langfuse.

Everything discussed below is implemented in this open-source repo, which you can run on your own videos: https://github.com/video-db/benchmark-vlms

The goal is simple: do not evaluate only the model, evaluate the full setup. For video workflows, the output depends on the segmentation strategy, frame sampling, video resolution, prompts, model choice, reasoning budgets, latency requirements, and post-processing.

The goal of the evaluation is not to declare a winner in the abstract. The goal is to decide what setup is right for your task, on your videos, at the quality, latency, and cost you can support.

## Define the task before touching the stack

Start by writing down what the system is expected to do.

That sounds basic, but it shapes almost everything that follows. Retrieval, monitoring, summarization, moderation, metadata extraction, and Q&A are different tasks. They produce different outputs, tolerate different errors, and usually require different extraction and evaluation strategies.

At this stage, the goal is not to answer every possible question. The goal is to narrow the problem enough that the benchmark reflects the real use case.

A useful way to do that is to get clarity on a few broad dimensions:

- What is the system expected to produce? A ranked clip list, an alert, a summary, an answer, or structured metadata all need to be evaluated differently.
- What does success look like in practice? In some workflows, false positives are the main problem. In others, missing an event is worse. This is where you define what "good enough" actually means for the product.
- What kind of signal does the task depend on? Some tasks depend mostly on static visual frame. Others depend on motion, spoken content, scene changes, visible text, or a combination of these. That directly affects extraction strategy, frame count, and model choice.
- What constraints does the system need to operate under? Real-time systems, batch pipelines, low-cost pipelines, and quality-first pipelines all push the setup in different directions.

Once these questions are clear, the rest of the setup becomes easier to design and much easier to interpret.

They also tell you where to start. If the task depends on short-lived actions, you will usually test denser sampling or more frames. If the video is mostly static, lighter extraction and smaller models may be enough.

If latency or cost is the main constraint, the benchmark should include lighter configurations early. If quality matters most, start with a stronger baseline and optimize down later.

## Build the dataset around the production decision

The dataset is the centre of the eval.

If the dataset does not reflect production, the results will not help much. Public benchmarks are fine for sanity checks, but they do not answer the question most teams actually care about: will this work on our data?

That means your evaluation set should include:

- Normal cases
- Hard cases
- Near-miss negatives
- Boring stretches
- Failure modes you already know about

For example, surveillance data should include occlusion, low light, motion blur, empty scenes, and crowded scenes. Meeting data should include crosstalk, screen shares, poor audio, quiet speakers, and long static sections. Retrieval tasks should include semantically similar wrong answers, not just obvious misses.

Do not build the set around what is easiest to label. Build it around the product decision you need to make.

## Define what accuracy means for the task

For retrieval, the real question is usually whether the right moment appears in the results, how high it ranks, and whether similar-but-wrong clips stay out.

For alerting, the question is usually whether the alert stream is usable. A detector that catches everything but raises an alert constantly may still be the wrong system.

For summarization, the useful question is whether the summary is factually correct, covers the important events, and avoids inventing things.

For metadata extraction, it is often better to score field by field. If you need `location`, `action`, `visible_text`, and `object_count`, score those separately.

This is also where precision and recall become product choices instead of academic terms. Decide early whether missed events or false alarms are more expensive for the use case.

## Compare setups, not just model names

Once the task and dataset are defined, compare complete configurations.

For video use cases, the main knobs are usually:

- Segmentation strategy
- Frame count
- Interval length
- Prompt
- Model family
- Reasoning or thinking settings
- Resolution and preprocessing
- Downstream validation logic

VideoDB already exposes several of these directly. Its scene extraction method supports shot-based and time-based extraction. Its indexing method supports custom scene indexes.

That is why the right benchmark unit is not "model A vs model B." It is "configuration A vs configuration B."

## Build the evaluation stack

We use VideoDB and Langfuse to run and track the evaluation workflow.

- VideoDB for ingest, segmentation, frame extraction, indexing, playback evidence, and running VideoDB-hosted models on scenes
- Langfuse for traces, datasets, experiment runs, and later analysis

## Start with VideoDB

The first job is to turn raw video into something benchmarkable.

That means uploading the asset, extracting scenes, choosing frame sampling, and optionally creating a baseline scene index. VideoDB's quickstart and scene methods support all of that directly.

```python
import os
import videodb

conn = videodb.connect(api_key=os.environ["VIDEO_DB_API_KEY"])
coll = conn.get_collection()
video = coll.upload(url="https://example.com/sample-video.mp4")
```

If you want natural scene boundaries, use shot-based extraction. If you want fixed windows for benchmarking, use time-based extraction. VideoDB's docs show both patterns.

### Shot-based extraction

```python
from videodb import SceneExtractionType

scene_collection = video.extract_scenes(
    extraction_type=SceneExtractionType.shot_based,
    extraction_config={
        "threshold": 30,      # Sensitivity (lower = more sensitive)
        "frame_count": 10     # Frames per detected shot
    }
)

for scene in scene_collection.scenes:
    print(scene.id, scene.start, scene.end)
    for frame in scene.frames:
        print(frame.url)
```

### Time-based extraction

```python
from videodb import SceneExtractionType

scene_collection = video.extract_scenes(
    extraction_type=SceneExtractionType.time_based,
    extraction_config={
        "time": 5,
        "frame_count": 3
    }
)

for scene in scene_collection.scenes:
    print(scene.id, scene.start, scene.end)
    for frame in scene.frames:
        print(frame.url)
```

At this point, you have the video units the benchmark will run on: scene boundaries with sampled frames.

These sampled frames are what the VideoDB-hosted model uses when you call `describe` on a scene. After the model returns descriptions, labels, or other structured metadata, pass that output to VideoDB's `index_scenes()` method. It turns the output into searchable indexes.

With that, the media side of the workflow is set up. The next step is to run the VLM over these scenes.

## Make the first request

The easiest path is to call `describe` directly on a VideoDB scene. That keeps the benchmark easy to reason about. The extraction step is explicit, the input is inspectable, and every output ties back to the exact scene and sampled frames that produced it.

A minimal first request looks like this:

```python
description = scene.describe(
    model_name="google/gemma-4-31B-it",
    prompt="Describe the scene.",
)

print(description)
```

This keeps model execution inside VideoDB while still letting you control the model and prompt used for each scene.

## Trace and compare with Langfuse

Once the execution layer is in place, the next job is to make the runs inspectable and reproducible.

In this workflow, Langfuse is the observability layer. It helps us trace each evaluation item, attach metadata, compare outputs, and define metrics. It also preserves enough context to understand why a result was good or bad.

This matters because a benchmark is not only about producing a score. It is also about being able to answer questions like:

- What exact input produced this output?
- Which configuration generated the result?
- How was it scored?
- What changed between two runs?

A useful trace for offline evaluation usually includes:

- Video ID
- Scene start and end timestamps
- Frame URLs
- Extraction config
- Prompt
- Model name
- Output
- Scores

That way, every result stays tied back to the exact media evidence and configuration that produced it.

## Define the right metrics before you compare runs

Not every task should be judged the same way, and not every score should be reduced to one overall number. A good evaluation pipeline should score the task in a way that reflects the actual product decision.

The important thing is to define those metrics early and keep them stable while comparing runs.

A simple trace structure is:

- One root span per evaluation item
- One child span for the model call
- Final output and metadata on the root

A simple example looks like this:

```python
import os
from time import perf_counter
from dotenv import load_dotenv
from langfuse import get_client

load_dotenv()

langfuse = get_client()

MODEL_NAME = "google/gemma-4-31B-it"
PROMPT = (
    "Describe the scene, key actions, and any visible text. "
    "Keep the description grounded in what is visible in the sampled frames."
)

for scene in scene_collection.scenes:
    frame_urls = [frame.url for frame in scene.frames]

    with langfuse.start_as_current_observation(
        as_type="span",
        name="video-evaluation",
        input={
            "video_id": video.id,
            "scene_id": scene.id,
            "scene_start": scene.start,
            "scene_end": scene.end,
            "frame_urls": frame_urls,
        },
        metadata={
            "model_name": MODEL_NAME,
            "prompt": PROMPT,
        },
    ) as root:
        with langfuse.start_as_current_observation(
            as_type="generation",
            name="scene-describe",
            model=MODEL_NAME,
            input={
                "scene_id": scene.id,
                "frame_urls": frame_urls,
                "prompt": PROMPT,
            },
        ) as generation:
            start = perf_counter()

            output_text = scene.describe(
                model_name=MODEL_NAME,
                prompt=PROMPT,
            )

            latency_ms = round((perf_counter() - start) * 1000, 4)
            score = evaluate_scene_output(output_text)

            generation.update(
                output=output_text,
                metadata={
                    "latency_ms": latency_ms,
                    "score": score,
                    "model_name": MODEL_NAME,
                },
            )

        root.update(
            output={
                "result": output_text,
                "latency_ms": latency_ms,
                "score": score,
            }
        )

langfuse.flush()
```

At this point, the trace contains the full context for each evaluation item: input, output, latency, and score. That makes the evaluation observable end to end.

## Use the output to make a decision

The output of this workflow should not be "model X won."

It should help you answer practical questions:

- Which configuration becomes the default path?
- Which lighter setup is good enough for easier cases?
- Which stronger setup should be reserved for harder slices?
- Where does the current system still fail?
- Should the next change be in the prompt, the extraction strategy, the thresholds, or the model itself?

That is the real purpose of the benchmark. It is not to produce a leaderboard. It is to help you decide what to deploy and what to improve next.

## Let the evaluation compound over time

A good evaluation run should not disappear after you make the first decision.

Over time, the traces, scores, and reviewed outputs start to become a high-quality dataset of real examples from your own domain. That makes future evaluations easier and helps catch regressions earlier.

It also gives you a stronger base for prompt iteration and dataset expansion. If that becomes the right next step, you can even train and adapt a smaller model on your own custom data.

## Run it on your own data

We have open-sourced the pipeline behind this workflow. You can run the same process on your own videos, define your own metrics, swap in your own models, and compare configurations without rebuilding the stack. You can find the repo here: [benchmark-vlms](https://github.com/video-db/benchmark-vlms).

A good first run is usually small and deliberate:

- Pick one use case
- Build a representative evaluation set
- Define the metric that matters for that task
- Compare a few meaningful configurations
- Review
- Choose a default path and if required a fallback strategy

## Make the benchmark useful

If your goal is best possible quality, start with a stronger baseline and optimize down later. Compare model choice, frame count, extraction interval, and resolution first, since those usually have a bigger impact on output quality than smaller model-running tweaks.

If your goal is lower latency, look first at lighter models, shorter context, fewer frames, lower resolution where acceptable, and model side optimizations like batching and caching.

If your goal is lower cost, test the same task with smaller models, quantized models, fewer frames, longer sampling intervals, and caching.

A practical way to think about it is:

- If accuracy or quality is the problem, start with the parts of the system that affect how much signal the model actually sees. In video workflows, that usually means segmentation strategy, sampling density, number of frames per segment, and resolution. If the benchmark is missing short actions, quick scene changes, or small visual details, the fix may not be a different model. It may be denser sampling, more frames, or better scene boundaries. If those changes do not move the result enough, compare stronger models or more reasoning-heavy settings.
- If latency is the problem, reduce the amount of work each request has to do. That usually means sending fewer frames, shortening context, lowering resolution where acceptable, or moving to a smaller or faster model. It can also mean tightening the model-running setup so requests are handled more efficiently.
- If cost is the problem, look for the parts of the pipeline that are easiest to simplify without breaking quality. That can mean fewer frames, longer extraction intervals, lower resolution, smaller models, quantized models, or caching repeated prompt and context patterns. It can also mean adding a lighter default path for common or easy cases. Then reserve the more expensive configuration only for the slices that actually need it.

By the end of a run, you should know which configuration becomes the default path and which setup gives you the best quality. You should also know which lighter configuration is acceptable when latency or cost matters more, and where the current system still fails. That is the output that matters.

## Citation

Please cite this work as:

```
S Nagaonkar, "How to Evaluate Multimodal VLMs for Your Video Use Case",
VideoDB Labs, May 2026.
```

Or use the BibTeX citation:

```bibtex
@article{nagaonkar2026evaluateMultimodalVlms,
  author = {S Nagaonkar},
  title = {How to Evaluate Multimodal VLMs for Your Video Use Case},
  journal = {VideoDB Labs},
  year = {2026},
  note = {https://videodb.io/blogs/how-to-evaluate-multimodal-vlms-for-your-video-use-case},
}
```

*First published on [VideoDB Labs](https://labs.videodb.io/research/how-to-evaluate-multimodal-vlms-for-your-video-use-case), May 15, 2026. This is the canonical home.*

---

<!-- source: /blogs/claude-chessboard-spatial-reasoning.md -->
# Strong VLMs can still fail on downstream vision tasks

> A 72-position chessboard-to-FEN evaluation where GPT-5.4 scored 100% and Claude Opus 4.7 peaked at 84.72%. The gap comes down to precise spatial localization, not chess understanding.

Category: Engineering
Published: 2026-05-19

---

Chess Lens needed the current board position in **FEN**, a standard notation for one exact chess position. FEN records the board as rows of piece letters and empty-square counts. For example, a row like `PP3P2` means two white pawns, three empty squares, one white pawn, and two empty squares.

![Understanding is not equal to precision in chessboard-to-FEN extraction](https://videodb.io/assets/blog-media/claude-chessboard-spatial-reasoning/cover-image.webp)

## FEN and what we measured

For this evaluation, we only measured the piece-placement field of FEN. The model received a board image and had to return which piece was on which square. We compared that output against the expected FEN for the same point in the game.

That makes the evaluation strict in the way Chess Lens needs it to be strict. If a pawn is shifted from one file to the next, the position may still look plausible in prose, but the FEN is wrong.

![Visual explanation of how chessboard-to-FEN extraction works](https://videodb.io/assets/blog-media/claude-chessboard-spatial-reasoning/chessboard-to-fen-explanation.webp)

## Setup

We evaluated a chess game across 72 board positions. Each evaluated model had to determine board orientation, identify every visible piece, assign each piece to a square, and emit the piece-placement field of FEN.

The score is exact-match accuracy against the expected FEN sequence for the game.

## Token budget changed the result

The first run used a 1024-token output limit. That was too low for high-reasoning configurations. Some models ran out of tokens before producing the final structured answer.

| Run | Accuracy | Exact / Eval | Notes |
| --- | --- | --- | --- |
| GPT-5.4 low summary | 100.00% | 72 / 72 | best overall |
| GPT-5.4 default | 93.06% | 67 / 72 | no reasoning summary |
| Claude Opus 4.7 xhigh summary | 76.39% | 55 / 72 | 6 parse errors |
| Claude Opus 4.7 high summary | 69.44% | 50 / 72 | 4 parse errors |
| Claude Opus 4.7 default | 66.67% | 48 / 72 | 1 parse error |
| GPT-5.4 xhigh summary | 4.17% | 3 / 72 | 69 parse errors |
| Claude Opus 4.7 max summary | 2.78% | 2 / 72 | 70 parse errors |

The low scores for GPT xhigh and Claude Opus 4.7 max were mostly truncation failures, not vision failures. The models used too many tokens and often did not reach the final answer.

We increased the output limit to 4096 tokens and reran the evaluation.

| Run | Accuracy | Notes |
| --- | --- | --- |
| GPT-5.4 low summary | 100.00% | best overall |
| GPT-5.4 default | 97.22% | strong, no summaries |
| GPT-5.4 xhigh summary | 94.44% | strong but expensive, still some parse errors |
| Claude Opus 4.7 max summary | 84.72% | best Claude Opus 4.7 run |
| Claude Opus 4.7 high summary | 80.56% | no parse errors, still mapping mistakes |
| Claude Opus 4.7 xhigh summary | 79.17% | similar to high |
| Claude Opus 4.7 default | 63.89% | weaker without thinking |
| Claude Opus 4.7 medium summary | 58.33% | worst Claude Opus 4.7 config |

## Aggregate scores were not enough

Claude Opus 4.7 improved after increasing the token budget, especially at max thinking. But the accuracy still did not catch up to GPT.

That pushed us to inspect the intermediate reasoning summaries and failed outputs instead of relying only on aggregate accuracy.

GPT's reasoning summaries usually stayed procedural and row-specific:

> For row 8, I see: a8 has a black rook, f8 has a black rook, h8 has a black king...
> Moving to rank 2, I see the white pawns at a2, b2, c2, with gaps where pieces have moved...

Claude Opus 4.7's reasoning summaries were more often global descriptions of the position:

> I'm looking at a chess board layout with pieces positioned across the rows, showing what appears to be a mid-game or puzzle position...
> Black has pawns scattered across the board with the king on f6, white has a rook on d1, a king on e3, and a few pawns positioned strategically.

This helped explain the remaining errors. Claude Opus 4.7 usually understood the board as a chess position, but it was less reliable at preserving the exact square-by-square layout.

## The failure pattern

Manual inspection showed mostly local errors. Pieces shifted by one file, rows had one extra or missing empty square, and moved-pawn gaps were missed. Other outputs had the correct piece family on the wrong square, or the correct piece in the wrong color or case.

For example:

| Expected | Claude Opus 4.7 Output | What Went Wrong |
| --- | --- | --- |
| `p1b1pk2` | `p2b1pk1` | empty-square counts shifted |
| `PP3P2` | `PP4P1` | pawn moved one file over |
| `1BNP1N1P` | `1BNB1N1P` | wrong piece at a specific square |

That pattern matters because FEN compresses each row using piece letters and empty-square counts. For example, `PP3P2` means two pawns, three empty squares, one pawn, and two empty squares. If Claude Opus 4.7 outputs `PP4P1`, the row still looks plausible. But the pawn has shifted by one file, so the FEN is wrong.

Claude Opus 4.7 was not failing at general chess understanding. It was failing at precise spatial localization.

![A single square mismatch invalidates the FEN row](https://videodb.io/assets/blog-media/claude-chessboard-spatial-reasoning/fen-row-mismatch.webp)

## The documented limitation

This matches the Claude vision documentation:

> "Spatial reasoning: Claude's spatial reasoning abilities are limited. It may struggle with tasks requiring precise localization or layouts, like reading an analog clock face or describing exact positions of chess pieces."

Source: https://docs.anthropic.com/en/docs/build-with-claude/vision

That line maps closely to this experiment. FEN generation requires exact localization across 64 squares. If a model recognizes the board but shifts one piece by one file, the FEN is still wrong.

## Takeaway

Chess Lens exposed a narrow but important failure mode: understanding the board is not enough when the output format requires exact coordinates.

Claude Opus 4.7 often produced outputs that looked reasonable at the position level, but FEN is evaluated at the square level. One shifted piece changes the board state. One wrong empty-square count changes the row. The model can be directionally right and still fail the task.

The takeaway is that downstream tasks need their own evaluations. The best general-performing model is not automatically the best model for the specific task you are trying to solve.

*First published on [VideoDB Labs](https://labs.videodb.io/engineering/field-notes/claude-chessboard-spatial-reasoning), May 19, 2026. This is the canonical home.*

---

<!-- source: /blogs/twelvelabs.md -->
# VideoDB × TwelveLabs: Search Any Moment Across Your Video Library

> TwelveLabs' multimodal understanding now plugs straight into VideoDB, so developers can turn live streams into searchable, alertable, playable moments.

Category: Partnerships

---

Human monitoring does not scale. A person can watch one feed for a while, maybe a few feeds with enough coffee, but fatigue always wins.

Think of a baby starting to climb, a riverbed changing from quiet to dangerous, or a restricted area suddenly occupied. Those are exactly the moments that get missed when the system depends on someone staring at pixels all day.

VideoDB's real-time infrastructure turns live streams into structured, searchable, actionable video data. With TwelveLabs' Pegasus 1.2 model available directly inside the VideoDB indexing pipeline, developers can build live video understanding apps without stitching together separate storage, streaming, model, and alerting systems.

The loop is simple: connect a stream, describe what the model should look for, index visual batches with Pegasus, define the event that matters, and receive a playable clip when the event is detected.

> "The camera stops being a passive feed. It becomes an event source your product can act on."

## The real bottleneck in video AI

The hard part is rarely one model call. It is everything around the model: ingesting RTSP or camera feeds, sampling frames, storing generated understanding, evaluating events, and delivering alerts fast enough to matter.

- **API sprawl:** video ingest, model inference, storage, playback, and notifications often live in different services.
- **Scaling pressure:** live streams produce continuous data, and visual understanding workloads are expensive if every frame is treated the same way.
- **Latency gaps:** if indexing, alerting, and playback are not part of the same pipeline, real-time products become slow review tools.

VideoDB collapses that workflow into one video-native layer. RTStreams handle live ingest and playback. Visual indexes convert batches of frames into scene descriptions.

Events define what the system should detect. Alerts deliver the moment as data, including confidence, explanation, and a stream URL.

## Introducing TwelveLabs inside VideoDB

The TwelveLabs integration adds Pegasus 1.2 to VideoDB's live visual indexing path. A stream index can opt into TwelveLabs frame understanding by setting one parameter: `model_name="twelvelabs-pegasus-1.2"`.

That small line matters. You keep the VideoDB primitives around the model: RTStream ingest, visual indexing, event definitions, WebSocket or webhook delivery, and playable clip generation. Pegasus handles frame-level understanding inside the index.

## How the pipeline works

```python
import videodb

conn = videodb.connect()
coll = conn.get_collection()

flood_stream = coll.connect_rtstream(
    name="Arizona Flood Stream",
    url="rtsp://samples.rts.videodb.io:8554/floods",
    store=True,
)

flood_scene_index = flood_stream.index_visuals(
    batch_config={
        "type": "time",
        "value": 10,
        "frame_count": 6,
    },
    prompt=(
        "Monitor the dry riverbed and surrounding area. "
        "If moving water is detected across the land, identify it "
        "as a flash flood and describe the scene."
    ),
    name="Flash_Flood_Detection_Index",
    model_name="twelvelabs-pegasus-1.2",
)

print("Scene Index ID:", flood_scene_index.rtstream_index_id)
```

The batch configuration controls how often VideoDB samples the live stream and how many frames go into each visual understanding pass. The prompt tells Pegasus what to look for. The model name selects TwelveLabs.

## From understanding to action

Events are reusable detection rules. Once an event exists, it can be attached to a stream index and delivered over WebSocket for live app experiences or webhook for server-to-server automation.

```python
event_id = conn.create_event(
    event_prompt="Detect sudden flash floods or water surges.",
    label="flash_flood",
)

ws_wrapper = conn.connect_websocket()
ws = await ws_wrapper.connect()

alert_id = flood_scene_index.create_alert(
    event_id=event_id,
    callback_url="",
    ws_connection_id=ws.connection_id,
)

async for msg in ws.receive():
    if msg.get("channel") == "alert":
        data = msg.get("data", {})
        print("Event:", data.get("label"))
        print("Confidence:", data.get("confidence"))
        print("Clip:", data.get("stream_url"))
        print("Why:", data.get("explanation"))
```

The alert payload is not just a notification. It carries the event label, confidence, explanation, timestamp, and stream URL for the detected moment.

## Demo 1: flash flood detection

[Open the Flash Flood Detection notebook in Colab.](https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/integrations/twelvelabs/Flash_Flood_Detection_TwelveLabs.ipynb)

The flash flood notebook connects a live RTSP sample stream, creates a Pegasus-powered visual index, and defines events such as `flash_flood`, `heavy_rainfall`, and `human_rescue`.

## Demo 2: baby crib monitoring

[Open the Baby Crib Monitoring notebook in Colab.](https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/integrations/twelvelabs/Baby_Crib_Monitoring_TwelveLabs.ipynb)

The baby crib notebook uses the same integration pattern with a different prompt and event rule. The stream index describes activity inside the crib and pays attention to standing, climbing, or escape attempts.

```python
crib_scene_index = crib_stream.index_visuals(
    batch_config={
        "type": "time",
        "value": 10,
        "frame_count": 6,
    },
    prompt=(
        "Describe the activity of the baby inside the crib. "
        "Notice if the baby stands up, climbs the rail, "
        "or attempts to climb out."
    ),
    name="Baby_Crib_Index",
    model_name="twelvelabs-pegasus-1.2",
)
```

## Why this is bigger than alerts

Real-time video intelligence is a new interface for software. Instead of asking users to review footage after the fact, the system can surface moments as they happen. The TwelveLabs integration strengthens the frame-understanding layer. VideoDB makes it productizable: ingest, index, search, alert, and replay all in one loop.

---

## Build real-time video understanding. Powered by TwelveLabs Pegasus + VideoDB.

Start from a working notebook, then adapt the stream, prompt, and event rule to your own product.

CTAs: [Open flood notebook](https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/integrations/twelvelabs/Flash_Flood_Detection_TwelveLabs.ipynb) · [Open crib notebook](https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/integrations/twelvelabs/Baby_Crib_Monitoring_TwelveLabs.ipynb)

---

<!-- source: /blogs/llama-index.md -->
# VideoDB × LlamaIndex: Plug Video Into Your RAG Pipeline

> The official VideoDB connector for LlamaIndex lets you treat video as a first-class source in any retrieval-augmented pipeline.

Category: Partnerships

---

Most RAG systems are fluent in text. They parse docs, chunk PDFs, embed web pages, and retrieve the paragraph that grounds an answer. But video is where a lot of the real world lives: product demos, lectures, security footage, training calls, sports clips, meetings, tutorials, news, livestreams. If your RAG pipeline cannot retrieve video moments, it is blind to one of the highest-bandwidth knowledge sources you have.

The VideoDB retriever for LlamaIndex lets a LlamaIndex app retrieve from VideoDB collections and videos, return timestamped nodes, and synthesize text answers. It can then turn the same retrieved moments into playable clips.

> "Video RAG should not end at a generated sentence. It should take you to the exact moment the answer came from."

## Why video breaks normal RAG

A video is not a long document. It is a timeline of speech, visuals, motion, and context.

- **Speech matters:** a training video may explain the answer verbally while the slide on screen stays static.
- **Visuals matter:** a product demo, sports play, or surveillance moment may be obvious on screen but never mentioned out loud.
- **Time matters:** a retrieved answer is only useful if you can jump to the moment and verify it.

LlamaIndex orchestrates retrieval and response synthesis. VideoDB stores, indexes, searches, and streams the video evidence.

## The architecture

VideoDB makes video look like a retrieval source without flattening it into text-only context. Spoken words and visual scenes become searchable nodes, each with metadata such as `video_id`, `start`, and `end`. LlamaIndex uses those nodes as retrieved context. VideoDB uses the timestamps to generate clips.

## Start with the retriever

```bash
pip install videodb llama-index llama-index-retrievers-videodb
```

```python
import os
import videodb

os.environ["VIDEO_DB_API_KEY"] = "YOUR_VIDEO_DB_API_KEY"

conn = videodb.connect()
coll = conn.get_collection()

video = coll.upload(url="https://www.youtube.com/watch?v=LPZh9BOjkQs")
video.index_spoken_words()
```

Now the video can participate in a LlamaIndex retrieval flow through `VideoDBRetriever`.

```python
from llama_index.retrievers.videodb import VideoDBRetriever
from videodb import SearchType, IndexType

spoken_retriever = VideoDBRetriever(
    collection=coll.id,
    video=video.id,
    search_type=SearchType.semantic,
    index_type=IndexType.spoken_word,
    score_threshold=0.1,
)

nodes = spoken_retriever.retrieve("Where does the speaker explain transformers?")
```

## Turn retrieved nodes into an answer

```python
from llama_index.core import get_response_synthesizer

query = "Where does the speaker explain transformers?"

response_synthesizer = get_response_synthesizer()
response = response_synthesizer.synthesize(
    query,
    nodes=nodes,
)

print(response)
```

## Then generate the clip

Every retrieved node includes a start and end time. For a single video, pass those intervals into `video.generate_stream()` and VideoDB creates a playable stream from the relevant moments.

```python
from videodb import play_stream

intervals = [
    (node.node.metadata["start"], node.node.metadata["end"])
    for node in nodes
]

stream_url = video.generate_stream(timeline=intervals)
play_stream(stream_url)
```

## Bring in visual understanding

Speech alone is not enough for video intelligence. VideoDB scene indexing turns visual moments into searchable scene descriptions.

```python
scene_index_id = video.index_scenes(
    prompt=(
        "Describe each scene with objects, actions, text on screen, "
        "and any visual context needed for retrieval."
    )
)

scenes = video.get_scene_index(scene_index_id)
```

Retrieve from that scene index with the same retriever interface:

```python
scene_retriever = VideoDBRetriever(
    collection=coll.id,
    video=video.id,
    search_type=SearchType.semantic,
    index_type=IndexType.scene,
    scene_index_id=scene_index_id,
    score_threshold=0.1,
)

scene_nodes = scene_retriever.retrieve("Show the part with a matrix or formula on screen")
```

## Multimodal RAG in practice

A practical multimodal flow retrieves from both indexes, combines the nodes, and lets LlamaIndex synthesize over the union.

```python
query = "Explain the section where the speaker discusses attention and shows a matrix."

spoken_nodes = spoken_retriever.retrieve(query)
scene_nodes = scene_retriever.retrieve(query)

response = response_synthesizer.synthesize(
    query,
    nodes=spoken_nodes + scene_nodes,
)

print(response)
```

For custom pipelines, fetch transcript and scene records from VideoDB, convert them into LlamaIndex `TextNode` objects, and build a standard `VectorStoreIndex`.

## Collection-level retrieval

VideoDB is not limited to one file. The retriever can target a whole collection. When retrieval spans multiple videos, each node still carries the `video_id`, `start`, and `end` metadata. You can use VideoDB timelines to compile clips from multiple source videos into a single stream.

```python
from videodb.timeline import Timeline
from videodb.asset import VideoAsset

timeline = Timeline(conn)

for node_with_score in spoken_nodes + scene_nodes:
    node = node_with_score.node
    timeline.add_inline(
        VideoAsset(
            asset_id=node.metadata["video_id"],
            start=node.metadata["start"],
            end=node.metadata["end"],
        )
    )

stream_url = timeline.generate_stream()
play_stream(stream_url)
```

## What builders can ship

- Support answers with video proof.
- Training libraries that answer questions.
- Meeting and lecture memory.
- Visual search for agent workflows.

[Open the Simple Video RAG notebook in Colab.](https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/integrations/llama-index/simple_video_rag.ipynb)

---

## Ground your agents in video. Answers, timestamps, and clips in one RAG loop.

Add VideoDB retrieval to LlamaIndex and make video a first-class knowledge source.

CTAs: [Open the notebook](https://colab.research.google.com/github/video-db/videodb-cookbook/blob/main/integrations/llama-index/simple_video_rag.ipynb) · [Read LlamaIndex docs](https://developers.llamaindex.ai/python/framework/integrations/retrievers/videodb_retriever/)

---

<!-- source: /blogs/tinyfish.md -->
# TinyFish x VideoDB: The Internet, Finally Visible to Your Agents

> TinyFish opens the web for agents. VideoDB lets them see inside video. Together, they close the loop between access and comprehension.

Category: Partnerships

---

Half of the internet's traffic today comes from AI agents. And yet, your agent can access barely 10% of what's actually out there, and understand even less of what it does find.

That's not a small gap. That's most of the internet.

Say you want to build an agent that follows a football tournament for you. Not just scores, but the actual story of each match. Who pressed in the second half, which goalkeeper was having a nightmare, the goal that came out of nowhere in the 87th minute. You don't want a Wikipedia summary. You want to feel like you watched it.

So you send your agent out.

It goes looking for the match. Half the good sources are behind login screens. The highlight reels it finds are on YouTube, returned as URLs it cannot see inside. It comes back with a final scoreline and maybe a text recap pulled from a surface-level search.

Technically, it did its job. But you wanted the 89th minute. The chip over the keeper. The red card that changed everything. Instead you got a box score.

This is the exact problem. Two walls, back to back.

The first wall is access. Most of the internet is not publicly scrapable. It lives behind login screens, paywalls, dynamic pages that need a real browser to render, forms that need to be filled. Your agent knocks, and the door doesn't open.

The second wall is comprehension. Even when your agent gets through, video is a black box. A YouTube URL is not information.

The transcript might tell you a commentator shouted "what a goal." It cannot tell you that the striker received the ball with his back to goal, turned two defenders, and curled it into the top corner. That lives in the frame. And your agent cannot see frames.

To close that loop, we built a World Cup agent with TinyFish and VideoDB. TinyFish gets the agent to the right parts of the web. VideoDB lets it understand what is inside the video.

Ask for goals, cards, fouls, or penalties from a match, and the agent turns open-web access plus video comprehension into a reel you can actually watch.

**Live app:** https://worldcup-video-agent.vercel.app/
**Repo:** https://github.com/video-db/worldcup-video-agent

## TinyFish: Opening Every Door

TinyFish gives your agent the ability to navigate the web the way a human does. Not scraping surface HTML, but actually going in. Logging into sites, clicking through pages, filling forms, handling dynamic content, and returning structured data from wherever it lands.

You give it a URL and a goal in plain English. It figures out the rest.

```python
from tinyfish import TinyFish, CompleteEvent

tf = TinyFish()

GOAL = """
Go to YouTube and search for "Manchester United full match 2024 2025".
Find up to 3 complete match videos. For each collect the title, URL,
channel, duration, opponent, and competition.
Return only JSON.
"""

with tf.agent.stream(url="https://www.youtube.com/", goal=GOAL) as stream:
    for event in stream:
        if isinstance(event, CompleteEvent):
            result = event.result_json
```

Your agent is no longer knocking on doors. It is walking through them.

## VideoDB: Actually Seeing What's Inside

Now your agent has the match URLs. But a URL is still just a pointer to a file it cannot read.

VideoDB ingests that video and turns it into something an agent can actually work with. It indexes the video at the scene level, running multimodal AI over sampled frames and generating descriptions of what is visually happening, moment by moment. Not what the commentator said. What was in the frame.

You craft a prompt that tells it what to look for, and it comes back with timestamped, searchable, retrievable scenes.

```python
from videodb import SceneExtractionType

scene_index_id = match_video.index_scenes(
    prompt="""
    You are analyzing a football match. Describe what is happening concisely.
    Use the prefix CHANCE: when the ball is near the penalty box with an attacker contesting it.
    Use YELLOW CARD: or RED CARD: when a card is shown.
    Use PENALTY: when a penalty is being taken.
    """,
)
```

Once indexed, your agent can search the video like a database.

```python
highlights = match_video.search(
    query="CHANCE ball near penalty box shot attempt",
    search_type=SearchType.semantic,
    index_type=IndexType.scene,
    scene_index_id=scene_index_id,
)

highlight_reel = highlights.compile()
```

It gets back playable clips. Actual video. Not a transcript, not a summary. The moments themselves, stitched into a reel, ready to watch or pass downstream to another agent.

## The Full Loop in Practice

Put them together and the agent's football research goes from a box score to this. TinyFish finds the match on YouTube. VideoDB ingests it, scenes it, and pulls every goal attempt, every card, every penalty.

The agent compiles a highlight reel, attaches match context, and delivers it. You get what you actually asked for.

We built a live version of exactly this. You can run any query, something like "Show me every yellow card from Brazil vs Morocco." The agent will find the match footage with TinyFish, process the video with VideoDB, and extract all the key moments. It returns a compiled clip alongside a structured match summary.

![World Cup video agent query asking for yellow cards from Brazil vs Morocco](/assets/images/Blog-post-preview/tinyfish-yellow-card-query.png)

The live app lets you turn a one-off request into a scheduled briefing. Set it up once, and the agent keeps bringing the right moments back on schedule.

First, you connect the two APIs the agent needs: TinyFish for access, VideoDB for video understanding. Then you choose the inbox where the result should land, whether that is Telegram, Discord, or Slack.

![World Cup video agent setup steps for adding API keys and choosing a delivery channel](/assets/images/Blog-post-preview/tinyfish-schedule-delivery.png)

From there, the prompt becomes a delivery rule. You pick the match query, the time, and the timezone. Every day on that schedule, the agent runs the loop on its own: find the right footage, understand the frames, cut the requested moments, and send the finished reel to your inbox.

![World Cup video agent setup steps for scheduling a daily reel and receiving it in an inbox](/assets/images/Blog-post-preview/tinyfish-schedule-keys-channel.png)

That is a new way to consume the internet. The agent does the searching, watching, and clipping in the background, then delivers the finished clip directly to your inbox at the scheduled time.

## It Is Not Just Football

These two walls appear everywhere agents go.

Imagine building a real estate research agent to find the best properties in a city. It can try to pull listings, but most property platforms require a login to see full details. TinyFish gets it in. Then the agent finds the walkthrough videos.

The listing description says "spacious, sunlit living room with modern finishes." The transcript of the walkthrough video says "here's the beautiful living area." Neither is useful.

What is useful is what VideoDB sees: a narrow kitchen, windows facing a concrete wall, flooring that is vinyl pretending to be wood. The agent that can only read would believe the listing. The agent that can see would not.

The visual layer is often where the truth lives. Transcript gives you what people said about a thing. Scene indexing gives you the thing itself.

Agents now make up half the internet's traffic. It is about time they could actually read it.

TinyFish and VideoDB do not just patch two separate problems. They close a loop. Access without comprehension is still blindness. Comprehension without access is still a locked door. Together, they are what it actually takes to send an agent into the internet and get something real back.

---

<!-- source: /blogs/why-agents-are-blind.md -->
# Why AI Agents Are Blind Today

> The gap between human perception and agent perception, and why it matters

Category: Philosophy

---

Your agent can summarize a 50-page document in seconds. It can write code, answer questions, and reason through complex problems.

But show it a 30-minute meeting recording and ask "what did the client say about the budget?" It fails.

## The Text-First Assumption

Modern AI agents are built on a text-first assumption. LLMs process text. RAG retrieves text. Tool calls return text. The entire agent architecture assumes the world is made of strings.

But the world isn't text.

* Your customer calls are audio
* Your security feeds are video
* Your user sessions are screen recordings
* Your meetings are multimodal streams

When agents encounter these inputs, they either:

1. Ignore them entirely
2. Attempt expensive full-video transcoding that doesn't scale
3. Hallucinate answers without verifiable grounding

None of these work.

## The Cost of Blindness

Consider what agents miss when they can't perceive:

**In enterprise workflows:**
* Customer sentiment from call recordings
* Visual context from screen shares
* Non-verbal cues in video meetings
* Timeline of events in incident recordings

**In monitoring applications:**
* Real-time security events
* Manufacturing quality issues
* Traffic and safety violations
* Drone and sensor footage

**In desktop assistants:**
* What the user is looking at
* Context from system audio
* Visual state of applications
* Multi-app workflows

An agent that can't perceive is an agent that hallucinates. It fills gaps with plausible-sounding fiction because it has no grounding in observable reality.

## Human Perception vs Agent Perception

Humans perceive continuously. We see and hear in real-time. We remember whole experiences. That means not just facts, but temporal sequences with sensory context.

When you recall a meeting, you don't remember a JSON object. You remember the moment: the screen, the voice, the pause before someone made a point.

Agents today have no equivalent. They have:
* Text-based memory (vector stores of embeddings)
* Text-based retrieval (semantic search over documents)
* Text-based reasoning (LLM inference over strings)

What they lack is perception. Perception is the ability to continuously take in video and audio, extract meaning in real-time, and ground responses in observable evidence.

## The Perception Gap

| Capability                | Human | Today's Agent |
| :------------------------ | :---- | :------------ |
| Continuous perception     | Yes   | No            |
| Real-time video/audio     | Yes   | No            |
| Episodic memory           | Yes   | No            |
| Evidence-grounded answers | Yes   | Partial       |
| Multimodal context        | Yes   | Limited       |

This isn't a minor limitation. It's a fundamental architectural gap.

## What Perception Enables

When agents can perceive:

1. **Grounded answers:** Every response can link to a playable moment
2. **Real-time awareness:** React to events as they happen, not after the fact
3. **Episodic recall:** "Remember the part where..." becomes answerable
4. **Multimodal reasoning:** Combine what was said with what was shown
5. **Continuous context:** Maintain awareness across sessions

The future of agents isn't just better reasoning. It's perception, the ability to see, hear, and remember.

---

## The agents that win will be the ones that can perceive. Give your agents eyes and ears.

[Read: Perception Is the Missing Layer](/blogs/perception-is-the-missing-layer)

---

<!-- source: /blogs/perception-is-the-missing-layer.md -->
# Perception Is the Missing Layer

> LLMs have reasoning. RAG has retrieval. What's missing? The ability to perceive the world as it happens.

Category: Philosophy

---

LLMs gave us reasoning. Vector databases gave us retrieval. Tool calling gave us action.

But when you look at the modern agent stack, there's a glaring gap: perception.

## The Current Agent Stack

Here's what a typical agent architecture looks like:

```
User Input (text)
    ↓
LLM (reasoning)
    ↓
Tools (retrieval, actions)
    ↓
Output (text)
```

Every layer is text-centric. Even "multimodal" models that accept images treat them as one-shot inputs: a single frame, processed once, discarded.

There's no:
* Continuous media processing
* Real-time event detection
* Temporal understanding
* Persistent perceptual memory

## What Perception Actually Means

Perception isn't just "can process an image." Perception is:

1. **Continuous:** Always on, not one-shot
2. **Temporal:** Understands time, sequences, causality
3. **Multi-source:** Video, audio, screen, mic, sensors
4. **Searchable:** Can be queried after the fact
5. **Actionable:** Triggers responses in real-time

When you perceive a meeting, you're not taking a screenshot. You're maintaining awareness of a time-evolving stream of visual and audio information, extracting meaning, and building memory.

## The Perception Stack

```
Continuous Media (screen, mic, camera, RTSP, files)
    ↓
Perception Layer (VideoDB)
    ↓
    ├── Indexes (searchable understanding)
    ├── Events (real-time triggers)
    └── Memory (episodic recall)
    ↓
Agent (reasoning + action)
    ↓
Output (grounded in observable evidence)
```

The perception layer sits between raw media and agent logic. It converts streams into structured context.

## Three Input Modes

| Mode                | Source              | Example                           |
| :------------------ | :------------------ | :-------------------------------- |
| **Files**           | Uploaded recordings | Meeting archives, training videos |
| **Live Streams**    | RTSP, RTMP, cameras | Security feeds, drones, IoT       |
| **Desktop Capture** | Screen, mic, camera | User sessions, support calls      |

Same architecture, same APIs, same mental model. Your agent can perceive a recorded file or a live stream the same way.

## From Batch to Real-time

Traditional video AI is batch-oriented:

1. Upload file
2. Wait for processing
3. Get results

Perception is real-time:

1. Stream continuously
2. Receive structured events as they happen
3. Act immediately

```python
# Events arrive in real-time
{"channel": "transcript", "text": "Let's talk about the budget..."}
{"channel": "scene_index", "text": "User opened the pricing spreadsheet"}
{"channel": "alert", "label": "budget_mention", "confidence": 0.95}
```

Your agent receives context as the world unfolds, not after processing completes.

## Searchable Memory

Perception includes memory. Not just current awareness, but the ability to recall.

```python
# What happened in this meeting about pricing?
results = video.search("pricing discussion")

for shot in results.shots:
    print(f"{shot.start}s - {shot.end}s: {shot.text}")
    shot.play()  # Play the exact moment
```

Every search result links to playable evidence. Your agent doesn't just claim something happened. It can show you.

## Why This Matters Now

Three trends are converging:

1. **Agents are going mainstream:** Not research demos, but production systems
2. **Edge devices have cameras:** Every laptop, phone, robot, and IoT device
3. **Users expect awareness:** "Why doesn't my AI know what I'm looking at?"

The agents that win will be the ones that can perceive. Text-only agents will feel blind in comparison.

## The Promise

When perception becomes a first-class layer:

* **Desktop agents** understand what you're doing, not just what you type
* **Support agents** see the user's screen, not just their description
* **Monitoring agents** react to events as they happen, not hours later
* **Meeting agents** know what was said AND shown, with timestamps

The future of AI agents is perception-first.

---

## The future of AI agents is perception-first. Build agents that see, hear, and remember.

[Read: MP4 Is the Wrong Primitive](/blogs/mp4-is-wrong-primitive)

---

<!-- source: /blogs/mp4-is-wrong-primitive.md -->
# MP4 Is the Wrong Primitive for AI

> Video files were designed for playback. AI agents need something different.

Category: Philosophy

---

MP4 was designed in 1998. Its job is simple: pack frames and audio into a file that plays sequentially from start to finish.

That's perfect for Netflix. It's terrible for AI.

## What MP4 Gives You

An MP4 file is a container. Inside:

* Compressed video frames (H.264, H.265, etc.)
* Compressed audio tracks (AAC, MP3, etc.)
* Timing information for synchronization
* Metadata (duration, resolution, codec info)

To access any content, you:

1. Decode the video stream
2. Extract individual frames
3. Process each frame through your model
4. Repeat for every second of footage

This works for short clips. It falls apart at scale.

## The Problem with Frames

Say you have a 1-hour video at 30fps. That's 108,000 frames.

To answer "what happened at 23:45?", your options are:

1. Decode and process all 108,000 frames (expensive, slow)
2. Sample frames and hope you don't miss anything (lossy, unreliable)
3. Process in real-time as the video plays (1 hour to process 1 hour)

None of these let you instantly query the content.

Compare this to a database:

```sql
SELECT * FROM meetings WHERE topic = 'pricing' AND timestamp > '23:40'
```

Instant. Indexed. Queryable.

MP4 doesn't give you this. It gives you a blob.

## What AI Actually Needs

AI agents don't watch videos. They query them.

The questions agents ask:

* "What was said about the budget?"
* "Show me the moment the error appeared on screen"
* "When did the person enter the frame?"
* "What happened between 10:30 and 10:45?"

These are queries, not playback commands. They need:

| Capability               | MP4     | What AI Needs |
| :----------------------- | :------ | :------------ |
| Random access by content | No      | Yes           |
| Semantic search          | No      | Yes           |
| Timestamped results      | Limited | Precise       |
| Multi-index queries      | No      | Yes           |
| Instant answers          | No      | Yes           |

## The Transcoding Trap

The common workaround: transcode everything.

1. Extract all frames
2. Run each through a vision model
3. Store the descriptions in a vector database
4. Query the database

This "works" but:

* **Cost**: Processing every frame is expensive
* **Latency**: Hours of processing before you can query
* **Storage**: Frame embeddings multiply storage costs
* **Staleness**: Live content can't be pre-processed
* **Loss**: Descriptions lose visual fidelity

You're converting video into text, then querying text. The video itself becomes a liability. You keep it around for playback, but you can't actually use it.

## Indexes as the Right Primitive

What if the primitive wasn't a file, but an index?

An index is:

* **Prompt-defined**: You specify what to extract
* **Timestamped**: Every result maps to exact moments
* **Searchable**: Natural language queries, instant results
* **Composable**: Multiple indexes on the same media
* **Playable**: Results link back to verifiable video

```python
# Create an index with a prompt
index = video.index_scenes(prompt="Identify product demonstrations")

# Query it with natural language
results = video.search("demo of the new feature")

# Get timestamped, playable results
for shot in results.shots:
    print(f"{shot.start}s: {shot.text}")
    shot.play()  # Verify by watching
```

The video file still exists. But you don't interact with it directly. You interact with indexes, which are semantic layers that make the content queryable.

## Multiple Perspectives

The power of indexes: you can create multiple on the same video.

```python
# Same video, different questions
safety_index = video.index_scenes(prompt="Identify safety violations")
activity_index = video.index_scenes(prompt="Track person movements")
text_index = video.index_scenes(prompt="Extract on-screen text")
```

Each index is a different lens on the same content. Query them separately or together.

Try doing that with an MP4.

## Beyond Files

The same model works for live streams:

```python
rtstream.index_visuals(prompt="Describe what user is doing")
rtstream.start_transcript()
```

No files, no pre-processing, no waiting. Indexes build in real-time as media flows.

## The Shift

| Old Model             | New Model                |
| :-------------------- | :----------------------- |
| File is the primitive | Index is the primitive   |
| Process then query    | Query without processing |
| Static, batch         | Dynamic, real-time       |
| One representation    | Multiple perspectives    |
| Playback-oriented     | Query-oriented           |

MP4 isn't going away. But for AI, it's the wrong level of abstraction.

---

## It's time to move beyond the file. Indexes are the right primitive for AI.

[Read: Why Video Was Built for Playback, Not Perception](/blogs/playback-vs-perception)

---

<!-- source: /blogs/playback-vs-perception.md -->
# Why Video Was Built for Playback, Not Perception

> 70 years of video infrastructure for human eyes, and why AI needs something different

Category: Philosophy

---

YouTube, Netflix, Zoom, Twitch. The entire video industry was built for one thing: putting pixels on human eyeballs.

That's a 70-year-old assumption. And it's the reason AI agents can't use video natively.

## The Playback Paradigm

Video infrastructure was designed around a simple model:

```
Source → Encode → Distribute → Decode → Display
```

Every piece of the stack optimizes for this:

* **Codecs** minimize bandwidth for sequential playback
* **CDNs** cache content for low-latency delivery
* **Players** buffer and render frames at the right framerate
* **Protocols** (HLS, DASH) adapt quality to network conditions

The end goal: a human watches a video from start to finish.

## What Playback Gives You

When you press play on a YouTube video:

1. The CDN delivers compressed chunks
2. Your device decodes frames in real-time
3. Frames render at 24/30/60 fps
4. Audio syncs with video
5. You scrub the timeline to navigate

This works brilliantly for entertainment. But notice what it doesn't give you:

* No way to query content
* No structured access to "what happened"
* No timestamp-level retrieval
* No semantic understanding
* No event detection

The video just... plays.

## What Perception Needs

AI agents don't watch. They query.

```python
# Agent question: "What did they say about the timeline?"
results = video.search("timeline discussion")

# Agent needs: timestamped, verifiable answers
for shot in results.shots:
    evidence = f"{shot.start}s: {shot.text}"
    playable_url = shot.stream_url
```

Perception requires:

| Capability     | Playback Model       | Perception Model            |
| :------------- | :------------------- | :-------------------------- |
| Access pattern | Sequential           | Random                      |
| Query type     | "Play from 10:00"    | "Find mentions of X"        |
| Output         | Pixels on screen     | Structured data + evidence  |
| Latency        | Seconds to buffer    | Milliseconds to query       |
| Scale          | One viewer at a time | Thousands of queries/second |

## The YouTube Gap

You can't ask YouTube:

* "What videos in my library mention competitor pricing?"
* "Show me every safety incident from last month"
* "When did this person appear in any of our recordings?"

YouTube has the content. But it has no semantic layer, so there is no way to query what's inside.

You can search titles and descriptions. You can't search content.

## The Zoom Gap

You can't ask Zoom:

* "What was the action item from yesterday's call?"
* "Show me the moment the client expressed concern"
* "When was the slide about Q4 projections shown?"

Zoom has recordings. But they're files, opaque blobs waiting for someone to watch them.

## The Enterprise Gap

Enterprise video is even worse. Security footage, training recordings, customer calls, manufacturing feeds.

All captured. None queryable.

The common workflow:

1. Something happens
2. Someone requests a recording
3. A human watches it (at 1x speed)
4. They manually note timestamps
5. Days later, you have an answer

This doesn't scale. And it definitely doesn't work for AI.

## Perception-First Architecture

What if video infrastructure was built for perception?

```
Source → Ingest → Index → Query → Evidence
```

Every piece optimizes for understanding:

* **Ingest** normalizes media from any source
* **Indexing** extracts semantic meaning with prompts
* **Query** returns timestamped, relevant moments
* **Evidence** provides playable verification

The end goal: an agent queries content and gets grounded answers.

## From "Play" to "Answer"

| Playback             | Perception                   |
| :------------------- | :--------------------------- |
| "Play the recording" | "What happened at 2pm?"      |
| "Skip to 10:00"      | "Find the product demo"      |
| "Watch this video"   | "Search across all videos"   |
| "Download the file"  | "Give me the relevant clips" |

Perception turns video from a thing you watch into a thing you query.

## Real-time, Not Batch

The playback model assumes recordings. You capture, then watch.

Perception works in real-time:

```python
# Live stream
rtstream.index_visuals(prompt="Detect intruders")

# Real-time alerts
{"channel": "alert", "label": "intruder", "confidence": 0.94}
```

No recording. No waiting. Events detected as they happen.

## The Platform Shift

For 70 years, video infrastructure optimized for:

* High visual fidelity
* Low latency playback
* Global distribution
* Human consumption

The next era optimizes for:

* Semantic understanding
* Instant queryability
* Real-time event detection
* Machine consumption

Video infrastructure is being rebuilt. Not for playback, but for perception.

---

## Video infrastructure is being rebuilt. Not for playback, but for perception.

[Read: What Episodic Memory Means for AI Agents](/blogs/episodic-memory-for-agents)

---

<!-- source: /blogs/episodic-memory-for-agents.md -->
# What Episodic Memory Means for AI Agents

> Humans remember experiences, not just facts. Your agent should too.

Category: Philosophy

---

When you remember a meeting, you don't recall a JSON object. You remember the moment: the room, the voice, the pause before someone made a key point.

Humans have episodic memory. AI agents don't. That's about to change.

## Two Kinds of Memory

Cognitive science distinguishes between:

**Semantic Memory:** facts and concepts
* "The capital of France is Paris"
* "Water boils at 100°C"
* Timeless, context-free, declarative

**Episodic Memory:** experienced events
* "I remember the meeting where we discussed the budget"
* "That call where the client mentioned timeline concerns"
* Time-stamped, contextual, experiential

Most AI memory systems are semantic. Vector databases store embeddings of facts. RAG retrieves documents.

But agents that perceive need episodic memory. They need to remember what they saw and heard, when it happened, and what the context was.

## Why Episodic Matters

Consider these queries:

| Query                                                  | Memory Type | What's Needed               |
| :----------------------------------------------------- | :---------- | :-------------------------- |
| "What is our pricing model?"                           | Semantic    | Retrieved from docs         |
| "What did the client say about pricing last Tuesday?"  | Episodic    | Retrieved from recordings   |
| "How many people attended the meeting?"                | Episodic    | Visual memory of the event  |
| "What was on screen when they mentioned the deadline?" | Episodic    | Multimodal temporal context |

Semantic memory can't answer episodic questions. You need memory of experiences, not just facts.

## Video as Natural Episodic Memory

Video is inherently episodic:

* **Time-indexed:** every frame has a timestamp
* **Multi-sensory:** visual + audio together
* **Contextual:** shows the environment, not just content
* **Continuous:** captures the flow of events

When you record a meeting, you're creating episodic memory. The challenge is making it retrievable.

## The Memory Problem

Raw recordings aren't queryable. You can't ask an MP4 file "what happened?"

Traditional approaches:
1. **Full transcription:** converts audio to text, loses visual context
2. **Frame extraction:** expensive, loses temporal flow
3. **Manual notes:** doesn't scale, subjective
4. **Just store it:** the recording exists but no one can find anything

None of these create true episodic memory. They create archives.

## Indexed Episodic Memory

The solution: indexes that understand what happened and when.

```python
# Create episodic memory from a video
video.index_spoken_words()  # What was said
video.index_scenes(prompt="Describe activities and events")  # What happened

# Query episodic memory
results = video.search("budget discussion")

for shot in results.shots:
    print(f"At {shot.start}s: {shot.text}")
    shot.play()  # Relive the moment
```

The index is the memory. It captures:
* What happened (semantic content)
* When it happened (timestamps)
* Evidence (playable links)

## Ephemeral vs Persistent

Not all perception needs permanent memory.

**Ephemeral:** process but don't store
* Real-time event detection
* Privacy-sensitive contexts
* Temporary sessions

```python
rtstream.index_visuals(
    prompt="Detect safety issues",
    ephemeral=True  # Don't persist
)
```

**Persistent:** store for later recall
* Meeting recordings
* Training content
* Compliance archives

```python
video.index_spoken_words()  # Stored by default
```

You control what your agent remembers.

## Desktop as Continuous Input

Desktop capture creates continuous episodic input:

```python
cap = conn.create_capture_session(end_user_id="user_123")

# What the agent "experiences":
# - Screen content (visual)
# - Microphone (spoken)
# - System audio (ambient)
```

The agent perceives the user's experience in real-time. With indexing, it builds memory.

Later:
```python
# Agent recall
"Remember when I was debugging that error? What file was I looking at?"

results = cap.search("debugging error")
shot.play()  # Show the moment
```

## Multi-Session Memory

Episodic memory spans sessions:

```python
# Search across all recordings
results = coll.search("product roadmap discussions")

# Results from any video in the collection
for shot in results.shots:
    print(f"Video: {shot.video_id}, Time: {shot.start}s")
    print(f"Content: {shot.text}")
```

The agent doesn't just remember one meeting. It remembers all meetings.

## Grounded Answers

Episodic memory enables grounded responses:

**Without episodic memory:**
> "I believe the pricing discussion happened last week..."

**With episodic memory:**
> "At 14:32 in yesterday's meeting, Sarah said 'We need to revisit the enterprise tier pricing.' Here's the clip: [play]"

The difference is trust. Episodic memory provides verifiable evidence.

## The Future

The agents we're building will:
* Perceive continuously (screens, mics, cameras)
* Index what they perceive (spoken, visual, events)
* Remember across sessions (episodic recall)
* Answer with evidence (playable proof)

This isn't science fiction. The architecture exists today.

---

## Agents that remember experiences. Not just facts, but moments they can prove.

[Read: Infrastructure that "Sees" and "Edits"](/blogs/infrastructure-that-sees-and-edits)

---

<!-- source: /blogs/infrastructure-that-sees-and-edits.md -->
# Infrastructure that "Sees" and "Edits"

> Explore VideoDB's multimodal infrastructure that combines vision and editing capabilities for intelligent video processing.

Category: Philosophy

---

The arc of video creation is shifting. For decades, video editing has been a manual, linear process. A human interacts with a timeline, making micro-decisions on every frame. But as Large Language Models (LLMs) transform how we write code and text, a new frontier is opening: **Agentic Video Editing.**

At VideoDB, we have been architecting the ecosystem required to support this shift. We believe that for AI to truly edit video, it cannot simply be a "co-pilot" offering suggestions. It must be an **autonomous agent** capable of seeing content, understanding context, and executing complex modifications programmatically.

This is the VideoDB vision: converting opaque video files into fluid, intelligent data that agents can manipulate in real-time.

## The Semantic Layer: Giving Agents "Sight"

Before an agent can edit a video, it must understand it. A standard MP4 file is a black box to an LLM. It is a stream of binary data without meaning.

VideoDB solves this by providing the **semantic infrastructure** for video. Through our [Visual Search and Indexing capabilities](https://docs.videodb.io/pages/understand/indexing-pipelines/create-an-index), we index video content into queryable data. This allows an AI agent to "watch" a video and instantly locate specific moments, objects, or actions. That turns a visual search problem into a database query.

Once the video is indexed, the agent moves from perception to action.

## From Prompt to Timeline: The Editing AI

We have built the **AI Video Editing Automation SDK** to bridge the gap between intent and execution. This allows developers to build agents that function like a human editor's brain, capable of:

1. **Scene Understanding:** Analyzing the mood, lighting, and context of a shot.
2. **Object Segmentation:** Identifying specific elements (like a person or a prop).
3. **Intelligent Overlays:** Inserting assets dynamically based on spatial awareness.
4. **Audio Analysis:** Syncing visuals to beats or speech patterns.

An agent can take a high-level command such as *"Add a No Smoking image overlay wherever anyone is smoking"* and execute the entire pipeline autonomously. It finds the cigarette, understands the spatial coordinates, and inserts the asset on the correct track, all without human intervention.

[*Explore the Timeline Architecture*](https://docs.videodb.io/pages/act/programmable-editing/timeline-architecture)

## Infrastructure for GenAI and Real-Time Compositing

Agentic editing isn't just about cutting existing footage. It's about generating new realities. The VideoDB ecosystem supports the seamless assembly of GenAI video, music, and audio in real-time.

Traditional workflows require expensive rendering and "MP4 rebuilds" for every change. VideoDB changes the physics of this process. We treat video as a dynamic canvas. Whether you are generating background assets or injecting hyper-personalized content, our infrastructure handles the compositing on the fly.

This server-side composition capability enables use cases that were previously impossible:

* **Hyper-Personalized Ads:** Injecting user-specific products into a video stream instantly.
* **Live GenAI Assembly:** Stitching together generated clips and audio without rendering latency.

[*Learn more about our Generative Media capabilities*](https://docs.videodb.io/pages/act/generative-media/index)

## Meet "Director": The Open Source Agent

To accelerate the adoption of agentic workflows, we have open-sourced **Director**.

Director is a reference implementation of a video editing agent built on VideoDB. It demonstrates how to orchestrate the "See, Understand, Modify" loop. It also serves as a blueprint for teams building their own automated post-production pipelines. Fork it, modify it, and deploy agents that act as autonomous video producers.

[*Check out the Director Open Source Repository.*](https://github.com/video-db/Director)

## Enterprise Scale and Advanced Workflows

While individual agents transform creation, **enterprise orchestration** transforms business models.

For media companies and platforms, VideoDB scales these agentic capabilities to handle millions of streams. Our enterprise solutions focus on sophisticated workflows where content must be adapted, personalized, and monetized in real-time across global audiences. From automated compliance editing to dynamic ad insertion that feels native to the content, we provide the backbone for the future of media delivery.

## Building the Future

We are moving past the era of rigid video files and manual timelines. We are entering the era of programmable media. VideoDB is the infrastructure that empowers developers to build agents that don't just watch video. They understand it, create it, and master it.

---

## Build agents that see and edit. Programmable media is the future.

[Explore Director on GitHub](https://github.com/video-db/Director)

---

<!-- source: /trust-centre.md -->
# Trust Centre — VideoDB

Current as of 18 April 2026.

VideoDB is SOC 2 Type II certified, ISO 27001 certified, and compliant with HIPAA and GDPR. Security certificates are available for review upon request — fill in the form on [/trust-centre](/trust-centre) and we'll reach out by email to verify the request and share the relevant documents.

## At a glance

- SOC 2 Type II — independently audited controls.
- ISO 27001 — certified Information Security Management System (ISMS).
- HIPAA — safeguards for protected health information.
- GDPR — EU data protection compliance.

## Security certificates

The current set of reports and certificates we share with prospective customers and partners under NDA.

### SOC 2 Type II Report (2025)

Independently audited report confirming our controls for security, availability, and confidentiality meet AICPA Trust Service Criteria.

Issued Nov 2025 · Audited by Intercert.

### ISMS SA1 Certificate (E-Copy)

Formal ISO 27001 certificate issued by E-Copy confirming VideoDB's certified Information Security Management System.

Issued Nov 2025 · Certified by E-Copy.

### ISMS SA1 Client Report

Detailed audit report for VideoDB's ISMS covering scope, controls, and findings against ISO 27001 requirements.

Issued Nov 2025 · Audited by Intercert.

### HIPAA Audit Report

Audit report confirming VideoDB's compliance with HIPAA safeguards for the protection of health-related data and PHI.

Issued Nov 2025 · Audited by Intercert.

### GDPR Audit Report

Audit report verifying VideoDB's adherence to the EU General Data Protection Regulation for handling personal data of EU residents.

Issued Nov 2025 · Audited by Intercert.

## Request access

Submit the form at [/trust-centre](/trust-centre#request-access) with your name, work email, and the certificates you'd like to review. We verify each request by email — typically within one business day — and share documents under NDA via a secure link.

Your information is used only to process the certificate request. We do not share details with third parties.

## Related

- [Security policy](/legal/security)
- [Privacy](/legal/privacy)
- [DPA](/legal/dpa)
- Contact: security@videodb.io

---

<!-- source: /legal/terms.md -->
# Terms of Service — VideoDB

> Terms and conditions for Spext Labs Inc.'s VideoDB Product. Current as of 30 October, 2025.

---

## Introduction

By accessing or using VideoDB ("Service"), a product of Spext Labs Inc., the company operating under the brand name VideoDB, you agree to be bound by these Terms and Conditions.

## Description of the Service

VideoDB is a proprietary software product that lets users upload video files, process them, and make them searchable via text-based queries using machine learning and natural language processing.

## Compliance and Certifications

Designed to comply with HIPAA, GDPR, ISO 27001, and SOC 2.

## User Responsibilities

You are responsible for ensuring uploaded data complies with applicable laws. Prohibited: content that violates laws/regulations, or sensitive personal information you lack the right to upload.

## Data Privacy and Security

Video content is processed for transcription, indexing, and searchability. All data is encrypted in transit and at rest. Access is restricted to authorized personnel. EU and similar-jurisdiction residents have rights of access, rectification, deletion, and data portability.

## Confidentiality

Uploaded content is treated as confidential and proprietary and protected from unauthorized access.

## Liability and Indemnification

The Service is provided "as is" and "as available" with no warranty. You agree to indemnify Spext Labs Inc. against claims arising from your use of the Service.

## Amendments, Governing Law, Contact

Terms may be amended at any time. Governed by the laws of San Francisco, California, USA.

Contact: **Spext Labs Inc.**, contact@videodb.io, 45 Lansing Street, #2111, San Francisco 94105 USA.

---

<!-- source: /legal/privacy.md -->
# Privacy Policy — VideoDB

> How VideoDB collects, uses, stores, and protects Personal Data, and your rights as a data subject. Current as of 30 October, 2025.

---

## Data Controller

Spext Labs Inc., the company operating under the brand name VideoDB, is incorporated in Delaware, USA, with offices in California and India, and determines the purposes and means of processing personal information.

## Information We Collect

- **Personal Data you provide** — collected only when voluntarily provided (sign-up, support, certain services). Personal Data in uploaded audio files is used solely to deliver transcription and improve algorithms; never sold, leased, or shared without authorization. No employee or third party views files during upload/transcription.
- **Non-identifiable data** — collected passively; cannot identify you. Cookies are used for functionality and analytics; no Personal Data is collected via cookies without permission.
- **Aggregated data** — anonymized research that does not identify you, may be shared with affiliates and partners.

## Use of Personal Data

Used consistently with this policy: to provide and monitor access to Services, improve content and functionality, understand users, and (with opt-out) send marketing communications.

## Storage & Safety

Data may be transferred and stored outside the EEA using best-in-class providers. Methods include TLS encryption in transit and between data centres, enterprise-grade secure data centres (24/7/365 monitoring), and continuous security monitoring.

## Disclosure

VideoDB does not sell your information. Limited sharing may occur for business transfers, related companies, consultants/vendors, and legal requirements.

## Your Choices, Exclusions, Children, Third-Party Sites, Security

You can visit the Site without providing Personal Data. Policy excludes unsolicited information. VideoDB does not knowingly collect data from children under 13. Policy does not cover third-party sites. Reasonable steps protect data, but no transmission is fully secure.

## Your GDPR Rights

Right to be informed, access, rectification, erasure, restriction, objection, data portability, and not to be subject to automated decision-making.

## Retention

Data retained as long as necessary for the Services and legitimate business purposes. **Data related to video files is kept for 180 days.**

## Contact / DPO

Data Protection Officer — Spext Labs Inc. (operating under the brand name VideoDB): ashu@spext.co, 45 Lansing Street, #2111, San Francisco 94105 USA. General support: support@videodb.io.

---

<!-- source: /legal/security.md -->
# Security & Compliance — VideoDB

> We act as stewards of your data with multiple layers of security. We do not sell, rent, or share your information with third parties for promotional use. Our security certificates (SOC 2 Type II, ISO 27001, HIPAA, and GDPR) are available for review on request — request access via our certificate request form at /trust-centre#request-access and we'll share them under NDA. The Service is operated by Spext Labs Inc., the company operating under the brand name VideoDB. Current as of 30 September, 2025.

---

## Security posture

- **Encrypted in transit** — SSL encryption during transit.
- **Encrypted at rest** — Industry-standard 256-bit AES encryption at rest.
- **Protected by SOC 2** — SOC 2 Type II compliant; HIPAA and ISO 27001 certified.
- **Privacy First** — We never store your files and adhere to sovereignty laws. HIPAA & GDPR compliant.
- **Protected** — Data and logs are untraceable back to an individual user.
- **Secure Vendors** — AWS, GCP, and Azure compliance checks regularly.

## Access Control

You always retain access to your files (view, export, download) and can delete them anytime; deletion removes audio, video, and transcription completely. Employees cannot access your audio and transcripts without permission.

## Data Encryption

All data between you and VideoDB's servers uses field-standard TLS, including transfers between data centres for backup and replication.

## Network Protection

Multiple layers including firewalls, intrusion protection systems, and network segregation.

## Secure Data Centres

Enterprise-grade hosting facilities with 24/7/365 monitoring and surveillance, on-site security staff, and ongoing security audits.

## Security Monitoring

The security team continuously monitors systems, event logs, notifications, and alerts to identify and manage threats.

Contact: support@videodb.io

---

<!-- source: /legal/dpa.md -->
# Data Processing Agreement (DPA) — VideoDB

> VideoDB Data Protection Addendum. Forms part of the Agreement between Customer (Controller) and VideoDB (Processor) to comply with EU GDPR for the processing of Personal Data.

---

## Key terms

Defines Data Transfer, EU GDPR, Standard Contractual Clauses, Controller, Processor, and Sub-processor. Where the Agreement and this DPA conflict, the DPA prevails.

## Obligations

- **Controller** — warrants it has rights and legal basis to provide Personal Data, supplies privacy notices, requests purges where required, and promptly notifies the Processor of complaints, data-subject requests, or legal process.
- **Processor** — acts only on documented Instructions, assists with data-subject and regulator requests, and ensures onward transfers meet equal-or-higher protection standards.

## Data secrecy, audit, transfers

Personnel are trained and bound by confidentiality. Controller may audit (15 days' notice, at its expense). Transfers outside the EEA follow Schedule 1 / Standard Contractual Clauses.

## Sub-processors, breach, deletion

Processor may engage approved sub-processors (listed in Annex III) and remains liable for them. Personal Data breaches are notified without undue delay. On termination, Personal Data is returned or deleted (within ~30 days), and all copies deleted as soon as practicable.

## Schedule 1 — Annex I (Parties & Transfer)

- **Data Exporter** — Customer (Controller); **Data Importer** — Spext Labs Inc., operating under the brand name VideoDB (Processor), 45 Lansing Street, #2111, San Francisco 94105 USA; contact Ashish Choithani, Lead Engineer, contact@videodb.io.
- **DPO** — Ashutosh Trivedi, ashu@spext.co.
- **Data subjects** — Customer's authorized users. **Data categories** — name, address, DOB, age, education, email, gender, image, job, language, phone, related person/URL, user ID, username. **Sensitive data** — none. **Frequency** — continuous.

## Annex II — Technical & Organisational Measures

Security management system, personnel security, access controls (least-privilege, MFA/SSO), and data-centre/network security (AWS, multi-AZ resiliency, disaster recovery, vulnerability management, TLS/HTTPS encryption, multi-tenant isolation, secure destruction). Aligned to ISO/IEC 27001:2022.

## Annex III — Sub-processors

Amazon Web Services, Cloudflare, Google Cloud, OpenAI, Slack, SolarWinds, AssemblyAI, Linear, and GitHub. See the page for nature of processing, data categories, location, and each vendor's security resources.
