# Claude + VideoDB skills edited our launch video. Then it watched its own cut.

> Raw takes, a hand-drawn deck, and a research paper went in. A subtitled, mastered, post-ready launch video came out. Every edit was decided by understanding.

Category: Product
Published: 2026-08-09

---

Last week we shipped the launch video for our research paper, [*Search over the Visual World*](https://labs.videodb.io/papers/search-over-the-visual-world.pdf). It runs 2 minutes 14 seconds across twelve narrative beats: talking-head takes intercut with a hand-drawn deck and a scroll through the paper itself, a generated music bed, and brand-styled captions burned in. Plus three vertical variants for social.

<!-- TODO: embed final cut (YouTube/CDN) -->

![A frame from the final cut: the research paper scrolling past its benchmark table, with the founder overlay](https://videodb.io/assets/images/Blog-post-preview/claude-edit-paper-scroll.webp)
*A frame from the final cut. The paper scroll is anchored on the benchmark table the agent found on its own.*

No human opened an editor. The editing was done by Claude in a chat session, using VideoDB's [agent skills](https://github.com/video-db/skills) as its eyes and hands. What follows is the receipts, real screenshots from the session. This is the clearest demonstration yet of what we mean when we say *the edit is a query*.

## What went in

Multiple talking-head takes (the usual founder-fumbling-lines footage, recorded on Loom), an animated HTML deck (12 slides, one per beat), the paper PDF, and a brief that fits in one breath: follow this script kit, overlay my takes, show the paper scrolling for credibility, generate a light music bed.

One scoping instruction is worth noting: *"for my videos you can just use the transcript index, there's nothing much visually to do"*. Take-picking is an audio problem, and the human knew it. The agent agreed and indexed the takes transcript-first.

## The first cut: 2:35, with judgment calls a junior editor would miss

The agent created a collection, then uploaded and transcribed the takes in the background. Meanwhile it rendered the deck headless and built a paper-scroll clip. On its own it found the exact results table in the PDF worth scrolling past: it went looking for the benchmark numbers and anchored the scroll on them.

Then it pulled timed transcripts, computed **word-boundary cut points for each beat**, and composited everything server-side on a [VideoDB timeline](https://videodb.io/blogs/infrastructure-that-sees-and-edits). The hook take runs full-frame, the paper cameo sits under the intro, and the deck carries beats 2 to 12 with the founder in a corner bubble. Slides advance on sentence starts, with the ambient bed at 12% volume.

![Session screenshot: the agent delivers the full 2:35 draft with a streaming link and its own notes](https://videodb.io/assets/images/Blog-post-preview/claude-edit-draft-delivery.webp)
*The draft delivery, in the session: 2:35, streamed for review before any MP4 download.*

Two details from the draft delivery that tell you this isn't template assembly:

- *"Math beat uses pass 1 — **pass 2 said '93,000'**."* Two takes covered the same beat. In one of them the founder said the wrong number. The agent read both transcripts, caught the discrepancy against the deck, and cut in the take where the number was right.
- *"Your recordings are HDR (that washed-out look) — all takes were **tone-mapped to SDR** before editing."* Nobody asked. It noticed.

It also flagged its own imperfections unprompted, each with a proposed one-line fix: the hook framing covering a title line, the bubble clipping a chart label.

## Directing an AI editor sounds exactly like directing a human one

The feedback round, verbatim:

> *"move the video overlay to the left bottom it won't mess with the diagrams on the right. The start of second slide starts randomly. The end closing slide and last 3 needs to choose correct take and correct starting points. The Hook should be only my video and no background… reverify the start and end of each cut in timeline and choose the exact start perfectly. Not bad for the first draft"*

No timecodes. No EDL. Director's notes. The agent triaged them the way a good editor would. Four of five were mechanical: bubble to bottom-left, hook full-frame, the slide-2 reveal anchored to a specific spoken line instead of popping mid-beat, and a **silence-aware re-audit of every cut-in and cut-out** so no segment catches the tail syllable of the previous sentence.

The fifth was the closing-take choice. It pulled the candidate takes from the transcripts and put the creative call back where it belongs, with the human.

![Session screenshot: the note 'bw slide 5 and 6 there's a lot of mess' and the agent's frame checks that followed](https://videodb.io/assets/images/Blog-post-preview/claude-edit-directors-note.webp)
*Another round, verbatim: "bw slide 5 and 6 there's a lot of mess." The agent went and looked.*

![Contact sheet of frames the agent sampled from its own render to check slide timing](https://videodb.io/assets/images/Blog-post-preview/claude-edit-frame-sampling.webp)
*How it looked: contact sheets of its own render, sampled to check the slide transitions frame by frame.*

## The part no other editing stack can do

Then came the instruction that makes this a different category of thing:

> *"whenever you have the final video, reupload and do visual index and audio index and verify that everything is perfect from editing and consumption point of view, if not try to replace the messedup portions. Do at least 1 pass of quality"*

![Session screenshot: the reupload-and-verify instruction, the agent's QA contact sheets, and the final cut delivered QA-passed on both indexes](https://videodb.io/assets/images/Blog-post-preview/claude-edit-final-cut-qa.webp)
*The instruction and the delivery: 40 commands later, the final cut ran 2:48, QA-passed on both indexes.*

So the agent rendered, then **uploaded its own cut back into VideoDB, indexed it for spoken words and visuals, and reviewed its own work.** And the pass earned its keep. It found that the draft's cut points were systematically early. Spoken-word timestamps were running about 1.3 to 2.6 seconds ahead of the file timeline, varying per file. That is exactly why some beats started mid-breath.

The fix: re-align every boundary against the real audio using silence fingerprints, re-cut in each file's own timeline, gate-check all seven talking segments with local ASR (first and last word verified), and re-stretch the deck piecewise so all twelve slide changes land on the corrected schedule.

QA on the shipped file. The transcript audit: every beat starts clean, every sentence completes. The visual pass: 12 of 14 windows with zero defects. The other two showed the corner bubble covering the paper's corner, which is inherent to any overlay. Final: **2:48, QA-passed on both indexes.**

Sit with the loop: *ingest → understand → decide → render → re-ingest → verify.* Every other programmatic editing tool stops at "render". The output ships un-watched, because assembly APIs have no eyes. Here, the same perception layer that picked the takes **watched the finished cut and fixed it.**

And because this was real dogfooding, the timestamp-offset finding went into a gaps doc for our product team, with repros. The edit session filed its own bug report.

![Session screenshot: commissioning the gaps doc for the team alongside the vertical output](https://videodb.io/assets/images/Blog-post-preview/claude-edit-gaps-doc.webp)
*Dogfooding for real: the gaps doc commissioned mid-session, with findings and repros straight to the product team.*

## Then the notes you'd give any editor

*"okay now 1.25x the whole video"* → 2:14, tempo-shifted with pitch preserved, dead-air verified gone. It also added an unprompted editor's caution: 1.25x on an already-brisk hook can read rushed. A variable-speed alternative was offered.

*"the audio of the section where I am on full screen the intro is low. can you fix it?"* → measured, not guessed. The hook came from a different recording session and was **7.7 dB quieter** than everything else. It boosted just that window (all beats now within 0.9 dB), then mastered the whole mix from −30 LUFS to **−16 LUFS**. That is the loudness X and LinkedIn feeds expect.

## Captions that became a brand asset

The subtitle request produced the detail we'd show any media team. Instead of shipping raw ASR, the agent rebuilt the caption text as a **corrected canonical script**. So the captions say *Qwen* (not "Quinn"), *TwelveLabs* (not "12 labs"), and *GPT* (not "GPD").

The text was chunked into 1 to 2 word pops (200 of them), butt-joined inside sentences so a word is always on screen. The styling is brand: bold uppercase, white with black outline, numbers in VideoDB orange, 70 ms fade. It was delivered as a custom caption track on the timeline, streamed for review *before* any MP4 was rendered.

![QA frames from the shipped cut with brand captions burned in, the keyword highlighted in VideoDB orange](https://videodb.io/assets/images/Blog-post-preview/claude-edit-final-frames.webp)
*QA frames from the shipped cut, with brand captions burned in and keywords in VideoDB orange.*

And then one more instruction: *"also store this subtitle style for future as videodb brand."* The style is now project memory: the template, the corrections list, and the review-first pipeline. Any future session reproduces it on any new video by asking for **"brand subtitles."** The edit session didn't just produce a video. It produced a reusable house style.

## Three verticals, art-directed in plain language

*"from timeline of the successful one, extract my video and audio but the deck build it specific to 9:16… you can also create a split variant where my video on top and the deck at the bottom (like social media has many such videos) do you get it?"* It got it. Variant A was deck-native 9:16, with a blurred deck texture, cinematic. Variant B was the creator-style split.

Then one more note: *"I want my video down and on 1/4th of the screen and give paper more"*. That produced Variant C: face as a full-width band in exactly the bottom quarter, captions in the seam, and the paper cameo **re-rendered as a portrait crop** so the results tables are actually readable in a phone feed instead of a letterboxed sliver. The deck was re-laid for vertical, not squeezed.

![A vertical variant streamed for review: the deck re-laid for 9:16 with the founder in a portrait band](https://videodb.io/assets/images/Blog-post-preview/claude-edit-vertical-variant.webp)
*A vertical variant streamed for review, with the deck re-laid for 9:16, not squeezed.*

<!-- TODO: embed Variant C vertical (YouTube/CDN) -->

## Why this works (and why "AI editing" tools can't)

Timeline-assembly APIs are hands without eyes: you compute every cut yourself and express it in JSON. Consumer AI editors have eyes but no API, so a human drives. Rough-cut copilots pick takes, but inside human NLE workflows.

The loop above requires something structurally different: **an index that decides, a timeline that renders, and the same index again to verify, because the output is just video.**

Understanding-driven editing. The edit is a query. The QA is a query too.

## Do this with your own footage

This wasn't a bespoke pipeline. It was Claude with VideoDB's agent skills (`npx skills add video-db/skills`). Any agent that can call tools can run the loop: upload takes, index, search for moments, compose, render, re-ingest, verify.

Start with the [quickstart](https://docs.videodb.io/pages/getting-started/quickstart), or read the [builder's guide to giving agents eyes](https://videodb.io/blogs/give-your-ai-agents-eyes). This post is what the "Act" verb looks like at full stretch.

## FAQ

**Can AI actually edit raw footage into a finished video?**
Yes, with understanding-driven infrastructure. The agent transcribes and indexes the takes. It selects the best take per script beat by querying transcripts, including catching factual slips between takes. Then it composes the timeline programmatically, renders, and re-indexes its own output to verify the edit. The session above shipped a QA-passed, mastered, subtitled 2:14 final plus three vertical variants.

**How does the agent pick the "best take"?**
By querying the indexes, not scrubbing footage. Transcripts show which take covers which beat cleanly. When two takes disagree (one said the wrong number), the transcript comparison catches it. Where the choice is genuinely creative, the agent surfaces candidates and asks.

**Can it check its own work?**
Yes, that's the differentiating loop. The rendered video is re-uploaded, indexed on both channels, and audited against the intended structure. Misaligned cuts get re-cut. Assembly-only APIs cannot do this, because they never perceive their output.

**Does it handle subtitles, loudness, speed changes, and vertical formats?**
All demonstrated in this one session: corrected-script captions in a stored brand style, per-beat loudness matching and −16 LUFS mastering, pitch-preserved 1.25x, and three art-directed 9:16 variants.

**Is this video generation?**
No. Every frame is real footage: takes, slides, paper pages. VideoDB is video understanding and editing infrastructure. It doesn't synthesize avatars or synthetic scenes. (The only generated element was the background music bed, which was requested.)

---

Run the loop on your own footage: start with the [quickstart](https://docs.videodb.io/pages/getting-started/quickstart) (free tier), or [talk to us](https://videodb.io/company#contact) about what an editing agent could do on your footage.
