The AI Post
Agents & CodingOpen ModelsEnterpriseFundraisingGenerative MediaGovernanceInferenceInfrastructureLegal & SafetySector Impact
← Front Page Research & Evals · Google · Google Research

Google Research shows a 10‑minute video from a multi‑agent pipeline

Four components handle storyboarding, generation and critique, and Google says the system holds characters and places steady across a video far longer than a clip.

Google Research published work on Thursday on a multi-agent system for long-form video, together with a 10-minute continuous video it says the system generated with visual consistency held throughout. The post is by Yale Song and Yiwen Song, research scientists at Google.

The system has four named parts, according to the post. AI Video Co-Director treats storytelling as a global optimisation and uses multi-armed bandit algorithms to choose among creative directions. CANVAS keeps a persistent visual memory of characters and of the state of objects, so transitions inside a setting hold rather than drift.

A²RD turns storyboards into minute-long video through a loop that retrieves, then synthesises and refines, switching between extrapolation and interpolation as it goes. VQQA generates visual questions about the output and feeds a vision-language model's critiques back as what Google calls semantic gradients for prompt optimisation.

The scores are Google's own. It reports 81.4 on GenAD-Bench for the Co-Director, continuity gains for CANVAS on HardContinuityBench, and improvements on T2V-CompBench and VBench2 from VQQA. The papers are on arXiv and the 10-minute video is on YouTube. No code or public demo appears to be offered.

Sources 1 source

  1. Source Google Research