Franklin AI Explainer

Google’s AI Video Co-Director Targets Long-Form Consistency

Key Takeaways

  • Long-form AI video needs continuity, not just higher-quality individual clips.
  • Persistent visual memory and automated critique could make multi-shot generation more reliable for creators and developers.
  • Public code for Co-Director and A²RD gives builders a starting point for testing agentic video workflows.

Google Research has introduced an AI video co-director designed to turn short generated clips into coherent, minutes-long videos. The system combines four agentic frameworks for creative planning, visual memory, long-form generation, and prompt refinement, targeting two persistent problems in multi-shot AI video: identity drift between scenes and cascading errors caused by flawed upstream assets, according to Marktechpost.
The orchestration layer runs on Gemini and Veo, although Google describes the approach as model-agnostic and says it can drive other video generators. Outputs inherit SynthID watermarking from the underlying models. The work is not presented as a single Google product; instead, it brings together research systems with different roles in long-form generation.

Four agents address different failure points

The first framework, Co-Director, treats creative planning as a search problem. Accepted at COLM 2026, it uses a multi-armed bandit to explore combinations of creative strategy, narrative mode, and aesthetic archetype.
An Orchestrator Agent selects a configuration, while a Pre-Production Agent develops the storyboard. Keyframe, video, and audio sub-agents then produce the media. An MLLM Judge evaluates the resulting cut and sends a factored reward back to the bandit. This allows the system to assess different creative directions rather than relying on one fixed sequence of handcrafted prompts.
The design addresses a difficulty in multi-agent video workflows: when the final cut fails, it can be difficult to identify which earlier prompt or generated asset caused the problem. By feeding evaluation results back into the planning process, Co-Director is intended to make the pipeline a global optimization task rather than a simple chain of independent generation steps.
The second framework, CANVAS, focuses on persistent visual memory. It tracks characters, locations, and object states as a story progresses, then retrieves visual anchors when the narrative returns to an earlier setting or entity.
Google tested CANVAS with a museum heist scenario. In the comparison described by the research team, AutoStudio lost the thief’s cap and Gemini-3.1-Pro changed the gemstone, while CANVAS preserved both details. The result illustrates the continuity problem the framework targets: a video can remain visually polished while still changing important story elements from one shot to the next.

Memory and iterative refinement for longer videos

The third system, A²RD, or Agentic Autoregressive Diffusion, generates a video segment by segment without additional training. Each segment passes through a Retrieve, Synthesize, Refine, and Update loop that consults a multimodal video memory.
The agent can extrapolate when the story introduces a new beat and interpolate when previously seen entities return. A long-form system needs to create new content while preserving the appearance and state of characters, locations, and props introduced earlier.
Google reports that A²RD was evaluated on videos ranging from one to 10 minutes and produced a continuous 10-minute film. The framework reported improvements of up to 30% in consistency and 20% in narrative coherence on those videos. Its code is publicly available through the project’s GitHub repository.
The fourth framework, VQQA, uses closed-loop prompt refinement. It generates visual questions for each prompt, then uses vision-language model critiques as “semantic gradients” to rewrite the text prompt. The method does not require access to the internal workings of the video generation model.
Rather than accepting the latest output automatically, VQQA includes a Global Selection step that chooses the strongest video from all iterations. Google reports absolute gains of 11.57% on T2V-CompBench and 8.43% on VBench2 compared with vanilla generation.
Together, the systems cover different stages of production. Co-Director searches over high-level creative choices, CANVAS maintains persistent visual facts, A²RD extends the story across segments, and VQQA revises prompts based on observed visual shortcomings.

Benchmarks measure continuity and narrative quality

Google introduced three benchmarks alongside the frameworks. GenAD-Bench contains 400 advertising scenarios covering 200 fictional products from 50 brands. HardContinuityBench tests whether systems maintain consistency when scenes reappear and when props change state. LVBench-C includes 120 scenarios in which key assets disappear for at least 10 segments before returning.
On GenAD-Bench, Co-Director achieved an average score of 81.4 and received a 3.96 out of 5 human rating. The reported baseline for random search was 75.7. The evaluation included comparisons with Veo 3.1, Kling 3.0 Omni, Wan 2.6, and MovieAgent.
CANVAS delivered reported gains of 21.6% in background continuity, 9.6% in character consistency, and 7.6% in prop consistency. Those measurements focus on visual details likely to drift when a generated story moves between locations or revisits earlier events.
The systems also differ from other long-video approaches in their implementation. StoryMem, developed by ByteDance and NTU, uses a memory-to-video diffusion model and fine-tunes its base model with LoRA. MovieAgent from Show Lab at NUS uses multi-agent planning and per-character customization. AutoStudio uses multiple language-model agents with a Stable Diffusion-based agent for image sequences rather than video.
Google’s reported approach does not fine-tune the underlying generators. It orchestrates Gemini and Veo through an external agentic layer, which could make the design adaptable to other generators, although the available reporting does not establish how broadly that portability has been demonstrated.

Code availability remains partial

Developers can access some components now. Google has made Co-Director and A²RD code available on GitHub, while CANVAS is listed as forthcoming. The complete four-framework pipeline is not described as a Google product, and the available code does not necessarily represent a single turnkey system for producing long-form videos.
The research points to a shift in how AI video generation is being structured. The central challenge is no longer only the quality of an individual clip; it is maintaining a shared world state while making creative decisions across many shots. Google’s combination of planning agents, visual memory, segment-level generation, and iterative critique offers one approach to that problem, with the 10-minute A²RD film serving as the clearest reported demonstration of the target workflow.

Our read

Franklin AI Take

Google’s work suggests that near-term progress in AI video may depend as much on orchestration as on larger generation models. Planning, memory, evaluation, and prompt revision directly address the production problems that make short clips difficult to assemble into coherent stories. The reported benchmark results and 10-minute demonstration are notable, but they remain research findings. The key open question for builders is whether this external agentic layer can maintain its benefits across different video generators and more varied real-world workflows.