Visual AI Digest: Unpaired Video Training, Arbitrary Resolutions, and Character Consistency

Multi · August 20, 2026 · 2 min read · 6 sources

Research

Boosting Text-to-Video Models with Unpaired Image-Video Data

This paper introduces a plug-and-play approach that allows you to train an image-to-video model using only a text-to-video dataset. It's a big deal because it removes the need for expensive, hard-to-find paired image-video data, making video model training far more accessible.

Unlocking Arbitrary-Resolution Generation in Diffusion Models

New research focuses on making positional encoding in diffusion transformers (DiTs) more flexible, allowing them to generate images at arbitrary resolutions and aspect ratios without retraining. This tackles a major practical limitation of current models like SD3 and Sora.

Consistent Character Animation from a Single Portrait

This paper presents a framework for generating long, coherent videos from text and a single reference character image, addressing the common issue of character 'drift' or 'morphing' in video generation. It's a critical step toward practical, character-consistent animated storytelling.

MasterShots: Cinematic Storytelling in Long-Form Video Generation

This research introduces a framework for extending text-to-video models into long-form, high-definition video generation with complex narratives and cinematic language. It's pushing the boundary beyond short clips into coherent, story-driven filmmaking.

News

Mochi 1: A New Open-Source Text-to-Video Model Arrives

The open-source community is making waves with Mochi, a new state-of-the-art text-to-video model that's now available for fine-tuning, inference, and research. It's one of the most capable open models to date, challenging the dominance of closed-source players.

Stability AI Releases Stable Video 4D for Dynamic Novel-View Synthesis

Stability AI continues to roll out its video tools, with the Stable Video 4D research paper detailing a method for generating dynamic novel-view videos from a single object video. It's another step in building a full-stack, practical video pipeline for creators.

Stay Ahead

Delivered each morning.