Visual AI Digest: Sora Rivals, Controllable Editing, and Scaling Laws

Multi · September 14, 2026 · 1 min read · 6 sources

Research

DESM: Story-Level Cartoon Video Generation

The Alibaba and ByteDance models represent a major new competitive frontier, but the paper (from Daniel Gao) introduces DESM, a training-free framework that generates a high-quality

FrameBridge: Efficient Video-to-Video Translation

FrameBridge tackles a major pain point in T2V conditioning. It provides a plug-and-play method for video models to learn from and react to temporal control signals (like user edits or external motion data) using just 1-2 reference frames, making consistent, controllable video editing much more practical.

Multi-Entity Self-Rectification for T2I

This paper attacks the notoriously difficult problem of generating high-quality images with multiple characters and complex spatial relationships. It introduces a novel self-rectifying mechanism that significantly improves prompt fidelity for instructive, detail-rich T2I generation.

Hierarchical Counterfactuals for Text-Grounded Video Generation

Most T2V models struggle with long, complex prompts. This paper introduces a hierarchical visual-semantic counterfactual learning approach to explicitly train models to better ground their generation in the detailed text instructions, a key step toward reliable cinematic control.

News

Alibaba, ByteDance launch new AI models to compete with OpenAI Sora

This news coverage directly connects to the DEep Story Model (DESM) paper, confirming that major players are now aggressively shipping

Analysis

Video Scaling Laws: Depth vs. Width in DiTs

A deep technical dive into scaling laws specifically for T2V models. The authors find that, counter-intuitively, investing compute in model depth (layers) yields better returns than width (parameters) for video tasks, providing a crucial practical guide for efficient architecture design.

Stay Ahead

Delivered each morning.