Visual AI Digest: China's Sora Challengers, Video Scaling Laws, and Long-Prompt Grounding

Multi · September 15, 2026 · 1 min read · 5 sources

News

Alibaba and ByteDance Launch New AI Models to Compete with OpenAI's Sora

The Chinese AI giants are directly challenging OpenAI's Sora dominance with new text-to-video models. This move signals that the high-end video generation market is rapidly becoming a multi-player field, not a US-centric monopoly.

Research

VideoTree: Adaptive Tree-based Visual Tokenization for Long Video Understanding

This paper introduces 'VideoTree,' a method to build adaptive visual tokens for long-form video understanding. It's a practical tool for anyone trying to get better performance out of their video models without massive compute costs.

Multi-Entity T2I: Improving Generation of Complex Prompts

The paper proposes a new approach to improve text-to-image generation for prompts with many distinct entities. It directly tackles a common pain point—generating complex scenes where objects don 't blend or disappear.

Long-Prompt Grounding in Text-to-Video Generation

This work focuses on grounding long, descriptive text prompts in video generation. It's a key step toward making video AI that can actually follow a detailed script rather than just a simple command.

Scaling Laws for Video Transformers

The paper analyzes scaling laws specifically for video Transformers. Understanding these laws is critical for anyone planning to train or fine-tune a large video model, helping predict performance gains from more data or compute.

Stay Ahead

Delivered each morning.