Open-Source Challengers, Physics Priors, and the Efficiency Frontier
ByteDance open-sources lightweight video model Seaweed to rival larger systems
This paper from ByteDance introduces Seaweed, a streamlined 7B-parameter video generator that rivals models 10x its size. It achieves high-fidelity generation in just 5 seconds by using a decoupled spatio-temporal architecture and training on a large-scale private dataset.
PhysGen grounds video generation in Hamiltonian mechanics
The paper presents a system that grounds video generation in physical laws, notably conservations of mass, momentum, and energy. This enables physically plausible simulations of fluid and cloth dynamics, moving beyond purely visual realism.
New distillation method preserves spatial details in fast diffusion
The authors introduce a method to distill diffusion models into a 4-step student while preserving fine-grained spatial details. It achieves this through a novel noise schedule and a coupled distillation objective, significantly accelerating inference without sacrificing quality.
Physics-based hallucination removal improves video realism
This work proposes a framework to automatically detect and remove 'hallucinations' (unrealistic artifacts) from generated videos using a physics-based renderer and a semantic consistency check. It establishes a new benchmark for evaluating hallucination in video generation.
Linear attention architecture improves diffusion efficiency
The paper explores replacing convolutional layers in U-Net-based diffusion models with linear attention to reduce memory use and improve speed. The resulting architecture, called LinFusion, achieves comparable performance while being significantly more efficient for high-resolution generation.
Stay Ahead
Delivered each morning.