We compare our video model with full-resolution Wan2.1 and token-reduction baselines alongside the Level-of-Token layout. Full video playback makes it possible to inspect motion, temporal consistency, and visual detail beyond the paper’s sampled frames.