Video and Beyond

Video is where multimodal AI gets genuinely hard. It's not just images — it's images over time, plus audio, at a scale that dwarfs a single picture. The temporal dimension adds motion, causality, and continuity that a still frame can't capture, and the sheer data volume strains everything. Video is also the frontier where the most impressive recent generation results have appeared, and where multimodal AI is actively pushing forward. Understanding video — and the other modalities beyond the core ones — shows where the field is heading.

This post covers video — the modality that adds time to vision — and other modalities beyond the core text/image/audio: video understanding and generation, why the temporal dimension is hard, and the broader landscape of modalities (3D, sensor data, and more). It’s the frontier-facing post, showing where multimodal AI extends beyond the established modalities. Video builds on everything prior (vision, generation, audio) while adding the challenge of time.

Video: vision plus time

Video is a sequence of images (frames) over time, usually with audio — so it combines vision with a temporal dimension (and often audio). This makes it both richer and much harder than still images:

Video is vision plus time (plus audio) — a sequence of frames capturing motion and change — making it the richest and hardest common modality, requiring temporal understanding beyond static vision. Its two capabilities, understanding and generation, both build on earlier posts while grappling with time and scale.

Video understanding and generation

The two core video capabilities — understanding video and generating it — extend the vision capabilities into the temporal dimension:

Video understanding (recognizing actions/events, reasoning about what happens over time) and video generation (creating temporally-coherent video from text, extending diffusion to motion) both extend vision capabilities into time — with temporal coherence the central, hard challenge. Video generation especially is a rapidly-advancing frontier. But video’s difficulty is compounded by its sheer scale.

Why video is hard: the scale challenge

Beyond temporal coherence, video is hard because of its scale — the sheer volume of data — which strains computation and training:

Video is hard because of scale (a video is thousands of images plus audio — enormous data straining compute and memory) on top of temporal coherence and long-range temporal modeling. These challenges are why video AI lags images but is a very active frontier — advances often come from efficiency innovations that make video scale tractable. Video is the hardest common modality, but not the only frontier.

Beyond the core modalities

Multimodal AI extends beyond text, image, audio, and video to other modalities — showing how general the approach is and where the field is heading:

Multimodal AI extends beyond the core modalities to 3D/spatial data, sensor and scientific data, and embodied/robotic multimodality — showing that the same recipe generalizes to essentially any modality. Video (vision plus time) is the hardest common modality (temporal coherence plus scale) and a fast-advancing frontier, and beyond it, multimodal AI keeps expanding. Next, the final post: building with multimodal AI and the any-to-any future.

Key takeaways

Further reading

Sources & References

Generation extended to video