Generative Image & Video AI
Diffusion models turned text prompts into images in 2022 and video by 2024-25; open models run on a home GPU, and Veo 3 added synchronised sound in May 2025.
Open in the interactive tree →A diffusion model learns to remove noise step by step; the idea dates from 2015 and became practical with the 2020 DDPM method. DALL-E 2 (April 2022), Midjourney (July 2022) and the open Stable Diffusion (22 August 2022, trained for about US$600,000 on 256 A100 GPUs and able to run on consumer GPUs) made text-to-image a mass tool. OpenAI previewed its Sora video model on 15 February 2024, and Google's Veo 3 (May 2025) generates dialogue, effects and ambient sound together with the picture.
As of October 2026
Image models keep turning over: OpenAI's GPT Image 2 arrived in April 2026 and Ideogram 4.0 in June 2026. Video is costly: OpenAI released Sora 2 with a social app on 30 September 2025, announced its shutdown on 24 March 2026 and closed the app on 26 April and the API on 24 September 2026, reportedly citing compute shortages and a cost of about US$1 million a day, while models such as Kling 3.0, Seedance 2.0 and Runway 4.5 rank higher on industry leaderboards. Copyright disputes continue: as of November 2025 Getty Images had largely lost its lawsuit against Stability AI.
Open steps
- Robust marking of synthetic media Medium AI leverageEU rules require machine-readable marking of AI-made images, audio and video from 2026, yet the Commission's code of practice says no single marking technique is enough.
- Physically consistent long video High AI leverageVideo models still break object permanence and physics. A January 2026 benchmark found none of the tested world models consistent enough for reliable real-world use.
- Cheaper video generation High AI leverageSora's app and API were shut down in 2026, reportedly at about US$1 million a day in costs. Distilled models now stream 720p video at 24 FPS. How far can cost fall?
- Rights to training data Low AI leverageCourts have not settled whether training image models on copyrighted works is lawful. Getty won permission to appeal its UK loss against Stability AI on a point of law.
Where AI could help
High AI leverage. The models are AI end to end; progress follows data and compute, while serving cost and training-data rights now limit use.
- Synthetic captions and filtered data to train better models
- Distillation that cuts the cost per image or clip
- Automatic scoring of prompt faithfulness and physical plausibility
- Provenance marks and detectors for synthetic media
Shown so far
Prerequisites
- Thermodynamics1824Diffusion image models are derived from non-equilibrium thermodynamics (2015)
- Graphics Processor (GPU)1999
- Deep Learning2012
- Generative Adversarial Networks2014GANs gave the first photorealistic generated images
- Transformer Architecture2017
- Diffusion Models2020Diffusion models power current image and video generators