Veo 3.1, Kling, Sora 2 and Grok: a test of five AI video generators with identical prompts
Habr authors compared five leading AI video generators — Veo 3.1, Kling v3 Omni, Sora 2 Pro, LTX-2.3 Pro and Grok Imagine Video 1.5 — across seven identical scenarios: a medieval knight, a juggler, a panicky cat and Will Smith with spaghetti. Same prompts, same settings, equal conditions — a fair stress test in mid-2026.
AI-processed from Habr AI; edited by Hamidun News
The Habr authors tested five leading neural networks for video - Veo 3.1, Kling v3 Omni, Sora 2 Pro, LTX-2.3 Pro, and Grok Imagine Video 1.5 - running each through seven identical scenarios with the same prompts and settings.
What was included in the test
Five tools from different ecosystems participated in the comparison:
- Veo 3.1 — the latest video model from Google DeepMind, third generation of the line
- Kling v3 Omni — the flagship product of Chinese studio Kuaishou
- Sora 2 Pro — OpenAI's advanced video generator, developing the original Sora from 2024
- LTX-2.3 Pro — open-source solution from Lightricks, a notable alternative to closed models
- Grok Imagine Video 1.5 — video model from xAI's ecosystem, integrated into the Grok platform
The selection covers different market segments: American AI laboratories, Chinese studios, and open-source - without bias towards a single player.
How the methodology works
The key principle of the test is strict reproducibility. Each model received the same prompt with identical parameters: resolution, duration, basic generation settings. Seven scenarios of increasing complexity.
Among them were a medieval knight in motion, a juggler with props, a panicking cat, as well as the classic stress test - Will Smith greedily eating spaghetti. The latter scenario became a meme in the AI community: early models generated frankly failed results here with distorted hands and unnatural physics.
The set specifically covers historically weak points of generative video: hand and finger movement physics, facial expressions, object interaction, spatial realism.
Why the comparison is relevant right now
Back in 2024, AI video was easily recognized by characteristic artifacts - drifting faces, extra fingers, jerky movements. Over the past year and a half, the market has gone from curiosities to tools actually used in marketing, film, and advertising.
"In previous years, neural network video was noticeable by characteristic flaws.
How are things in mid-2026?" — the authors frame the central question of the test.
Today, the user faces a different challenge: not "distinguish from real," but "choose the right tool from dozens available." Each model promotes its own benchmarks and marketing videos — an independent test under equal conditions remains rare.
What it means
Comparison using single prompts is one of the most honest formats for evaluation: manufacturers' marketing claims are not taken into account. The results are useful for those choosing a tool for specific tasks - marketing videos, prototypes of visual scenes, content for social networks. The AI video market has definitively left the demonstration phase — now it's a practical choice with real stakes.
Which neural networks for video were tested?
Veo 3.1 (Google DeepMind), Kling v3 Omni (Kuaishou), Sora 2 Pro (OpenAI), LTX-2.3 Pro and Grok Imagine Video 1.5.
Which neural network for video is the best?
In the Habr test, five tools were compared: Veo 3.1 (Google DeepMind), Kling v3 Omni (Kuaishou), Sora 2 Pro (OpenAI), LTX-2.3 Pro and Grok Imagine Video 1.5. Each model went through seven identical scenarios for objective comparison.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.