Apple ML Research presented a method for generating video with sound from text using Text-to-Sounding-Video
Researchers at Apple ML Research presented work on the Text-to-Sounding-Video direction — generating video with synchronized sound directly from text. The authors identify two unresolved issues: interference between video and audio signals under a common text condition and an unclear optimal mechanism for merging features from different modalities.
AI-processed from Apple ML Research; edited by Hamidun News
Apple ML Research presented a method for generating videos with synchronized sound from text descriptions — a direction the paper calls Text-to-Sounding-Video (T2SV).
What is Text-to-Sounding-Video
T2SV is the task of simultaneously generating video sequences and audio tracks from a single text description, where both modalities — video and sound — must be coordinated with that specific text condition, not just with each other. Previously, researchers focused mainly on joint learning of video and audio, but according to the authors, two key problems in this field remain unsolved.
What problems the researchers are solving
The first problem is a bottleneck in text conditioning. If the same shared text prompt is used for generating both video and sound — denoted in the paper as TV=TA — interference arises between modalities: the signal for video and the signal for audio begin to interfere with each other. The situation is further complicated by the gap between detailed, information-rich descriptions used to train the model and short, concise prompts that users actually write when making requests.
The second problem is the lack of clarity about which feature fusion mechanism is optimal for the interaction between video and audio. The authors note that to address the first problem — the text conditioning bottleneck — they propose a new approach to condition processing and modality interaction.
- The research direction is called Text-to-Sounding-Video (T2SV)
- The goal is to generate video and sound simultaneously from a single text description
- Two key unsolved problems are identified: modality interference with shared text and an unclear fusion mechanism
- A separate gap is noted between detailed training descriptions and short user prompts
Why this is a complex task
Unlike ordinary video generation without sound, T2SV requires the model to simultaneously hold the meaning of text in two different modalities — visual and auditory — while preventing one signal from drowning out the other. This is why, as the authors note, simply using a shared text description for both modalities turns out to be insufficient: it provokes the very interference discussed in the paper. The problem is compounded by the fact that real users almost never write prompts as detailed and structured as those the model was trained on — meaning a solution that works well on training data may not work as well on a short request from an actual end user.
This gap between how the model is trained and how it is actually used is what the authors identify as one of the main obstacles to quality video-with-sound generation.
What this means
Apple ML Research's work highlights specific technical barriers on the path to truly coordinated video-with-sound generation from text — and the fact that a major research lab explicitly names fundamental problems as unsolved shows that the industry still has several steps to go before convenient mass-market text-to-video-with-sound tools become available.
Frequently asked questions
What does the term Text-to-Sounding-Video mean?
It is the generation of video together with synchronized sound directly from a text description, where both output signals — video and audio — must correspond to the text itself, not just coincide in time with each other.
What problems prevent such generation today?
The authors identify two: interference between video and audio signals when using a shared text condition, further complicated by the gap between detailed training descriptions and short user prompts, as well as the lack of a clear optimal mechanism for fusing features from different modalities.
Want to stop reading about AI and start using it?
AI News is a curated feed of AI/tech news. Hamidun Academy teaches you to use AI systematically in your work.
The AI world, distilled — once a week
Seven stories that actually mattered, hand-picked. No noise, no reposts, no press releases.
Done! Check your inbox for a confirmation.