AI First: Should Unity Generate Music and Video? Yes
TL;DR
Yes. Unity should generate music and video inside the editor. VberAI Studio already does both from a game UI plate: Music / Song and AI Video. Godot already has addons that run generation in the editor. After an engine adds the entry, the plate, the track, and the clip still have to share one context.
The picture on screen is already an input
The Godot Asset Library addon ElevenLabs API by wchc takes an API key, then generates speech from a prompt or from text already in the project. The current addon is text-to-speech. The setup is the same shape as music: generation happens in the editor, and the input is a prompt or material the project already has.
In VberAI Studio the input is the game UI plate itself. AI Sound writes a track card. AI Video writes a clip. Both land back on the canvas.


The boundary keeps moving out
Code and art are already common generation jobs. Music and video are the next ones to add. One platform can grow several automated generation passes, and take generation jobs that used to stay outside the editor further in.
After the entry exists, context still has to sync
Godot’s follow-up is already a plugin. A matching editor entry for Unity is the open question. Once that entry exists, the plate, the track, the clip, and the scene still have to carry the same context. If the UI changes and the earlier music or video stays on the old description, the next pass is out of date.
Closing
Yes. Unity should support music and video generation in the editor. Godot already runs generation there through a plugin. The plate, the track, and the clip still have to stay on one context.