TL;DR

Yes. Unity should generate music and video inside the editor. VberAI Studio already does both from a game UI plate: Music / Song and AI Video. Godot already has addons that run generation in the editor. After an engine adds the entry, the plate, the track, and the clip still have to share one context.

The picture on screen is already an input

The Godot Asset Library addon ElevenLabs API by wchc takes an API key, then generates speech from a prompt or from text already in the project. The current addon is text-to-speech. The setup is the same shape as music: generation happens in the editor, and the input is a prompt or material the project already has.

In VberAI Studio the input is the game UI plate itself. AI Sound writes a track card. AI Video writes a clip. Both land back on the canvas.

Music card on the canvas after AI Sound

AI Video clip on the canvas, 832 by 480

The boundary keeps moving out

Code and art are already common generation jobs. Music and video are the next ones to add. One platform can grow several automated generation passes, and take generation jobs that used to stay outside the editor further in.

After the entry exists, context still has to sync

Godot’s follow-up is already a plugin. A matching editor entry for Unity is the open question. Once that entry exists, the plate, the track, the clip, and the scene still have to carry the same context. If the UI changes and the earlier music or video stays on the old description, the next pass is out of date.

Closing

Yes. Unity should support music and video generation in the editor. Godot already runs generation there through a plugin. The plate, the track, and the clip still have to stay on one context.