Engines Warming...
2026.08.18 // 09:00
@cy520569
One thing we've been thinking about while working on XMK Seedance is how quickly an AI video workflow becomes complicated once references are involved.
A text prompt is simple. You write an instruction, generate, review, and try again.
But real projects rarely stay that simple.
Someone might start with a product image, add a visual reference for lighting, another reference for camera movement, and then use an existing clip to explain the pacing they want. At that point, the challenge isn't just "generate a better video." It's making all of those inputs understandable and manageable.
That has changed the way we're thinking about the Seedance workflow.
We're experimenting with a more reference-led approach around XMK Seedance(https://www.xmk.com/seedance), where images, video, audio, and written instructions can work together instead of forcing the user to describe everything through increasingly complicated prompts.
The interesting product question for us is what happens before generation.
Which reference controls the subject?
Which one is only there for style?
What happens when two references suggest different directions?
And how much control should we expose before the interface itself becomes another creative obstacle?
We're still learning where that balance should be.
The biggest lesson so far is that supporting more inputs doesn't automatically create a better workflow. The product also has to help people understand what each input is doing.
For other builders working on generative products: how do you add more control without making the creation process feel heavier?
