Why the two products belong together
A character model is good at establishing who someone is — appearance, wardrobe, manner. A video model is good at making that person move. Neither does the other's job well, which is why the handoff between them is the interesting part.
Step one: fix the character
Settle the character description first: build, hair, wardrobe, setting. Write it down as a fixed block of text you reuse. Any variation here becomes a visible inconsistency later.
Step two: produce a reference frame
Generate a still image of that character in the exact framing you want the clip to open on. Portrait framing for identity hold; a plain background for stability. This still is the anchor for everything that follows.
Step three: animate the anchor
Take the reference frame into an image-to-video tool and choose a low-magnitude motion template. The clip inherits the character from the image, so the video model only has to solve movement.
Step four: judge and re-anchor
If identity drifts, do not roll the same clip again — regenerate the reference frame with a tighter crop and animate that. Fixing the anchor is faster than fixing the animation.
Animate an image, not a description. The image is the character; the prompt is only the movement.