Why reference to video, not image to video
Image to video starts from one frame, so the scene is fixed by that image. Reference to video treats the portrait as an identity and lets the prompt set any scene — cafe, gym, city — while keeping the face. That's what makes a series of posts possible.
Talking heads vs scenes
Short scenes with a line or two work well in Wan 3. Long, to-camera posts — tutorials, reviews — are better in the talking-photo tool, which lip-syncs a full script to the portrait.







