Both modes produce a clip from a prompt, so they look interchangeable in a picker. They are not. One decides what the subject looks like and then moves it; the other takes a subject that already exists and only decides how it moves. Choosing wrong is the most common reason a generation comes back technically fine and completely unlike what you had in mind.
Browse image-to-video templatesIn text-to-video the model invents everything: the face, the room, the wardrobe, the light, and the motion. In image-to-video you have already fixed the face, the room, the wardrobe and the light by supplying a frame, and the only thing the model still decides is what happens next. That is the whole distinction, and every practical consequence follows from it. If you care what the subject looks like, you are in image-to-video territory whether or not you realised it.
Across the model families available here, image-to-video endpoints outnumber text-to-video by roughly two to one. The industry converged there because consistency is the hard part of video, not motion. Given a starting frame, a model has an anchor for identity, lighting and framing across every subsequent frame. Given only words, it has to re-derive all of that from scratch, and small drifts compound over the length of a clip.
A text-to-video prompt has to carry appearance, setting, camera and motion at once, and it will be read as a whole. An image-to-video prompt should carry almost nothing but motion and camera — describing the subject again is at best wasted and at worst actively harmful, because you are inviting the model to reinterpret something it can already see. The most common image-to-video mistake is pasting in a text-to-video prompt and wondering why the face changed.
Ask whether you can name the person in the shot. If the answer is "a woman in a red dress" — a description, not an individual — text-to-video is fine and gives you more variety per attempt. If the answer is "this person, from this photo", image-to-video is the only mode that can deliver it, and no amount of prompt detail will make text-to-video hold a specific likeness. A useful middle path is to generate a still first, keep the one you like, and animate that.
Text-to-video disappoints on repeatability: run the same prompt twice and you get two different people. Image-to-video disappoints on ambition — it will not relocate your subject to a different room, change the time of day, or add a second person who is not in the source frame. Asking it to do those things usually produces a clip that ignores the instruction rather than one that fails visibly, which is why it wastes credits quietly.
Some model families accept a start frame and an end frame, and generate the transition between them. It is worth knowing this exists because it solves a problem neither of the other modes handles well: getting a clip to arrive at a specific final state. With a single start frame you are asking the model to invent where the motion ends up, and it will pick something plausible rather than something you chose. With both ends fixed, the model is interpolating rather than improvising, and the result is far more predictable — at the cost of needing you to produce a second image that is consistent with the first.
Motion instructions fail most often because they describe a feeling rather than a movement. "Dynamic", "cinematic" and "lively" are not movements; they are adjectives the model has to translate into one, and it will pick the average of everything it has seen. Movements are things with a direction, a subject and a rate: the camera pushes in slowly, the hair lifts in a light breeze, the head turns left and returns. A few model families also expose an explicit amplitude control, which is worth using when it exists — it is the difference between asking for less motion and hoping for it.
One thing that does not change between the two modes is what drives the bill: duration first, resolution second. It is tempting to assume image-to-video is cheaper because you supplied half the information, and that is not how the pricing works — the model still renders every frame either way. Choosing a mode is a creative decision; choosing a length is a budget decision. Making them separately, in that order, is the habit that keeps iteration affordable.
Not reliably. Text-to-video re-invents the subject every run. If a particular likeness matters, generate or supply a still and animate it.
Usually the opposite. Once the frame is supplied, extra description of the subject invites reinterpretation. Keep the prompt to motion, camera and pacing.
Drift accumulates frame by frame, so it shows up later in the clip. A shorter duration is the direct fix; a different prompt usually is not.
Neither mode is inherently more expensive — duration and resolution drive cost far more than the choice of mode. See the guide on choosing duration and resolution.
Only slightly, through camera movement and lighting shifts. It cannot move your subject to a different location. That needs a new still.











