Turn a Photo Into a Video
By the OFMAI Team · Updated July 2026
Short-form feeds favour video over stills, which means a library of good photographs is an underused asset rather than a finished one. Turning them into short clips is the least effortful way to increase what you publish without producing anything new.
There is one rule that separates results that look real from results that look obviously synthetic, and it is the opposite of most people's instinct. It comes from the source material and it is the most valuable thing in this lesson: start with subtle motion. A smile forming, a slow turn of the head, a hair flick. Complex actions produce artefacts.
Why subtle motion works and ambitious motion does not
Generative video models are extrapolating — inventing frames that did not exist. Every frame they invent is an opportunity for something to go wrong, and the amount that can go wrong scales with how much has to change between frames.
A slow head turn changes very little per frame. The face stays broadly where it was, the lighting stays consistent, the background does not move much. Errors are small and land inside the range a viewer reads as normal.
A complex action — a full body turn, hands manipulating an object, walking across a room — changes a great deal per frame. Limbs have to be tracked as they occlude and reappear, objects have to keep their shape while being handled, the background has to remain coherent as the camera relationship shifts. Each of those is a place where the output can fail visibly: fingers that merge, an object that changes shape mid-motion, a face that drifts away from the person it started as.
The asymmetry matters more than it seems, because these failures are not graceful. A subtle motion that comes out imperfect still reads as a slightly odd video. A complex motion that comes out imperfect reads as fake — and once a viewer has registered that, they have stopped watching and will not be persuaded back by the second half.
So the rule is not about being timid. It is about spending your risk budget where it pays. Ask for one small thing and you will get it reliably. Ask for three big things and you will usually get none of them convincingly.
The motions that reliably work
A short catalogue worth working through, roughly in order of how safe they are.
The safest is facial motion with no body movement at all: an expression forming, a blink, a slow smile, eyes moving to the camera. The frame stays almost entirely still, so almost nothing can break, and the effect on a viewer is large relative to how little is happening — a still photo that comes alive by a few degrees is genuinely arresting.
Next, slow head and shoulder motion: a turn towards the camera, a tilt, a glance away and back. Still a small change per frame, still no limbs crossing each other.
Then hair and fabric: a hair flick, hair settling, a garment moving in the air. These read as very natural because they are exactly the kind of secondary motion a real photograph implies but cannot show — and because there is no rigid structure that has to stay consistent, small imperfections are invisible.
Then ambient environment: light shifting, something moving softly in the background, steam or smoke. The subject barely moves at all and the frame still feels alive.
Beyond that, the risk rises quickly. Anything involving hands doing something specific, anything where the subject changes position substantially, anything requiring an object to be held and manipulated — these can work, but expect to spend several attempts getting one usable result, and price that expectation in before you start.
Producing the clip
There are two routes, and which one you want depends on whether you are working from your own photo or from a reference someone else's post gave you.
For your own material, the video generation form is the direct route. You describe what should happen and supply reference images — your character's likeness, and optionally an outfit, a setting or a prop — and get a clip built to that description. This is the path for "I have this look and I want a few seconds of it moving." Keep the description to the single motion you decided on; every additional instruction is another thing the render has to get right.
For a reference that came from someone else's post, the replication path from earlier in this module is usually better, because it does not require you to describe the motion at all — the source already contains it. Faithful mode in particular is a natural fit here: the source's opening frame becomes your character's starting image, and the source video supplies the movement. If the outlier you found is a still image rather than a video, image replication at 3 credits is the cheap way to get your character into that composition first, and you can decide afterwards whether it deserves a video.
Format the output for where it is going. Vertical is the default for short-form feeds and there is rarely a reason to deviate. Duration should follow the motion rather than the other way around: a hair flick does not need eight seconds, and padding a short motion out to fill a longer clip produces exactly the dead frames at the end that make a viewer leave.
What it costs, and how to spend it sensibly
Video generation is priced by resolution and duration together. The baseline is 480p at five seconds for 15 credits. 720p costs twice the baseline and 1080p five times, with duration scaling linearly on top of whichever resolution you pick — so a 1080p clip is a substantially larger commitment than a 480p one of the same length.
That pricing shape has a clear implication for how to work: test at the baseline, finish at whatever resolution the destination justifies. Whether a motion is going to look convincing is entirely visible at low resolution — an artefact does not hide at 480p and appear at 1080p, it does the reverse. Running your first attempts cheaply and only committing to a higher resolution once you have a version that works is the difference between an affordable habit and an expensive one.
If you are going through replication instead, the costs are the ones from lesson two: 3 credits flat for an image, and 15 credits per five seconds for faithful mode at standard quality. The image route in particular is worth abusing while you are learning what suits your character, because it is cheap enough to be wrong repeatedly.
Judging the result
Watch it once at full size and once at the size it will actually be seen — a phone, held at arm's length, scrolled past. The second viewing is the one that counts, and things that look wrong on a large screen often disappear at feed size while other things become more obvious.
Look specifically at the places these renders break: hands and fingers, the boundary where hair meets background, anything held or touched, and whether the face is still recognisably the same person at the end as at the beginning. Identity drift across a clip is the failure that does the most damage to a character account, because consistency is the entire proposition.
If something is wrong, reduce before you retry. The instinct is to add detail to the description to correct the error; the better move is usually to remove the part of the motion that broke. A clip of a smile forming that works is worth more than a clip of a smile-and-turn-and-reach that nearly works.
And accept a hit rate below one hundred percent. Even well-chosen subtle motions will occasionally come back wrong. Testing cheaply is what makes that acceptable rather than expensive.
Animate a still
- Pick a photo that has somewhere to go — Images already suggesting motion — a look towards the camera, hair caught mid-movement, an expression about to change — animate far better than fully static poses.
- Decide on one motion — A smile forming, a slow head turn, a hair flick. Write it in a few words. If you need a comma, it is probably two motions.
- Choose your route — Your own photo, describe the motion yourself on the generation form. A reference from someone else's post, use replication and let the source supply the movement.
- Set it vertical and short — Vertical for short-form feeds, and a duration that matches the motion rather than padding it out with dead frames.
- Test at the baseline resolution — Whether the motion works is visible at 480p for 15 credits. Do not spend five times that finding out.
- Review at feed size — Watch on a phone. Check hands, the hair-to-background edge, and whether the face is the same person at the end as at the start.
- Reduce, then re-render — If something broke, remove the part of the motion that broke rather than adding description to correct it. Then re-render at the resolution the destination justifies.
Motions ranked by how reliably they render
- Facial only — a blink, a slow smile, eyes moving to camera. Safest; almost nothing changes per frame.
- Head and shoulders — a slow turn, a tilt, a glance away and back. Still small per-frame change, no limbs crossing.
- Hair and fabric — a hair flick, hair settling, a garment moving. Reads as natural; no rigid structure to keep consistent.
- Ambient environment — shifting light, soft background movement. The subject barely moves and the frame still feels alive.
- Hands doing something specific — risky. Fingers merge and objects change shape mid-motion.
- Full body repositioning or walking — expect several attempts per usable result. Price that in before starting.
Frequently asked questions
- Why do complex actions produce artefacts?
- Because the model is inventing every frame after the first, and the amount that can go wrong scales with how much has to change between one frame and the next. A slow head turn changes very little: the face stays roughly in place, the lighting holds, the background is stable, so any errors are small enough to sit inside what a viewer reads as normal. A complex action changes a great deal per frame — limbs occlude and reappear, objects have to hold their shape while being handled, the background has to stay coherent as the camera relationship shifts. Every one of those is a place the render can fail visibly. And the failures are not graceful: an imperfect subtle motion still reads as a slightly odd video, while an imperfect complex motion reads as fake, at which point the viewer has already gone.
- How long should the clip be?
- Long enough for the motion to complete and no longer. This is a real constraint rather than a stylistic preference: duration multiplies the cost linearly, and padding a short motion out to fill a longer clip leaves dead frames at the end, which is precisely where a viewer decides to scroll. A hair flick or a forming smile resolves in a couple of seconds; forcing it to fill eight means six seconds of your character doing nothing after the interesting part has finished. Let the motion set the length rather than picking a length and finding motion to fill it. If you genuinely need a longer clip, that is usually a sign you want a sequence of short ones cut together rather than one long render — which is also cheaper, since you can discard the attempts that failed instead of re-rendering the whole duration.
- Should I test at a low resolution first?
- Yes, and the pricing makes the case on its own. The baseline of 480p at five seconds is 15 credits; 720p costs twice that and 1080p five times, with duration scaling on top. Since whether a motion is going to look convincing is fully visible at the baseline — artefacts do not hide at low resolution, they are if anything more obvious — testing high is spending five times as much to learn the same thing. Work out at the baseline whether the motion holds, whether the identity stays consistent across the clip, and whether the duration is right. Only once you have a version that works is it worth re-rendering at a resolution the destination actually justifies, and for many short-form placements the baseline is already sufficient. The habit of testing cheap is what makes a below-100% hit rate affordable rather than painful.
- Can I animate a photo that is not of my own character?
- You can work from someone else's post as a reference, but the output should be your character rather than theirs — that is the whole basis of what this module teaches. The replication path handles exactly this: it takes the source's composition and movement and puts your established character into it, so what you publish is your own person in a proven format rather than a copy of someone else's footage. Image replication at 3 credits flat is the cheap way to test whether a composition suits your character before committing to video; faithful video replication then supplies the movement from the source clip itself, so you do not have to describe it. Working the other way round — animating a photograph of a real person who has not consented — is not something to do, regardless of the tooling involved.
Related reading
You have the content. Now place it.
Producing is only half of it — each platform rewards different cadences and formats. The platform playbooks module covers what changes where.
Start free