Make Your Character Talk

By the OFMAI Team · Updated July 2026

Among all the short-form formats worth learning, the one where your character looks into the lens and says something is the most reliably useful — not because it gets the most views, but because it gets a disproportionate number of comments, and comments are the engagement signal that is hardest to fake and most expensive for a platform to ignore.

The source material this lesson draws on makes the case with numbers rather than assertion. The instructor points to a talking post with 72 comments sitting on an account whose other recent posts had three, two, five, twelve, five and fifteen. That is the neighbour test from the first lesson, applied to a format rather than a single post — and the gap is large enough that it is not noise.

Why the format works

Two things are happening at once, and the format only works when both are present.

The first is that a person addressing the camera directly is harder to scroll past than a person doing something in the third person. Direct address reads as being spoken to, and that recruits attention in the first half-second — which is the only window that matters, since the platform's decision about whether to widen distribution is made largely on whether people stayed.

The second is that the script asks for something. A statement gets watched; a question gets answered. The specific pattern the source demonstrates is a direct question paired with a small offer for responding — something along the lines of asking people to comment a particular word in exchange for something. That combination converts passive watchers into commenters at a much higher rate than the video alone would, because it removes the two frictions that stop people commenting: not knowing what to say, and not seeing a reason to bother.

Keep it short. The reference clip in the source material is nine seconds, and the generated voiceover for it came out at eight — that is the right neighbourhood. A talking clip is a delivery mechanism for one sentence, and the temptation to add a second one is the main way the format gets ruined.

One honest caveat: the source presents its claims about the algorithm's response as fact, and we have not independently verified them. What we can verify is the comment gap on the account shown, which is real and is the reason to try the format. Treat the mechanism as plausible and the result as measurable — then measure it on your own account.

What the traditional production chain looks like

The source lesson walks through producing one of these end to end, and the workflow is worth laying out precisely, because it is representative of how this kind of work is normally done and because the step count is the entire argument.

The creator finds a reference post and downloads it. The spoken line is extracted using a transcription service — an account, pasting the post's address, copying the resulting text out. The line then goes into a voice service, where a voice is selected, the text pasted in, an audio file generated and downloaded to disk. Separately, the opening frame of the reference is screenshotted and taken to an image generator, where character reference photos and that screenshot are uploaded together with a prompt, an image is generated, and it is downloaded. The audio file and the image are then both uploaded to a lip-sync engine, a resolution is chosen, a prompt is written, the job is run, and the resulting video is downloaded. That file goes into a video editor to fix a visual glitch at the frame's edge and is exported. Finally, because all of this happened on a desktop and the post has to go up from a phone, the file is moved across using a file-transfer service before the platform's own composer adds text overlays and publishes.

That is five separate services plus a file-transfer step. Four of them need their own account and their own balance. And between them sit six manual download-and-re-upload handoffs, each of which is a place to grab the wrong file, lose a version, or discover at the end that the audio and the image do not match.

None of this is a criticism of the approach — it is how the work gets done when the pieces live in different places, and the creator demonstrating it produced a perfectly good result. The point is only that the step count is what it is.

The same output in one place

OFMAI collapses the middle of that chain, because the pieces it needs already exist on your character. The reference photos that define your character's appearance are attached to it. So is the cloned voice, if you have set one up. Neither needs to be fetched, uploaded or matched to anything — they are what the character is.

On the generation screen, a video can carry a spoken dialogue field and a character-voice toggle. You write the line, switch the voice on, and the clip is produced with your character delivering it in her own voice. There is no separate audio file to generate, download and re-upload, no base image to produce in one place and carry to another, and no moment where the audio and the visual are two files you have to keep aligned by hand.

The video replication path covered in the previous lesson handles the same job from the other direction. If your reference is an existing talking post, adapt mode reads the delivery out of the source and rebuilds it with your character — and because the dialogue passes through a written stage, you can have it come out in a different language. The character-voice toggle applies there too.

Two things stay outside the tool, and it is worth being straight about them. Transcribing a reference clip's spoken line is not something OFMAI does — you type the line yourself, which for a nine-second clip is a matter of seconds and removes an account and a step rather than adding one. And the on-screen text overlay that makes the clip watchable with sound off is added in the platform's own composer at posting time, exactly as the source describes, because that is where it belongs.

Writing the line

Production is the easy half. The line is what determines whether the format does anything, and most attempts fail here rather than in the render.

Keep it to one sentence and under about ten seconds spoken. Say it out loud with a timer before you commit — text that looks short reads long, and a line that overruns forces the clip to keep going past the point where the viewer has decided.

Make the ask explicit and trivially easy. "Comment a single word" outperforms "let me know what you think", because the second one requires the viewer to compose something. The lower the effort you request, the more people clear the bar.

Give a reason. The source's pattern pairs the request with a small offer for responding, and that pairing is doing real work — it converts the ask from a favour into an exchange. Whatever you offer, make sure you can actually deliver it, because a promise you do not honour costs you more than the comments were worth.

And write it in the language your audience speaks, not the language your reference was in. If your reference clip is in a language your followers do not use, adapt mode's language selection solves that directly, and it is one of the strongest reasons to reach for that mode over faithful.

What it costs

Generating a video with dialogue is priced by resolution and duration. The baseline — 480p at five seconds — is 15 credits. Higher resolutions scale from there: 720p costs twice the baseline and 1080p five times, with duration scaling linearly on top. Switching on the character voice at generation time adds 2 credits.

If you are instead rebuilding an existing talking post through replication, the costs are the ones from the previous lesson: 21 credits per five seconds for adapt mode, plus 2 for the character voice.

For a format that lives or dies on whether the line lands rather than on pixel count, the baseline resolution is usually the right starting point — the source material makes the same call and sticks with the lower setting deliberately. Get the script working first, then decide whether resolution is what is holding the post back. It usually is not.

Produce a talking clip

  1. Find the format, not just the postApply the neighbour test to comments rather than views. A talking post pulling many times its account's usual comment count is the signal you want.
  2. Write one sentenceA direct question with an explicit, low-effort ask and a reason to respond. Say it out loud with a timer — under ten seconds.
  3. Confirm the character has a voiceThe character-voice option needs a cloned voice attached to the character. Set that up first if it is not already there.
  4. Fill the dialogue fieldOn the generation form, enter the line in the spoken dialogue field and switch the character voice on.
  5. Keep the resolution modestStart at the baseline. The format depends on the script, not the pixel count, and higher settings multiply the cost.
  6. Generate and reviewWatch it once with sound and once without. If it does not hold you for the full duration, the line is too long.
  7. Add the on-screen text at posting timeUse the platform's own composer to time the text overlay to the audio, so the clip works with the sound off. Then publish.

The step count, honestly

The same nine-second talking clip, produced two ways. Tool categories are named rather than vendors; the counts come from the source lesson's own walkthrough.

The multi-tool chainIn one place
Separate services involved5 — a transcription service, a voice service, an image generator, a lip-sync engine, a video editor1
Accounts and balances to maintain41
File download / re-upload handoffs60
Extra file-transfer step to reach a phoneYesYes — unavoidable either way
Getting the voiceGenerate elsewhere, download, re-uploadAlready attached to the character
Getting the base imageGenerate elsewhere, download, re-uploadAlready attached to the character
Aligning audio to visualManual — two files kept in sync by handHandled in one job
Transcribing the reference lineAutomated by a serviceTyped by hand — seconds for one sentence
On-screen text overlayPlatform composer at posting timePlatform composer at posting time

Frequently asked questions

Do I need a cloned voice to use this format?
For the character-voice option, yes — that toggle requires a voice already attached to the character, and it will not let you submit without one. It is worth setting up if you intend to produce talking clips regularly, because a voice that stays the same across every post is part of what makes a character read as a person rather than a series of unrelated renders. Viewers notice inconsistency in voice faster than they notice it in appearance. That said, you can generate a video with dialogue without the character-voice option switched on, so the format is not gated behind it. If you are testing whether the talking format suits your account at all, run a few without it first and add the voice once you know the format earns its place. Adding a voice later does not require redoing anything else on the character.
How long should the spoken line be?
Under ten seconds, and ideally closer to eight. The reference clip in the source material runs nine seconds and the generated voiceover for it came out at eight, which is a good target to aim at. The constraint is not arbitrary: a talking clip has one job, which is to deliver a single ask, and every second past that point is a second in which the viewer can decide to leave before reaching it. Long lines also compound a second problem — the longer the delivery, the more chances there are for a small artefact in the render to become noticeable. Write the sentence, read it aloud against a timer, and if it overruns, cut words rather than speeding up the delivery. The most common failure in this format is a script that tries to say two things.
Can my character speak a language other than the reference clip's?
Yes, through the adapt mode of video replication. Because adapt passes the source video through a written description stage rather than transferring its motion directly, the dialogue is reconstructed rather than copied, and you can nominate the output language — around thirty are available, with the option to keep the original. This matters more than it first appears: the strongest talking formats often surface first in one language and reach audiences in others long after, so being able to bring a proven format into your audience's language is frequently the whole opportunity. If you are writing your own line rather than rebuilding someone else's clip, the question does not arise — you simply write it in whichever language you want and generate from the dialogue field directly.
Why does the on-screen text still have to be added at posting time?
Because that is where it works best, not because it is a gap. A large share of short-form viewing happens with the sound off, so the text overlay is what makes the clip legible to those viewers — and platforms treat text added in their own composer as native, with their own fonts, timing controls and placement. Adding it there also lets you adjust the timing against the audio while previewing the actual post, which is the only place you can see exactly what the viewer will see. The source lesson does the same thing for the same reason, typing and timing the overlay in the platform's composer immediately before publishing. It is a minute of work at the end of an otherwise automated pipeline, and it is the one step where doing it by hand is genuinely better than automating it.

Related reading

Not every post needs speech

A still photo you already have can become a short clip with very little work — and the craft rule for doing it well is counter-intuitive. Continue to the last lesson in this module.

Start free
    Make Your Character Talk (And Why It Pulls Comments) | OFMAI Academy