A vertical video can have a useful idea and still feel hard to watch. The opening takes too long. The captions cover the demonstration. The edit moves quickly, but nothing important happens.
When I edit, I treat captions and pacing as one problem: how much work am I asking the viewer to do?
My goal is not to make every second louder or faster. It is to make the next second worth watching. Here is the editing process I use for short explainers, product demonstrations and founder videos.
Start with the moment that matters
Before changing captions, I check whether the video starts in the right place.
Imagine a coffee shop filming a 30-second tutorial about bitter espresso. An opening like “Hi everyone, today we wanted to share a few tips” uses time without giving the viewer a reason to stay.
I would start with: “Bitter espresso? Try a slightly coarser grind.” Then I would show the adjustment.
That opening identifies the problem and begins answering it. It does not withhold the useful part behind a promise.
I use a simple test: if I remove the first sentence, does the video lose meaning? If not, I cut it. This is especially useful with generated scripts, which can add introductions that sound polished but contribute little.
Make captions easy to read once
Captions help people follow speech without relying entirely on audio. They also compete with everything else on screen.
I start with accurate transcription. Product names, prices and technical terms deserve a manual check. A caption that turns “15 grams” into “50 grams” changes the instruction, not just the spelling.
Next, I break the text into meaningful phrases. For the espresso example, I might display “Bitter espresso?” followed by “Try a slightly coarser grind.” I would avoid splitting “coarser” from “grind” just to make every caption the same length.
I usually start with one or two lines and roughly three to seven words per caption. That is a working limit, not a rule. Longer words, smaller screens and a busy background may require shorter chunks.
I choose a readable typeface, strong contrast and a consistent position. Then I preview on a phone. If I have to pause to read, I change the timing or reduce the competing visual information.
Keep captions away from the evidence
In a demonstration, the viewer needs to see the action. Captions should not cover the grinder setting, the software button or the finished result.
I leave room around important details and check the video against the intended platform's interface. Buttons, descriptions and other overlays can obscure text near the edges. There is no single placement that works perfectly everywhere.
If the action sits in the lower half of the frame, I move captions higher for that section. I keep those changes deliberate rather than letting text bounce around with every sentence.
I also separate spoken captions from editorial labels. A caption transcribes “Move the setting one step coarser.” A label might identify the adjustment dial. If both are needed, I avoid making them compete for the same space.
Edit for progress, not constant movement
Fast cutting is not the same as good pacing. Five cuts that repeat the same point can feel slower than one clear shot showing a useful action.
I look for changes in information. The viewer sees the problem, understands the adjustment and watches the result. Each shot earns its place by moving that sequence forward.
For a hypothetical 25-second coffee tutorial, I might sketch this structure:
- Seconds 0 to 3: show the espresso and name the problem.
- Seconds 3 to 8: show the grinder adjustment.
- Seconds 8 to 18: explain the next step with relevant footage.
- Seconds 18 to 25: explain what to check when tasting again.
Those timings are an editing draft, not a retention formula. If the adjustment needs another two seconds to be understood, I give it that time and cut elsewhere.
I remove accidental pauses, repeated phrases and unnecessary setup. I keep pauses that let someone inspect a detail or absorb an instruction.
Check the video with sound off and on
I watch once with sound off to check whether the captions and visuals carry the core explanation. Then I listen with sound on to check whether the cuts interrupt natural speech.
If every spoken pause has vanished, the result can feel breathless. If captions reveal the next sentence too early, they can pull attention away from the current demonstration.
At Filmotion, I think of automation as a first pass, not the final judgement. Generated captions and suggested cuts still need checking against the actual footage.
Use retention to ask a better question
After publishing, I look at the retention curve where the platform provides one. I compare videos of similar length and format rather than treating every completion rate as directly comparable.
A drop near a caption-heavy section gives me a question, not a verdict. Was the text difficult to read? Did the explanation repeat itself? Did the opening attract people expecting something different?
For the next comparable video, I change one meaningful element, such as shortening the opening or simplifying caption chunks. Audience differences still make that comparison imperfect.
I cannot edit my way to guaranteed retention. I can remove avoidable friction. Clear captions, visible evidence and enough time to understand each step are where I start.
Filmotion