← THE VIDMOAT JOURNAL

Editing Course Videos for People Who Are Trying to Do Something

ONLINE COURSES12 MIN READ

Here is the thing that most course editing advice misses. A person watching a tutorial is not an audience. They are a participant. Their hands are on a keyboard, their eyes flick between your video and their own screen, and they are pausing every forty seconds to try what you just did.

That changes the job completely. Entertainment editing is about keeping someone in a chair. Course editing is about not wasting the time of someone who is already committed. They are not going to leave because they got bored. They are going to leave because they got lost, or because you made them sit through eleven seconds of you hunting for a menu.

So the priorities invert. Pacing matters, but not the kind of pacing a vlog needs. Retention matters, but the metric is whether they finish the module and can do the thing. Personality matters least of all, which is uncomfortable news for anyone who has spent money on a good camera.

What follows is the craft, in rough order of how much it improves a course: cutting dead air, structuring the first fifteen seconds, deciding when to zoom, building the repeatable furniture of a series, chapters, and captions. Then the parts a tool can do for you.

01

Dead air is the whole job

If you fix nothing else, fix this. In an unedited screen recording, somewhere between fifteen and thirty per cent of the runtime is nothing happening. You are looking for a menu item. You are waiting for a build. You are saying so, and, um, let me just, while your cursor circles the top of the screen. None of it teaches anything, and all of it costs the student real minutes.

There are three separate species of dead air and they get cut differently.

The first is silence, the gaps where you are not speaking and the screen is not changing. Cut all of it down to a beat. Not to zero, because speech with every pause removed sounds unhinged, but to roughly a quarter of a second between sentences and half a second between topics. Automatic silence trimming handles this: the clip is replaced by a set of clips covering only its non-silent portions, effects and keyframes intact, and you are left with the same lesson three minutes shorter.

The second is filler inside sentences. The ums, the false starts, the sentence you began twice. These are not silence, so a silence pass will not touch them. The way to remove them at scale is through the transcript: transcribe the lesson, find the word timings for every filler, and cut those exact ranges out of the timeline. In Vidmoat that is a ripple cut by default, so the surviving pieces compact and everything after them pulls left. On a forty minute lesson this is the difference between an evening of scrubbing and a couple of minutes.

The third is the worst one: waiting. A build that takes twenty seconds, a page that loads slowly, a form that submits. Do not cut it out entirely, because the student needs to know the wait exists. Speed it up instead, to something like eight or ten times, and let the on screen progress visibly move. Ten seconds of real time becomes one second of visible progress and the student learns that the step takes a while without paying for it.

  1. 01Run a silence pass over the raw recording first, before you do anything creative.
  2. 02Transcribe, read the transcript, and cut filler words and false starts by their word timings.
  3. 03Find every wait and ramp it up rather than removing it.
  4. 04Watch the result at normal speed once. Anywhere you feel the urge to skip forward, cut.
02

The first fifteen seconds belong to the screen, not to you

A talking head opening exists for one reason: to establish that a person is here and what this lesson covers. That takes eight to twelve seconds. Anything longer is you enjoying yourself.

What it should contain is one sentence naming the specific thing they will be able to do at the end, and nothing else. Not a welcome back to the channel. Not a recap of the previous lesson, which they either just watched or deliberately skipped. Not a request to like anything. A student who has paid for a course does not need to be sold the course again at the top of every video.

Then get to the screen. The transition from face to screen is the single most important cut in a lesson, and the smoothest version is not a cut at all: keep talking over the top of the transition so your voice carries continuously while the picture changes. The student never experiences a gap.

There is a useful middle state worth using more than most people do. Rather than face, then screen, go face, then screen with your webcam small in a corner, then screen alone once the work gets detailed. Your presence fades out as the difficulty fades in. When you come back to talk about what just happened, the webcam can grow again.

Vidmoat has named layouts for exactly this shape, so framing a clip as a stacked screen and camera arrangement, or into a rounded card in a corner, is one instruction rather than a manual crop and position on every clip. The value is less the time saved and more that lesson nine is framed identically to lesson two.

03

Zoom when the detail is small, and only then

The default for a screen recording should be the whole screen. The student is trying to build a mental map of an interface, and they cannot do that if you keep showing them fragments of it without context.

Zoom in for three situations. When you are typing into a specific field and the text would be unreadable at full screen size on a phone. When you are pointing at something small, a checkbox, a setting, an error message in a log. And when there is a moment of change worth emphasising, a value updating or a state flipping.

Zoom out again immediately afterwards. The mistake is staying zoomed, so that three steps later the student has no idea where on the page you are. A zoom should be a sentence long, not a paragraph.

How you get there matters. A hard cut to a zoomed view is disorienting because the student has to re-locate themselves. A move that takes about half a second and eases out at the end reads as the camera leaning in, and the eye follows it without effort. Same when you pull back.

The related question is what size you recorded at. A 1080p screen recording zoomed to two hundred per cent is a soft mess. Record at the highest resolution your machine will give you, and increase the font size in your editor or terminal before you hit record. Bumping an editor to sixteen or eighteen point costs you nothing and removes most of the reasons you would have needed to zoom in the first place.

If you want the emphasis to follow something that moves, a label or a highlight box can be stuck to a tracked region of the clip so it travels with a menu, a cursor path or a piece of UI as it moves. Used once per lesson to call out the thing being clicked, it beats an arrow you have to keyframe by hand.

04

The furniture that makes twelve videos feel like one course

This is the part that separates a course from a playlist, and it is almost entirely mechanical.

Every lesson opens with the same card in the same position with the same animation, showing the module number, the lesson number and the lesson title. Every lesson closes with the same card, showing what comes next by name. The typeface, the colour, the timing and the placement do not vary. Ever.

The reason is not aesthetics. It is that a repeated frame becomes a signal. After three lessons the student recognises the intro card without reading it and knows a new unit has started. When they come back a week later, the closing card tells them exactly where they stopped. Consistency is doing navigational work.

Keep the cards short. Two seconds for the intro card, three for the outro. A ten second animated logo sequence at the top of every lesson in a thirty lesson course costs your students five minutes of their lives collectively, and it is the first thing they learn to skip, which trains them to skip the beginning of your videos in general.

These cards are the obvious use for custom HTML elements. You describe the card once, including the layout, the weights, the background and the animation, and it becomes a real clip on the timeline. For the next lesson it is the same element with two fields changed. That is why a series edited this way looks designed rather than assembled, without you designing anything twice.

  1. 01Build one intro card and one outro card before you edit lesson one.
  2. 02Fix their duration, position and animation and write those numbers down.
  3. 03Reuse them unchanged for every lesson in the series, changing only the text.
  4. 04Keep them under two and three seconds respectively.
05

Chapters, because nobody watches a lesson once

Course videos have an unusual viewing pattern. The first watch is linear. Every watch after that is a search. Someone is stuck on step four, and they are dragging the scrubber back and forth trying to find the forty seconds where you did the thing they cannot get to work.

Chapters fix that, and they are cheap. Break the lesson at the points where the task changes, not on a timer. A typical fifteen minute lesson has four to seven of these. Name them by the action, not by the concept: Installing the CLI beats Setup, Fixing the CORS error beats Troubleshooting. The student searching is thinking in verbs.

Write the chapter list before you edit, if you can. If you know a lesson is five chapters long, you will naturally record it as five segments, and the edit becomes assembling five clean takes rather than salvaging one long one. This is the highest leverage change available to most course creators and it happens before the camera is on.

Chapter segments detected from the content can be persisted on the project itself, which means the same breakdown feeds your video description, your course platform's lesson outline and your own edit structure without you writing it out three times.

One more thing the chapter structure gives you: a rerecording unit. When a tool changes its interface eight months from now, you rerecord chapter three, not the whole lesson. Courses that are structured this way stay current. Courses recorded as one continuous forty minute take rot.

06

Captions, and the student on the 7:42 train

A meaningful share of course watching happens somewhere the sound is inconvenient. On a commute, in an open office, next to a sleeping child, in a second language. Captions are not an accessibility checkbox on a tutorial, they are a core delivery format.

Auto-generated captions are close but not close enough on technical content, and the failure mode is specific: proper nouns and jargon. The transcription will render your library name as three ordinary English words. Every time. Budget five minutes per lesson to fix the terms that matter, and keep a list of the ones that recur so you fix them faster each time.

Style them conservatively. Course captions want one or two lines, centred low, high contrast, a plain sans face, and nothing else. Word by word karaoke highlighting is available and looks great on a short vertical clip, but on a twenty minute lesson the constant motion at the bottom of the frame competes with the screen recording you are asking people to read. Save it for the promotional cut.

Watch the collision between captions and interface. Screen recordings often have important content near the bottom of the frame: a terminal, a status bar, a form footer. If your captions sit on top of it, either raise the captions or frame the recording so the bottom strip stays clear. There is a layout for this that lifts the clip and keeps the lower band free, and using it from lesson one is easier than discovering the problem at lesson fourteen.

The transcript is worth more than the captions, incidentally. It is the searchable text of your course, the raw material for lesson notes, and the thing that let you cut every filler word in section one. Generate it early and it earns its keep three times.

07

A working order of operations

The sequence matters more than it looks like it should, because several of these steps are much cheaper when they come after the others.

Transcribe first, before any cutting, so you have word timings for the whole recording. Then trim silence, then cut filler by transcript range. Doing the cuts before the transcription means transcribing again afterwards for captions, which is wasted effort and wasted time.

Do the structural edit next: pick your chapter boundaries, cut the segments, and reorder if a step makes more sense earlier. Only then add the visual layer, the zooms, the layouts, the intro and outro cards. Adding graphics before the structure settles guarantees you will move all of them later.

Captions come last, over the final cut, because captions timed to a version you subsequently trim are captions timed to nothing.

Then watch it once, all the way through, at normal speed, without touching anything. Not to admire it. To notice the two places where you still get bored, and cut those.

  1. 01Transcribe the raw recording.
  2. 02Trim silence automatically.
  3. 03Cut filler words and false starts by transcript ranges.
  4. 04Speed up every wait rather than deleting it.
  5. 05Set chapter boundaries and cut the structural segments.
  6. 06Add layouts, zooms, and the standard intro and outro cards.
  7. 07Add and correct captions over the finished cut.
  8. 08Watch it end to end once and cut whatever bores you.
TRY IT YOURSELF
Try it on your own footage

Vidmoat is free to start — connect your agent, or just use the editor.

Start editing free →
Q&A

Frequently asked

How long should a single course lesson be?

Between six and fifteen minutes for most technical material. Under six and you are usually splitting one idea across two videos; over twenty and the student cannot find their place when they come back. If a lesson wants to be thirty minutes, that is usually a sign it is three lessons wearing a coat.

Should the webcam be on screen the whole time?

No. Bring it in for the opening, shrink it to a corner once you are working, and remove it entirely during detailed steps where the student needs the full screen. Bring it back when you are explaining rather than demonstrating. A permanent floating head over a code editor covers something important eventually.

Is it worth removing every um?

Not every one, but most. A few natural hesitations keep you sounding human. What actually hurts is the cluster: three fillers in one sentence while you think. Cutting by transcript ranges makes this quick enough that you can afford to be selective rather than either exhausting yourself or leaving them all in.

How do I stop a screen recording looking soft after zooming?

Record at the highest resolution you can and raise your font sizes before recording rather than zooming afterwards. An editor at sixteen or eighteen point read at full screen is sharper than a twelve point one blown up two hundred per cent, and it removes the reason for most zooms.

Do chapters actually help, or are they busywork?

They help the second and third viewing, which for a course is most of the viewing. A student stuck on one step needs to find that step in ninety seconds. Naming chapters after the action rather than the concept is what makes them findable.

Can I fix bad recording audio in the edit?

Only partially, and it is worth being honest about that. Noise reduction here is a high-pass filter that removes low rumble, and level normalisation is peak based rather than loudness based. Both help a decent recording sound tidier. Neither rescues a laptop microphone in a room with an air conditioner. Spend the money on the microphone, not the camera.

TOOLS MENTIONED
Video to Text TranscriptionSilence RemoverFiller Word RemoverAI Auto Captions
MORE FROM THE JOURNAL
Your AI agent is editing video blindWho the modern creator actually isHow to edit video with Claude using MCPEdit video from Telegram: send a clip, get it back cut