- Latch Infinity
- Education & EdTech
From SLM to Screen: Turning a Word Document into a Lecture Students Finish

Somewhere on a university server right now sits a Self Learning Material document. Forty pages. Times New Roman. A table of contents, five units, learning outcomes at the top, MCQs at the bottom, references formatted with real care.
It is academically sound. It has passed review. And someone has just been asked to “turn it into videos by the end of the month.”
This is where most e-learning production goes sideways - not because the source material is bad, but because a document and a lecture are fundamentally different animals wearing similar clothing. A document is scanned. A lecture is endured. Readers skip, re-read and control their own pace. Viewers get one pass, in one order, at a speed someone else chose.
Reading a document aloud over slides is not a lecture. It’s a hostage video with better typography.
Here’s the process that actually works.
Step one: audit the document as content, not as a script
Before a single slide is designed, read the SLM with one question running: what is the actual teaching here?
Most units contain three types of material, and they need three completely different treatments.
Conceptual content - definitions, frameworks, models, cause-and-effect. This is what video is genuinely brilliant at, because it can animate a process the page can only describe.
Procedural content - steps, calculations, methods. Video handles this well when demonstrated, badly when read.
Reference content - tables, extended lists, statutory provisions, bibliographies. Video is the worst possible medium for these. They belong in a downloadable, not on screen at 1080p where nobody can pause fast enough.
Sort the document into these three buckets first. Typically 50 to 60% of an SLM survives into video, 20% becomes an on-screen job aid, and the rest stays happily as a document. Trying to film 100% of it is the original sin of e-learning production.
Step two: restructure around attention, not around headings
The document’s chapter structure was built for a reader with a table of contents and a highlighter. Video needs a structure built for someone whose attention is under active attack from a phone in the same hand.
The unit of video is not the chapter. It’s the segment - six to ten minutes covering one idea, with a clear opening question and a clear resolution. A 40-page unit typically becomes five to eight segments, not one 55-minute recording.
This isn’t only about attention spans. Segmented content is searchable, re-watchable, and re-usable across programmes. When a syllabus changes, you re-record one eight-minute segment rather than an entire lecture.
Within each segment, restructure from heading order to question order. Academic documents typically move definition → explanation → example → application. Video works far better in reverse: a concrete situation, the problem it poses, then the concept that resolves it, then the formal definition. Same content, same rigour - just sequenced so the learner has a reason to want the definition by the time it arrives.
Step three: write a script, not a narration
The most common failure in the entire pipeline is handing a voice artist the document and pressing record.
Written academic prose has long sentences, subordinate clauses, passive constructions and nominalisations (“the implementation of the framework requires the consideration of…”). All of it is perfectly readable and completely unlistenable. The ear can’t hold a 40-word sentence with three dependent clauses. It gives up around word 25 and starts thinking about lunch.
Spoken script rules that make a measurable difference:
• Sentences under 20 words. Most under 15. • Active voice, almost always. • One idea per sentence. • Signposting the page doesn’t need: “There are three of these. Here’s the first.” • Contractions. People say “it’s.” Formality lives in your rigour, not your syllables. • Read every line aloud before it’s approved. If you run out of breath, so will the
presenter.
Critically, the script must preserve the academic substance exactly. Simplifying the language is required. Simplifying the content is a compliance problem, and in a regulated programme it’s the kind that surfaces during an audit.
Step four: design visuals that carry meaning, not decoration
Slides in most recorded lectures are the script, in bullet form, in a smaller font. This is the worst configuration available, because it forces learners to read and listen simultaneously - two channels competing for the same processing capacity. Learning outcomes suffer measurably.
The alternative: the visual carries what words can’t.
• Processes become animated flows that build step by step as the narration reaches them
• Comparisons become side-by-side structures, not bulleted paragraphs • Data becomes a chart with one highlighted takeaway • Abstractions become a concrete metaphor, drawn • Definitions get four words on screen, not the full sentence being spoken
A useful discipline: if a slide would still make sense with the audio muted, it’s probably a document page in disguise. Slides should be incomplete without the narration, and the narration incomplete without the slides.
Step five: production choices that survive scale
Most institutions need dozens of hours, not one showpiece. Production design should optimise for repeatability.
Presenter on camera or voice-only? On-camera presence improves engagement and trust, particularly in the opening and closing of each segment. A workable hybrid: presenter on camera for the first 45 seconds and the last 30, voice-over graphics in between. You get the human connection without needing a faculty member to be film-ready for eight straight hours.
One consistent set. A single lighting and camera setup, locked and documented, means episode 40 matches episode 1. Photograph the setup. Write down the settings. Future you will be extremely grateful.
Teleprompter, without the stare. Faculty reading from a prompter often develop the distinctive glassy gaze of a hostage. Position it just under the lens, rehearse once, and let them paraphrase. Slight imperfection reads as authority. Perfect delivery reads as a recording.
Record audio properly. Viewers forgive mediocre picture and abandon bad sound within seconds. A lavalier and a treated room outperform an expensive camera in a reverberant hall, every time.
Step six: build for the LMS from day one
“LMS-ready” is not a file format - it’s a workflow. Decide before production: the video platform and its encoding requirements, the SCORM or xAPI packaging standard, caption and transcript requirements for accessibility, chapter markers, the assessment items that
sit alongside each segment, and the file naming convention that maps every asset back to its unit and outcome.
Decide this after production and you’ll spend as long repackaging as you spent shooting. Ask anyone who has renamed 340 files by hand.
Step seven: quality-check against the source
The final pass is the one people skip. Put the original SLM next to the finished video set and verify: is every learning outcome addressed? Is every mandated topic covered? Are definitions verbatim where regulation requires it? Do the assessment items still map to what was taught?
The video is a derivative of a compliance-approved document. If it drifts, the compliance approval doesn’t travel with it - and in a regulated programme, that’s not a content problem. That’s an audit finding with your name on it.
Accessibility and language: decide early or pay twice
Two requirements get treated as post-production extras and are, in practice, structural decisions that need making before the first script is written.
Accessibility. Captions and transcripts are mandatory in most regulated programmes and non-negotiable in good practice. But real accessibility goes further than a subtitle file. Any information conveyed purely visually - a chart, a diagram, a colour-coded distinction - must also exist in the narration or a text alternative, or a learner using a screen reader receives an incomplete unit. This is a scripting decision. If the voiceover says “as you can see here,” the visual has just become load-bearing and inaccessible in the same breath. Write narration that stands alone, and the accessibility work largely takes care of itself.
Language. If the programme will be delivered in more than one language, the entire production changes shape. On-screen text must live in editable graphics layers rather than baked into the video. Narration should be recorded to a locked picture so alternate voice tracks can be swapped without re-editing. Idioms and culturally specific examples need flagging at script stage, because they don’t translate - they need replacing, and that’s an academic decision, not a translator’s.
Retrofitting either of these into a finished library is roughly as expensive as producing it again. Deciding both at the brief stage costs a conversation.

