Voxsmith started as a PowerPoint tool. Speaker notes in, narrated audio out, attached right back onto the slides. That’s still the core of it.
But a lot of what instructional designers and eLearning developers write isn’t a deck. It’s a script. A facilitator guide. A policy document that needs to become an audio walkthrough. If your content lives in Word, you don’t need PowerPoint just to get ElevenLabs narration out of it — Voxsmith handles .docx files directly, on both Windows and Mac.
Here’s how it works.
The short version
- Write your Word document with
[Chapter Title]markers to break it into sections - Point Voxsmith at the
.docxfile - Pick a voice and generate
- Get one
.wavfile per chapter, ready to use
No COM automation, no slides, no notes pane. Just text in, audio out.
Why chapters instead of one giant file
If Voxsmith generated a single audio file for an entire document, you’d have no way to swap out a section, regenerate just the part you edited, or drop pieces into different parts of a course. Chapters solve that. Each one becomes its own file, so you can mix and match in your LMS, video editor, or wherever the audio ends up.
The chapter marker is simple: a line by itself, in square brackets.
[Introduction]
Welcome to the course. In this module, you'll learn the fundamentals of...
[Module One: Getting Started]
Let's start with the basics. Before we go further...
Each bracketed line starts a new chapter. Everything underneath it, until the next bracket, becomes that chapter’s narration text. Anything written before your very first [Chapter Title] marker — a working title, a note to yourself, a draft heading — gets ignored, so feel free to leave scratch notes at the top of the doc without worrying about them ending up in the audio.
A couple of things worth knowing:
- Markdown-style links and footnotes won’t trip it up.
[text](url),[^1], and reference-style links like[text][ref]all look like they could be a chapter marker, but Voxsmith only treats a bracketed line as a marker if it’s alone on its own line with nothing else around it. - Empty chapters get skipped. If you add a marker and don’t write anything underneath it before the next one, that chapter just won’t generate — no empty audio file, no error.
- The title is for your reference, not the filename. Output files are named by sequence (
chapter_01.wav,chapter_02.wav, and so on), so name your chapters whatever’s useful to you while you’re writing.
Setting it up
Open Voxsmith and browse to your .docx file the same way you’d browse to a PowerPoint deck. Voxsmith reads the file extension and automatically routes it down the document pipeline instead of the slide pipeline — there’s nothing to toggle.
Once it’s loaded, the top bar will show you a chapter count instead of a slide count, so you can confirm it parsed the way you expected before you spend any API credits.
From there, it’s the same generation flow you already know:
- Pick your voice
- Adjust speed, stability, and similarity if you want to tune the read
- Set your output folder (or let Voxsmith default to a new folder next to your source file)
- Generate
When it’s done, you’ll have one .wav per chapter sitting in that output folder.
Where this is actually useful
A few places this comes up more than you’d expect:
Facilitator and SME scripts. If a subject matter expert wrote their talking points in Word rather than building slides, you can narrate their script directly without ever touching PowerPoint.
Course transcripts that need to become audio. Written training content that’s getting repurposed into an audio or video format — no slide deck involved at all.
Policy and compliance documents. Long-form text that needs an audio option for accessibility, without anyone building a 40-slide deck just to have something to attach narration to.
Drafting and iteration. Sometimes it’s just faster to write and edit a script in Word than to manage it across dozens of speaker-notes fields in a deck. Write it as one document, mark your chapters, and let Voxsmith split the audio for you.
A note on what this isn’t
Voxsmith’s Word support is read-only, by design. It extracts paragraph text and generates narration from it — it doesn’t write anything back into the document, and there’s no equivalent of the PowerPoint auto-attach step. You’ll always get standalone audio files, which you then bring into whatever tool builds your final course, video, or podcast episode.
If you’re working from .txt or .md instead of .docx, the exact same chapter-marker convention applies. Pick whichever format is easiest for you to write in — Voxsmith treats them identically once it gets to the chapter-parsing stage.
Try it
If you’ve already got Voxsmith, this works today — no update needed, no setting to find. Just open a Word document with a couple of [Chapter Title] markers in it and see what comes out.
Haven’t picked it up yet? Voxsmith is a one-time $89 purchase that covers both the Windows and Mac builds under a single license.