I’ve been creating narrated PowerPoint presentations for years.
If you’ve ever built training content, you know the drill. Write your speaker notes, record audio, attach it to slides, test animations, make edits, re-record audio, re-attach audio, test again.
It’s not difficult work. It’s repetitive work.
When AI voices started getting good, I thought the process would become dramatically easier. In some ways it did. Tools like ElevenLabs can produce incredibly natural narration in minutes.
The problem is everything that happens before and after the audio is created.
My original workflow looked something like this:
- Copy notes from PowerPoint
- Paste them into ElevenLabs
- Generate audio
- Download audio files
- Organize files
- Attach audio to slides
- Rebuild animations when something broke
- Repeat for every revision
The voice generation was fast. The production process wasn’t.
That’s ultimately why I built Voxsmith.
Instead of treating PowerPoint and ElevenLabs as separate tools, Voxsmith connects the two workflows. It reads the speaker notes directly from a PowerPoint presentation, generates narration using your ElevenLabs voice, then attaches the audio back to the appropriate slides while preserving animations and timing.
The result is a narrated PowerPoint presentation that’s ready for review, editing, or production.
For me, the biggest benefit isn’t saving a few minutes. It’s removing the repetitive work between writing and delivery.
I can focus on the content instead of managing hundreds of audio files.
If you’re already using ElevenLabs and PowerPoint together, Voxsmith was built specifically to eliminate the busywork between those two tools.
Want to see it in action? Watch the demo.