The voiceover on our launch and product videos comes out of ElevenLabs. When a script needs narration and we are not booking a voice actor, we generate it here — fast enough to fix a line without re-recording, clean enough to sit under motion graphics. This is the beginner setup: the account, picking a voice, and the two settings that actually change how it sounds.
It is the narration layer in the stack we run on, downstream of the video research we do with vidIQ and upstream of the page we ship the finished video on.
What ElevenLabs is
Text-to-speech that does not sound like text-to-speech. You paste a script, pick a voice, and it renders an audio file. There is a library of ready-made voices and the option to design or clone one, but for launch videos the stock library is more than enough. The output is a clean WAV or MP3 you drop straight into the edit. You never touch a microphone, a booth, or a schedule.
Setup, start to finish
- Make an account. The free tier gives you a monthly character quota, enough to test voices and narrate a short video or two.
- Open the Text to Speech editor. Paste your script into the box.
- Pick a voice from the library. Preview a few. The right voice for a product launch is usually calmer and lower-energy than the one you first reach for.
- Generate, listen, and download.
The free tier includes a monthly character quota you can test all of this with.
Visit ElevenLabsThe settings that actually matter
There are a handful of sliders. Two of them do most of the work.
Stability
Stability controls how consistent the delivery is from render to render. Low stability is more expressive and more variable — good for character work, risky for narration because a re-render can shift the read. High stability is steadier and flatter. For product voiceover we run it fairly high so the tone stays even under graphics and so fixing one line does not change the feel of the rest.
Similarity
Similarity is how closely it tracks the original voice's character. Push it up and you get more of that specific voice; too high and artifacts creep in. We keep it high but back off the moment we hear any digital edge.
The rest
Style exaggeration and speaker boost are situational. Leave them low for narration. The biggest quality lever is not a slider at all — it is the script. Punctuation is direction. A period is a beat; a comma is a shorter one. Write the way you want it read and the model follows.
The real workflow
Here is where it earns its slot. We paste the near-final script, pick the voice, set stability high, and render. Then we drop the audio under the motion graphics and watch the timing. A line runs too long? We do not re-record a human and re-book a session — we edit the text and re-render that line in seconds. That iteration speed is the actual reason it is in the pipeline. It turns voiceover from a scheduling problem into a text edit.
One caveat we hold to: it is one voice option, not the whole sound design. Music, pacing, and the mix still matter, and a bad script read cleanly is still a bad script. ElevenLabs handles the narration; it does not handle the taste. We still write the words, cut the timing, and sit the voice in the mix by hand.
Pricing and licensing, an honest take
Billing is by characters — how much text you turn into speech per month — plus the plan tier. The free tier is enough to evaluate it and narrate short pieces. Paid tiers raise the character quota and, importantly, sort out commercial rights: if you are putting the audio in client work or ads, you want a plan that grants commercial use and clarifies ownership. Read that part before you ship, not after.
Honest take: the common complaints are real and worth knowing. It can mispronounce names and unusual terms, so proofread by ear. Very long or very expressive scripts are where the seams show. And character-based pricing means a talky, long-form channel burns quota faster than a studio doing short launch videos. For our use — tight, narrated product videos — it is well inside the value. Check the current pricing and license terms on their site.
Verdict
ElevenLabs is the fastest way we know to get clean narration under a video without booking a session. Pick a calm voice, run stability high, treat punctuation as direction, and re-render lines as text edits. Keep it as one instrument in the mix, not the whole thing. When the video is cut, it goes live on a page we ship on Netlify — the same pipeline, end to end.
Generate your first voiceover free and hear the settings for yourself.
Visit ElevenLabsFrequently asked questions
Is ElevenLabs free?
There is a free tier with a monthly character quota. It is enough to test voices and narrate a short video. You upgrade when you need more characters or commercial rights.
Can I use ElevenLabs voices commercially?
On the right plan, yes. Paid tiers grant commercial use and clarify ownership of the audio you generate. If the voiceover is going into client work or ads, confirm the license on your plan before you publish.
What do people actually complain about with ElevenLabs?
The honest ones: occasional mispronunciation of names and rare words, seams on very long or very expressive scripts, and character-based pricing that adds up for long-form use. None of it is a dealbreaker for short narrated videos, but proofread by ear.
Does ElevenLabs sound robotic?
Far less than older text-to-speech, but delivery depends on the voice and settings. High stability and clean punctuation get you an even, natural read. Chasing maximum expressiveness is where it can wobble.