Chapter II

How to cast and direct an AI narrator

Single voice or full cast, cloning your own voice, directing tone and pacing, and how much music a book can carry.

Introduction

Casting and direction matter more than the engine. Pick the wrong voice or leave tone vague, and no amount of production fixes it later.

This chapter covers style, casting, cloning your own voice if you want it, directing tone and pacing, dialogue with multiple characters, music and sound design, and how to review in the delivery room.

If you haven't heard a finished sample yet, start with a free first chapter. Turnaround is about 24 hours.

Choose a narration style

Decide single narrator or multi-voice before you cast. Single narrator is simpler and fits most fiction and nonfiction. Multi-voice works when distinct characters carry long stretches of dialogue and you want each speaker to sound like a different person.

Within that, pick a register: documentary and calm for memoir, history, and how-to; more dramatic for thriller and romance. Match the genre. Listeners notice when a cozy mystery gets a horror-movie read.

You're not directing a stage play. A little energy goes a long way. Theatrical overacting wears thin over ten finished hours.

  • Single narrator: one voice carries narration and dialogue
  • Multi-voice: separate voices for major speakers; narrator may still handle exposition
  • Calm: steady pace, neutral emphasis, facts and scenes do the work
  • Dramatic: sharper shifts in pace and tone, still grounded in the text

Cast a voice

Audition short passages from your book, not generic demos. Use a scene with dialogue if the book has a lot of it. Use exposition if the book is mostly narrative.

Listen for seams: odd pauses, flat energy on emotional lines, rushed dialogue. A voice that works in one paragraph may not hold for a full chapter.

Pick for the whole book, not the loudest scene. You'll live with this voice for hours.

Placeholder: voice audition with short book passages

Clone your own voice

Clone when you want the book to sound like you without recording every line yourself. Most production tools can take a short sample and use it for the full read.

Record in a quiet room, one mic, natural read. WAV preferred. Avoid background noise, room echo, and clipping. Speak at normal volume; don't whisper or shout.

Usable speech floor is about 10 seconds for a quick clone path. Better results around 30 seconds. Longer production samples are a different brief; see the blog if you need that level of detail.

  • One take, steady pace, no music or effects on the recording
  • Read a passage that sounds like your book, not a sales pitch
  • Check levels before you upload: peaks shouldn't hit the red

Direct tone and pacing

Give direction at the scene level when it matters. "Slower here," "dry humor," "she's hiding something" is enough. You don't need a note on every sentence.

Mark pauses where the text needs a beat: after a reveal, before a hard cut, when a character stops mid-thought. A simple [pause] or line break in your notes works.

Default to trusting the narrator for connective tissue. Micromanaging every comma slows production and rarely sounds better in the final mix.

Handle multiple characters and dialogue

Tag speakers clearly in the manuscript before production. Consistent labels (CHARACTER:, or quoted dialogue with attribution) keep cast and editor aligned.

Voices should be distinct enough to tell apart in a fast exchange. They don't need cartoon accents or exaggerated pitch shifts. Subtle differences in pace and weight usually do the job.

Keep narrator and dialogue separate in the listener's ear. When the narrator reads a line in a character's voice, mark it. When dialogue is tagged, the cast can stay in role without the listener losing track of who's speaking.

Add music and sound design

Score is seasoning, not the meal. Keep music and ambience well under half the runtime. Most of a book should stay dry narration and dialogue.

Favor ambience over wall-to-wall music. A room tone or quiet bed can set a scene without fighting the words. Put music under dialogue sparingly; when two people talk, leave the mix dry so nothing gets buried.

Intimate dialogue often works best with no bed at all. If the words carry the scene, leave them alone.

  • Place cues at scene opens, turns, or exits, then pull them out
  • Duck or mute beds under dense dialogue
  • Check licensed music for commercial audiobook use before you lock the mix
  • Export a dry narration stem and a scored mix if you might need both later

Review and request retakes

Proof with the manuscript open. A synced transcript (Follow along style) is the fastest way to catch wrong words, mispronunciations, and pacing that feels off.

Placeholder: delivery room with Follow along and retake request

Flag one clear note per issue. "Chapter 4, paragraph 12: should be 'breath' not 'breathe'," not a vague "sounds weird." Retake that line when you can; you don't need to re-listen to the whole book each time.

When narration is clean, move on to publish for proofing, assembly, and final files.

Common questions

How long does a voice sample need to be to clone a voice?
About 10 seconds of usable speech is the floor for a quick clone. Around 30 seconds gives better results. Record in a quiet room, one mic, WAV preferred.
Should an audiobook use one narrator or multiple voices?
Single narrator is simpler and fits most fiction and nonfiction. Multi-voice is worth it when distinct characters carry long stretches of dialogue.
How much music should an audiobook have?
Keep music and ambience well under half the runtime. Place cues at scene opens, turns, or exits, then pull them out, and duck or mute beds under dense dialogue.
How do I ask for a narration retake?
One clear note per issue with a location and the fix: "Chapter 4, paragraph 12: should be 'breath' not 'breathe'." Retake the line, not the whole book.