Blog
Voices6 minutes read

Twenty Seconds Is Enough: What Our Best Author Voice Clones Were Recorded On

Published
DM

David Mainayar

Founder

Share this article


Summary

Four author clones, input and output. The 20-second clip each author sent, the clone reading their book, and the recording rules that held.

Our current clone path needs about twenty seconds of you talking. Not two hours, not a booth. Below are four author clones we shipped this month: the exact clip each author gave us, and the clone reading their book. Then the short list of what those clips had in common.

Key numbers: one clip, 20 to 25 seconds, one voice, no music. Peaks under −3 dB. A phone is fine. Levels on the four clips below ranged from −29 to −18 dB RMS and every one of them cloned.

Four clones, input and output

Left is what the author sent. Right is the clone, voice only, reading the opening of their book. Shared with each author's permission.

Jeremy Dyer, The Fundamental Investor

Twenty seconds of Jeremy talking. Full chapters are now in production on this voice.

Jeremy's clip, 20 s
Clone reading chapter one

Mitra Manesh, The Attentionist

Twenty-one seconds cut from a welcome video she had already published. No new recording.

Mitra's clip, 21 s (from her video)
Clone reading the sample

Alan Berg, AI For The Real World

Twenty-one seconds from his podcast. The same clone then read his Spanish title, Ghosting, ¿por qué me ignoran?

Alan's clip, 21 s (from his podcast)
Clone reading in English
Same clone reading in Spanish

Antony Hylton, Biblical Digital Marketing

Twenty-one seconds, British English. The hottest clip of the four (peaks at −0.1 dB). It still cloned, and we would still ask for 3 dB less.

Antony's clip, 21 s
Clone reading the preface

What the four clips had in common

  • One voice, start to finish. No host, no intro music, no kid in the hallway.
  • Continuous talking. Twenty seconds of speech, not twenty seconds with ten seconds of gaps.
  • The register of the book. Jeremy talks like a founder. Alan talks like a speaker. The clone copies the delivery it hears, so talk the way you want the book read.
  • Nothing fancy. Phone recordings and straight cuts from existing media. Two of the four came off a podcast and a video, and the model still learned the person, not the codec.

Level was the least consistent thing across the four and it mattered least. Noise and music are what break a clone. Level is a fader.

How to record it on a phone

This is the direction our engineer gives every author. It fits on a sticky note.

  1. Use the phone, not earbuds. Hold it 20 to 30 cm from your mouth, USB-C or charging port pointed at your face. The good mic is at the bottom.
  2. Avoid small hard rooms. A bathroom or an empty closet gives you slap. A bedroom or living room with a bed, sofa, curtains, rug soaks reflections. Sit near the soft stuff. Do not chase a perfect room; a bit of reverb on a well-placed mic is fixable.
  3. Turn everything off. Fan, AC, dishwasher, notifications. Windows shut.
  4. One take, 20 to 30 seconds, talking like you would read the book. A paragraph from your first chapter is ideal. Do not whisper, do not announce.
  5. Listen back once. If you hear the room, the fridge, or your voice pumping up and down, move and redo it. Do not try to fix it in an app.
  6. Send the original file. Email, Drive, or our intake link. Do not forward it through a messaging app; most of them recompress voice notes.

Have a real mic? Point it at your mouth

Last week an author sent us a podcast recorded on a broadcast dynamic mic that would embarrass most studios. Our engineer watched thirty seconds of the video and found the problem in one line: the mic was pointed at his desk. Everything else followed from that. He had cranked the input gain to make up for it, the gain pulled in the room, and the voice sat thin under it.

The fix, verbatim from that thread:

  • Mic pointed at your mouth, 5 to 7 inches away. Two to three inches sounds better and clips more. Play it safe.
  • Lower the input gain. If you have been recording off-axis, your gain is set for a mic that is pointing away. Point it at you and bring the gain down until peaks sit around −6 dB.
  • The furnished room is fine. Bed, sofa, curtains, done. If moving to another room is a hassle, do not move. Placement is the problem, not the room.
  • Reverb is fixable. Placement is not. A well-recorded voice with some room on it cleans up into a good clone. A voice recorded off-axis with the gain cranked does not.

Same rule for your phone or a USB mic. The capsule has a front. Find it and talk into it.

Already on tape? Use that

If you have a podcast, a YouTube intro, a webinar, a voice note to a friend, cut twenty seconds where only you speak and nothing plays under you. Two of the four clones above were built that way. The cut matters more than the source: no music bed, no crosstalk, no applause.

What kills it

  • Music or a bed under your voice.
  • A second voice anywhere in the clip.
  • A mic pointed somewhere other than your mouth.
  • Bathroom or empty-room echo. Mild reverb in a furnished room is fine.
  • Auto-gain pumping (phone voice memo apps do this when you move).
  • Whispering half and projecting half.
  • Reading in a voice you would never narrate in.

Twenty seconds means one bad second is five percent of the sample. Cut it out, or record again.

Longer samples

Twenty seconds gets you a production clone, and it is what every author above started with. If you want to record more, the long-form brief still applies and helps on pronunciation and retakes. It is not required to start.

Send your manuscript chapter and a twenty-second clip through the author portal and you hear the first chapter in your own voice. No clip yet? Start with a free first chapter on a cast voice and record later.

Hear your first chapter free

Send the manuscript. We cast, direct, and score the opening and send it back in about a day. No card, yours to keep.

More to read