← All notes

Guides · Sep 13, 2026

Cloning a voice from 30 seconds, without uploading one

A field guide to recording a clean reference sample at home, and how on-device voice cloning keeps the clip on your own Mac instead of someone's server.

usovoThe studio·5 min read·Narrated in 5:12
Abstract gradient: olive green and burnt orange, softly marbled

Listen instead · voiced on a MacBook Air, offline

Cloning a voice from 30 seconds, without uploading one
0:00 / 5:12

When RCA introduced the 44 ribbon microphone in 1931, it made radio announcers worse at their jobs for about a year. The old carbon microphones had been so insensitive that performers essentially shouted into them, and shouting is forgiving. The ribbon mic heard everything: the breath before a line, the chair, the announcer’s wedding ring against the script. An entire craft had to be invented in response, and broadcasters started calling it microphone technique. Stand here, not there. Turn your head, do not turn the page. Nobody was born knowing it, and everybody who sounded effortless on air had been taught.

Voice cloning has quietly recreated the same situation in your kitchen. The model listening to you is extraordinarily sensitive, and the thirty seconds you give it are not a formality. They are the entire brief. Everything the model knows about your voice, it learns from that one clip, including the things you did not mean to teach it.

The room matters more than the microphone

This is the part people get backwards, and it is worth saying plainly before you buy anything. A modest microphone in a good space beats an excellent microphone in a bad one, every single time. The model cannot separate you from your room. It hears one signal, and if that signal has a hard slap of reflection off a bare wall, the model learns that the slap is part of how you sound and will faithfully reproduce it in every line you ever generate.

So find soft things. A bedroom with a made bed, curtains and a wardrobe is better than any living room. A room with a rug is better than one without. If you want the oldest trick in the book, sit on the floor at the open door of a hanging wardrobe and read into the clothes. It looks absurd and it works, because coats are excellent acoustic treatment that you already own.

Sit a hand’s width to a hand and a half from the microphone, and speak just past it rather than straight into it.✓
Kill the refrigerator, the air conditioning and any fan you can reach, then wait ten seconds for the room to settle.✓
Put your phone face down in another room, not on the table.✓
Read standing up if you can. It changes your breath support and it audibly changes the result.✓
Do a warm-up read that you throw away. The first time through anything, everybody sounds like they are reading.✓

Read the script exactly as written

Usovo gives you a fixed passage to read rather than letting you say whatever you like, and that choice surprises people, so here is the reasoning. The model is not only given your audio. It is given the transcript of that audio, and it uses the pairing to work out how your particular voice maps onto particular sounds. When the transcript and the recording agree, that mapping is clean. When they disagree, it is not merely slightly worse. It degrades badly.

We found this the unglamorous way, by A/B testing it after some clones came out stuttering for no apparent reason. The culprit was never the microphone or the room. It was a speaker who had improvised a word, skipped a clause, or restarted a sentence and carried on. The passage is short, it is phonetically varied on purpose, and reading it verbatim is the single highest-leverage thing you can do for the quality of the result.

“The recording and the transcript have to agree. Everything else is a matter of taste.”Notes from the A/B

Read it the way you actually want to sound, too. If the voice is for an audiobook, read it at narration pace, warm and unhurried. If it is for short explainer clips, read it brighter and quicker. The model will copy your energy as readily as it copies your timbre, and a bored reference produces a bored voice for as long as that voice exists.

Where your voice sample goes afterwards

Nowhere. That is the short version, and it is a literal answer rather than a reassuring one.

Your microphone feeds the page, the audio is decoded and converted to a plain uncompressed WAV right there on your machine, and it is handed to a model process running locally on your own hardware. The reference clip and the voice built from it are written into your user library as ordinary files. There is no upload path in the application, no storage bucket, and no copy of your voice on any server we operate, which also means there is no way for your voice to end up in anything we train later. Deleting the voice is deleting a file.

The radio announcers of the 1930s ended up with a skill that outlived the equipment that forced them to learn it. Microphone technique still works, on any microphone, in any decade. Thirty honest seconds in a soft room will carry you further than anything you can buy, and the fact that those seconds stay on your own disk is not a feature we added. It is just what happens when there is nowhere to send them.

· · ·

Voice cloningPrivacyOn-deviceHow-to
Abstract gradient: warm amber and gold, folded like lamplight

Next note →

An 80,000-character audiobook, overnight

Guides · 4 min read

Painting: a blue car in a field of flowers

Habere et usu.
Own, and use.

Download for Mac — Apple silicon