When people hear "voice training" in the context of AI sales tools, they often picture something like fine-tuning: feeding a language model thousands of examples and having it update its weights to write in a specific style. That is not what we do, and it is not the right framing for what the problem actually requires.
Fine-tuning at the rep level is impractical for a few reasons. The data volumes needed to fine-tune meaningfully on a specific person's writing are in the tens of thousands of examples. A rep does not have that. A rep has a sent folder with a few hundred to a few thousand emails if you are lucky, most of which are replies to specific situations that do not generalize cleanly to new contexts. And even if you had the data, running per-rep fine-tunes at the model level does not scale to a product used by teams of five to fifty people.
The actual problem is a prompting and retrieval problem, not a weights problem. The question is: given a rep's actual writing samples, how do you construct a prompt that produces output matching that rep's voice with enough fidelity to pass a reading by someone who knows them?
What carries voice in an email
Before we can extract voice from examples, we need to know what we are extracting. In written sales emails, the signals that carry the most voice information are structural, not lexical. Two reps can share most of their vocabulary and still sound quite different because of differences in the following dimensions.
Sentence length distribution. Some reps write in short, punchy sentences that rarely exceed 15 words. Others write longer sentences with subordinate clauses and multiple qualifications. This is one of the most consistent style signals across writing samples from the same person.
Opening structure. How does the rep start a reply? Do they echo something from the prospect's message? Do they open with a direct statement? Do they acknowledge a shared context? The opening sentence pattern is highly consistent for most writers and very effective for voice anchoring.
Hedging level. Does the rep state things directly ("we can do X"), or do they hedge ("we should be able to do something like X for you")? Some reps are naturally assertive in writing; others are more tentative. Neither is wrong. Both are distinctive.
Sign-off habits. The closing line and signature are often surprisingly individualistic. Some reps always close with a question to maintain forward motion. Others close with a soft statement and let the prospect respond on their terms.
The 20-example rule
We have found that a voice profile built from 20 to 30 well-chosen examples is sufficient to anchor on these structural signals with enough fidelity for AI-drafted first replies to pass as the rep to prospects who do not know the rep well. "Passing to prospects who know the rep well" is a harder bar, and a few reps in our pilot flagged specific phrases or openings that did not quite sound like them. We treat those as edge cases to refine over time.
The choice of examples matters more than the quantity. We ask for a mix of thread types: a few first replies to cold or inbound inquiries, a few follow-up nudges on stalled threads, a few objection responses, and a few replies that the rep considers unusually effective. The diversity of thread type ensures the voice profile is not overfit to a single communication context.
One mistake we see often is curating only the rep's "best" emails. This tends to produce a profile that over-indexes on formal, polished writing and loses the natural rhythms that show up in everyday correspondence. Those rhythms are often what makes the voice feel authentic rather than polished. A rep whose casual Wednesday follow-up sounds different from their carefully crafted proposal-stage email is going to have a better-fitting profile if both types are included.
How the profile is applied at inference time
When a new lead arrives and the system drafts a reply, the voice profile is incorporated through a structured prompt that includes annotated examples from the rep's library, a set of inferred style rules extracted from the examples, and specific guidance on what to avoid based on patterns that the rep's writing consistently does not include.
The system does not generate the reply from scratch and then style-match it afterward. The style guidance is baked into the generation prompt from the start, which produces output that is structurally consistent with the rep's patterns rather than content-correct but style-wrong.
We also do not try to match voice at the level of specific phrases. Phrase reuse is actually a signal of template-based writing, which is the opposite of authentic. The goal is to produce output that follows the same structural logic the rep uses, not to reproduce their exact sentence constructions.
When voice matching fails and why
Voice matching degrades in predictable scenarios. The clearest case is when the rep's communication style varies significantly by context in ways the example set does not capture. A rep who writes one way in initial outreach and a very different way in deal-closing conversations will have a profile that blends both modes, and the output will feel inconsistent in contexts that should trigger one mode or the other.
The second failure mode is rep writing that is itself highly variable. Most people have consistent writing patterns even when they think they do not, but some reps genuinely shift style based on factors that are hard to encode. For these reps, the voice profile captures something like an average that feels slightly off in both directions. The honest answer for these cases is that the voice fidelity ceiling is lower, and the review step becomes more important.
The third failure mode is example set staleness. A rep's writing style evolves, and a profile built from emails from 18 months ago will drift from their current voice over time. We recommend refreshing the example set periodically: adding recent emails and removing old ones that no longer represent how the rep writes.
What this is actually solving for
The point of voice matching is not to create a convincing impersonation. It is to make sure that the AI-drafted reply is low-friction for the rep to review and send. If the draft already sounds like the rep, the review becomes confirmation rather than editing. That is a much faster workflow than receiving a generically drafted reply and rewriting it into the rep's voice before sending.
The more work the rep has to do to the draft before it is sendable, the closer the workflow gets to just writing the reply from scratch. Voice fidelity is the variable that determines whether AI-assisted replies are a genuine time saving or just a starting point that still requires substantial human effort.