We ran eight pilot accounts over roughly 14 weeks before we opened our waitlist more broadly. The intention was to validate that the core loop worked: inbound lead arrives, AI drafts a reply in the rep's voice, rep reviews and sends, the thread converts at a higher rate than before. We got that validation. We also got a number of things we did not expect, and a few assumptions corrected in ways that changed how we think about the product.
This is an honest account of what happened, including the parts that were more difficult than we anticipated.
What surprised us: rep resistance was smaller than we expected
We went into the pilot expecting pushback from sales reps about AI handling their correspondence. The concern we had internalized from sales tech conversations was that reps would feel their voice was being appropriated, or that the automation would make them feel replaceable, or that they simply would not trust the output enough to send it without heavy editing.
In practice, the resistance was much lower. Most reps adapted within the first week. The pattern we saw consistently was: first day, reps reviewed every draft carefully and edited roughly 60 to 70 percent of them significantly. By day five, significant edits dropped to around 20 to 25 percent. By week three, the reps who were using the tool heavily had settled into a rhythm where reviewing a draft took under 90 seconds and most reviews resulted in a send with minor or no changes.
The reps who remained resistant through the full pilot were not opposed to AI assistance in general. They were concerned about specific reply types where they felt the AI's draft was missing context that existed in their head but not in the CRM. That is a real gap. We are working on it.
What we got wrong: the voice profile quality was uneven at first
Our onboarding process for voice profile creation was not as rigorous as it needed to be. We asked pilot accounts to share 20 to 30 email examples from each rep and built profiles from whatever they sent. The quality of the profiles varied considerably based on the quality of the examples, and we did not catch that until we had already run a few weeks of live replies and reps were flagging that some drafts did not sound like them.
The root cause was that some accounts sent us their "best" examples, which skewed the profile toward formal and polished writing that did not reflect the rep's everyday tone. Other accounts sent examples that were mostly from a single reply type (follow-ups), and the profile did not generalize well to different thread contexts.
We rebuilt the onboarding process around explicit example curation: we now specify the mix of example types we need, review the examples before building the profile, and flag cases where the examples are too homogeneous to produce a well-rounded profile. This added two days to onboarding but cut the profile quality complaints by a significant margin in later accounts.
The objection we had not anticipated: "what if the prospect figures it out"
About half the pilot accounts, at some point during the first few weeks, raised a version of this concern: what happens if a prospect realizes they are getting AI-drafted replies and feels deceived? This concern did not come up in any of our pre-pilot conversations. We were not prepared for how seriously some teams took it.
Our current position on this: the reply is drafted by an AI system and reviewed and sent by the rep. It is not misrepresenting who is sending it. It is a tool the rep is using to write faster, similar to using a template or a spell checker. The rep reviews every reply before it goes out, and they can and do change it. The risk of a prospect "figuring it out" is not zero, but it is not materially different from the risk that a prospect realizes you have a template library.
This conversation did reshape how we think about product communication. We now make the review step much more prominent in how we describe the workflow, both because it is accurate and because it addresses this concern directly. The rep is not a passive conduit. They are choosing to send each reply.
The north star metric that emerged
We went into the pilot tracking reply time, reply quality (rated by reps), and ultimately booked meeting rate. All three moved in the right direction. But the metric that became most meaningful to the accounts themselves was median first-reply time as a function of inbound volume.
Before the pilot, median first-reply time for the accounts that used AI-assisted replies was around 2 to 3 hours. After the first month, it was under 10 minutes. More importantly, the variance collapsed: the standard deviation on reply time dropped from around 90 minutes to around 12 minutes. The team stopped having outlier leads that waited 6 or 8 hours because the right rep was in a meeting or working through a backlog.
That consistency is what the accounts cared about more than the average. Knowing that every inbound lead is getting a reply in under 15 minutes regardless of what else is happening changed how the team thought about their lead response workflow. It removed a class of conversation that used to happen regularly: "this lead came in three hours ago and no one has replied yet."
What did not work the way we intended
The low-confidence flagging system, which holds drafts for human review when the system detects uncertainty, generated more interruptions than we expected in the first few weeks. The threshold was too conservative. Reps were getting flagged drafts for threads that they would have preferred to handle automatically, and the review queue was creating friction rather than reducing it.
We recalibrated the threshold based on feedback from two of the pilot accounts. The current version flags fewer cases and the ones it does flag are more clearly edge cases that genuinely need review. Getting that threshold calibration right is an ongoing process. The right answer depends on the account's tolerance for occasionally imperfect automated replies versus their tolerance for review queue overhead, and that varies.
We also learned that accounts with very different inbound lead types (different personas, different product areas, different geography) need separate voice profiles per rep and per lead segment, not a single unified profile. Building out that segmentation is a product road map item we moved up substantially based on pilot feedback.
The one thing we would do differently
We would start with the CRM and notification workflow audit before building voice profiles. Three of the eight pilot accounts had notification routing delays of 15 to 25 minutes before a rep was even aware of a new lead. The AI reply system was working, but the leads were sitting unattended for a significant window before the system even saw them, because the CRM webhook was slow or misconfigured.
Cutting the notification lag is a prerequisite for the reply speed improvements to reach their potential. A tool that drafts a reply in 90 seconds on top of a 20-minute notification delay still produces a 22-minute first reply. The order of operations matters.