To judge whether an AI rollout works, compare teams using it with teams not using it over the same period, against metrics fixed and baselined before deployment.1 The note’s author considers this measurement design the part that transfers to any internal tooling rollout.1
Rules
- Baseline first. Pick one activity metric and one revenue metric and baseline both before deployment; without a baseline and a concurrent control group, the readout becomes anecdotes.1
- Compare cohorts, not calendars. Compare pilot and non-pilot teams in the same quarter (pipeline per rep, cycle length, win rate) rather than one team before and after, so market and season affect both.1
- Don’t lead with hours saved. Time saved is capped at what the team is paid for that time; lead with expansion and new capabilities.1
- Read spend against output. High spend alone is not a signal: a heavy user who also produces daily is the program working; a heavy user producing little gets coaching first, and a lower cap only if coaching fails.1
Where returns show up
Returns appear in order: usage, then output, then CRM outcomes, then cost.1
| Group | Meaning | Example |
|---|---|---|
| Efficiency | Same work, faster | Deal prep from 3–4 hours to 45 minutes |
| Expansion | More output from the same team | More accounts covered, more pipeline per rep |
| New capabilities | Work that did not happen before | Every account in a book scored overnight |
As described in the guide.1