Every sales AI conversation starts in the same place: the CRM. Feed it to a model, the argument goes, and you will get scoring, forecasting and next best action out the other end. It is a reasonable idea built on a shaky assumption.
Your CRM is a compensation artifact before it is a dataset. Stage transitions cluster around forecast calls rather than around what buyers actually did. Loss reasons are dropdown selections made under time pressure at the end of a quarter. Close dates move because someone needed them to move. None of that is dishonest, it is just how the system was designed, and a model trained on it will learn your reporting behaviour rather than your buying process.
That is the starting point for AI for modern businesses looking at sales. At iSpark we test CRM fields for signal before anyone models them, which routinely saves a quarter of wasted work. This article covers where AI in sales genuinely helps, the application most teams overlook, a documented example from a US company that published both its numbers and its methodology, the mistakes that waste the first year, and how to begin.
Test Your Fields Before You Model Them
The practical exercise takes about a week and it is the highest value thing a sales operations team can do before any AI project. Take a field you intend to use, such as loss reason or forecast category, and check whether it predicts anything. If deals marked lost to price close at the same rate as deals marked lost to product, that field is noise wearing a label.
Do that across your CRM and you will typically find a handful of fields carrying real signal and a long tail carrying almost none. Build on the handful. Our work on AI for sales teams starts here because the alternative is a scoring model that reports impressive accuracy against a target that was never real.
Where AI Powered Sales Delivers
| Application | What AI contributes | What determines success |
|---|---|---|
| Account research and call prep | Context assembled before every conversation | Whether sellers actually read it |
| Proposal and document production | First drafts from approved content | Quality of your content library |
| Lead scoring | Ranking against capacity to follow up | Whether your CRM fields carry signal |
| Forecasting | Pipeline estimates with uncertainty | Willingness to accept an unwelcome number |
| Call analysis | Patterns across conversations at scale | Coaching capacity to act on findings |
| Follow up drafting | Personalised messages at volume | Deliverability and suppression discipline |
Proposal production is the win nobody asks for and almost everybody benefits from. Sellers lose hours to assembling documents from scattered approved content, that time comes directly out of selling, and the output has a checkable right answer. It is low risk, quick to prove and rarely the thing on the request list.
A Real Example: Microsoft’s Own Sales Organisation
The challenge. Microsoft’s commercial sales organisation faced the problem every large sales team has. Sellers spent significant time on research, CRM admin and document preparation rather than in front of customers, and the company wanted evidence rather than optimism about whether AI changed that.
The solution and implementation. Microsoft deployed Copilot for Sales into the existing seller workflow, summarising account history before calls, drafting follow ups, preparing meeting briefs from CRM data and surfacing next best actions. It then measured the result against revenue metrics rather than adoption metrics, tracking revenue per seller and win rate cross referenced against actual daily usage.
The outcome. Across a cohort of 687 sellers between January and June 2024, Microsoft reported 9.4 percent higher revenue per seller, roughly 20 percent more deals won, and about 5 percent more opportunities per seller.
The business impact. More selling time and better prepared conversations, with results visible in revenue rather than in usage dashboards. The methodology deserves as much attention as the numbers. Microsoft compared sellers who chose to use Copilot daily against sellers who used it lightly. That is a correlational comparison, not a controlled trial, and high performing sellers may simply be the kind of people who adopt new tools early. Microsoft has said as much. Treat 9.4 percent as evidence of what happened in one organisation, not as a forecast for yours.
Common Mistakes in AI Sales Automation
- Building a scoring model on CRM fields nobody has tested for signal.
- Scoring more leads than the team has capacity to work, which changes nothing.
- Measuring adoption instead of revenue per seller and win rate.
- Letting outbound volume rise without suppression rules, which damages deliverability.
- Overriding an unwelcome forecast and then blaming the model for being wrong.
- Sending AI drafted follow ups that sellers have not read.
Best Practices Checklist
- Test CRM fields for predictive signal before modelling anything.
- Set scoring thresholds by follow up capacity, not by the model’s optimal cut off.
- Measure revenue per seller and win rate, cross referenced with actual usage.
- Enforce volume ramps and suppression in the system, not in a policy document.
- Keep a seller accountable for every message that goes out under their name.
- Record the forecast the model produced, including the ones you overrode.
How to Get Started
- Spend a week testing whether your key CRM fields predict outcomes.
- Start with proposal production or call preparation, where the answer is checkable.
- Baseline revenue per seller, win rate and time spent selling.
- Roll out to a cohort rather than everyone, and compare honestly.
- Move to scoring and forecasting only after the data question is settled.
Future Trends in AI Sales Enablement
Three developments look real. Research and preparation agents are moving from summarising to assembling full briefs across multiple systems, as Microsoft’s own account of its sales deployment describes. Buyers are using AI too, arriving better informed and less tolerant of a seller who knows less than they do. And deliverability is becoming the binding constraint on automated outbound, because everyone now has the capacity to send more than anyone wants to receive.
Key Takeaways
- Your CRM records reporting behaviour as much as buying behaviour. Test fields before modelling.
- Proposal production is the underrated win: low risk, checkable, and it returns selling time.
- Score to your follow up capacity, not to the model’s preferred threshold.
- Microsoft’s 9.4 percent came from a usage comparison, not a controlled trial.
Frequently Asked Questions
Why do AI lead scoring projects underperform?
Usually because CRM fields record reporting behaviour rather than buying signals, or because more leads are scored than the team can actually follow up.
What should sales automate first?
Proposal production and call preparation. Both return selling time immediately, carry low risk, and have an answer you can check against approved content.
Can AI forecast a pipeline accurately?
It can produce estimates with stated uncertainty. Accuracy depends on whether stage data reflects buyer behaviour or internal reporting rhythms.
Does AI outbound still work?
Only with strict volume ramps and suppression enforced by the system. Deliverability, not message quality, is usually what determines outcomes now.
What should we measure?
Revenue per seller, win rate and time spent selling, cross referenced against actual tool usage. Adoption figures on their own prove nothing.
Where to Take This Next
Sales teams getting real value from AI did something unglamorous first. They checked whether their data meant what the field names claimed, started with work that returns selling time rather than work that sounds impressive, and measured revenue rather than adoption. The teams that skipped that step usually spent a year proving a model was accurate against a target that was never real.
If you want an independent view of whether your CRM data would support a scoring or forecasting build, iSpark runs fixed scope sales assessments that end with a written recommendation, including recommending against it where the data does not hold up. A week of field testing is the place to start.
