All articles
Implementation3 min read

What to check in a conversation intelligence pilot before scaling it

A five-store pilot can go well and prove nothing. The four questions that actually decide whether it is worth rolling out across the chain.


Almost every pilot of this kind ends with a report saying it worked. The problem is that most of them did not measure the thing that mattered, and the decision to scale gets made on the gut feel of the team that ran it.

These are the four questions that actually decide.

1. Did associates use it without being asked?

It is the only adoption metric that matters and almost nobody looks at it.

If usage came from a supervisor reminding people in every huddle, what you proved is that your supervisor is persistent, not that the tool is useful. In a two-hundred-store chain there are not enough supervisors to go around.

What to look at: the share of associates who opened their own feedback unprompted in the last week of the pilot, compared to the first. If it drops, do not scale yet.

2. Did a concrete behaviour change, or just the mood?

"The team feels more focused" is not a result. It is pleasant and it cannot be audited.

What to look at: pick a specific behaviour before you start (the close attempt, for example) and measure it in week one and in the final week. If close attempts went from 30% to 55% of conversations, you have a result. If you did not pick the behaviour before starting, the pilot can no longer prove this and no later analysis will fix it.

3. Does the effect survive the novelty?

The first two weeks of any measurement improve the numbers. That is the Hawthorne effect and it has nothing to do with the product: people behave differently when they know they are being watched.

What to look at: a four-week pilot cannot answer this. Only from week six or eight can you see whether the behaviour stuck or whether it was opening-week enthusiasm. A short pilot can prove the tool works; it cannot prove the change lasts.

4. Do the pilot stores look like the chain?

The most common mistake: pilot stores get chosen for having the best manager, because those are the ones who will cooperate.

They are also the ones with the least room to improve and the least resemblance to the rest. A pilot that works in your three best stores says nothing about the hundred in the middle, which is where the money is.

What to look at: include at least one underperforming store and one with high staff turnover. Those two tell you whether the result is repeatable.

A pilot that goes well in the best stores with the supervisor standing over it does not prove the product works. It proves it works under the best possible conditions, which are exactly the ones you will not have when you scale.

What to agree before you start

Three things, in writing, before the first device goes in:

  • The behaviour you will measure, chosen and defined.
  • The decision threshold. What number means scale and what number means stop. Defining it afterwards means defining it to justify what you already wanted to do.
  • The legal framework and in-store signage, reviewed by your legal team. This is the point that stalls the most rollouts when left to the end, and costs the least when settled at the start.

A well-designed pilot can give you a no, and that is a good result too: it costs five stores instead of two hundred.

Get these in your inbox

What we learn deploying conversation intelligence in retail. No schedule, only when there is something worth sending.

One click to unsubscribe. We never share your address.

See this on your own stores.

We walk through real conversations from a store like yours and what Vera found in them.