Skip to content
Start a conversation
AI and Analytics

Before you buy an AI analyst, four preconditions

Conversational analytics has become plausible in a way it was not two years ago. Ask a question in English, receive an answer computed from your warehouse. The demonstrations are impressive and the underlying capability is real.

What has not changed is that the answer is only as good as the substrate it is computed from. A system generating queries against a messy warehouse produces confident, fluent, wrong answers, and it produces them faster than a human would and with less hedging. The failure mode of a bad analyst is a slow wrong answer with visible uncertainty. The failure mode of a bad AI analyst is an immediate wrong answer stated plainly.

Before buying, we would want four things true. Not because the technology is unready, but because these determine whether it produces value or liability in your specific estate.

One definition per concept

Ask most warehouses for revenue and there are several defensible answers. Gross or net. Recognised or billed. Including or excluding intercompany. A human analyst navigates this by knowing which one the asker means, usually from context the question did not contain.

A generated query picks one. It may pick differently on Tuesday than it did on Monday, based on phrasing that the user did not intend as meaningful. Two people asking what they believe is the same question get different numbers and have no way to see why.

The precondition is a semantic layer where each business concept resolves to exactly one definition, and the assistant queries that layer rather than the underlying tables. Without it you have automated the production of plausible numbers, which is not the same as automating analysis.

Access control that survives generation

Warehouse permissions are frequently coarser than people assume, with fine-grained control implemented in the BI layer through row-level filters and restricted datasets.

An assistant querying the warehouse directly bypasses all of it. The question of whether a regional manager can ask about another region’s performance, or whether anyone can ask about individual salaries, has to have a clear answer that lives below the assistant rather than beside it.

Relying on the model to decline is not a control. Prompt-level restrictions are guidance, not enforcement, and the gap between the two is exactly the sort of thing that surfaces during a security review rather than before.

A basis for judging the answer

The most under-discussed risk is that fluency is a poor proxy for correctness, and users do not have a way to tell the difference.

An analyst delivering a number brings implicit context: they know it looks high, they mention the promotion that ran, they flag that one region reports late. The assistant delivers the number. The user, who asked because they did not know the answer, has no basis for scepticism.

Practical mitigations exist and should be requirements rather than nice-to-haves. Show the generated query. Show which definitions were used. Show data freshness. Show the comparison period unprompted, so an anomalous figure looks anomalous. Give the user something to be suspicious with.

A clear scope of question

These tools are strong on retrieval and aggregation. What was revenue last month by region. How many accounts churned in Q3. Questions with a determinate answer that a well-formed query returns.

They are considerably weaker on causal and comparative questions. Why did churn increase. Whether the campaign worked. These require judgement about confounders, about what a valid comparison group is, about whether an observed difference means anything. A confident answer to a causal question from a system that cannot reason about confounding is worse than no answer.

Scope the deployment deliberately, and be explicit with users about where the boundary sits. The organisations getting value from this are the ones that positioned it as a fast way to get facts, freeing analysts for the questions that need judgement. The ones struggling positioned it as a replacement for asking an analyst anything.

The uncomfortable implication

Every precondition above is data platform work. A semantic layer. Coherent access control. Reliable freshness metadata. Documented definitions.

This is the work organisations have been deferring for years, and it is precisely the work an AI analyst is often positioned as a way to avoid. The pitch is that natural language removes the need for a well-modelled warehouse. The reality is closer to the opposite: a conversational interface raises the cost of a poorly-modelled warehouse, because it removes the human who was quietly compensating for it.

Our advice is not to wait. It is to sequence honestly. If the semantic layer exists, this technology will pay back quickly and visibly. If it does not, building it is the prerequisite, and it is worth doing regardless of whether an assistant ever sits on top.

How we would pilot it

Assuming the preconditions are broadly met, the pilot design matters more than the vendor selection, and most pilots are designed to produce a positive result rather than an informative one.

The standard approach is to gather enthusiastic users, let them ask questions, and collect satisfaction feedback. This measures whether people enjoyed the experience. It does not measure whether the answers were right, and satisfaction correlates with fluency, which is exactly the thing that is untrustworthy here.

A more useful design is adversarial. Collect a set of real questions with known correct answers, including some where the obvious query is wrong: a metric with a non-obvious exclusion, a period affected by a definition change, a question that is ambiguous between two defensible interpretations. Run them. Score accuracy, and score separately whether the system signalled uncertainty where uncertainty was warranted.

The second score is the one that predicts whether this will be safe in production. A system that is right eighty percent of the time and flags its own shaky cases is deployable. A system that is right ninety percent of the time with uniform confidence is not, because users have no way to identify the ten percent and will find them by acting on them.

Who it actually helps

One observation from the deployments we have seen work. The beneficiary is usually not the executive the demonstration was aimed at. Executives ask relatively few data questions directly and have people to ask.

The real gain lands with the operational middle: people who need a number several times a week, currently either wait for an analyst or approximate it, and whose questions are genuinely of the retrieval kind these systems handle well. That is a large population making a lot of small decisions with worse information than they should have.

Aiming the deployment there is less impressive in a steering meeting and considerably more likely to produce something people still use in six months.

Ready to turn complexity into your next advantage?

Book a discovery call