What Should CPG Teams Look for in an Analytics Agent?
Judge an analytics agent on five things: provider rulebook coverage, verification that runs on fixed rules, brand memory, data ownership, and the right level of autonomy with a human in the loop. Do not judge it on whether it has a chat interface. Every product in this category has a chat interface. The five criteria are where they actually differ.
The skepticism is warranted, and the analysts agree. Gartner forecasts that over 40% of agentic AI projects will be canceled by the end of 2027 and warns about "agent washing," the rebranding of existing chatbots and automation as agents; by its estimate only about 130 of the thousands of self-described agentic vendors are real [1]. Your evaluation process is the filter.
The five criteria
1. Provider rulebook coverage
Syndicated measures carry aggregation rules that differ by provider, and an agent that does not enforce them produces confident wrong numbers (the mechanics are in the aggregation trap). Good looks like named, per-provider rulebooks: the vendor can tell you exactly how SPINS, Circana, and Nielsen math is enforced, separately. Sous ships rulebooks for all three, so cross-provider brands run one workflow instead of three. A vendor who says "our model understands retail data" has answered a different question than the one you asked.
2. Verification you can audit
Both major AI labs are explicit that agents need ground truth from their environment at each step [3] and clear guardrails around what they do [2]. For analytics, that cashes out as two testable properties: every query is checked against the provider's rules by fixed logic that runs the same way every time, before the answer ships, and every number remains traceable to the query that produced it. In Sous, the queries live inside the workbook, so a human can audit the math behind any figure. If a vendor cannot show you the query behind a number in the demo, treat the number as unverified. The benchmark case for why this matters is in "how accurate are AI agents on syndicated data?"
3. Brand context and memory
A tool that starts every session from zero will give you generic answers forever. Good looks like a system that accumulates brand context: your products, categories, retailers, competitive set, and vocabulary, sharpened by every question and correction. The test plays out over time: ask in the demo what the product will know about your brand in month 6 that it did not know in week 1, and how you can see and edit that knowledge.
4. Who owns the data and context
The Sous position is blunt: brands should own their data and context, not rent dashboards. The data half is familiar contract diligence. The context half is new and easy to miss: after a year of use, the accumulated brand context (definitions, corrections, competitive framing, history) is one of your most valuable data assets. Ask what happens to it if you leave. If the answer is "nothing exportable," you are not buying a tool. You are being acquired by one.
5. Autonomy level and guardrails
More autonomy is not automatically better. OpenAI's own guidance says high-risk and hard-to-reverse actions should trigger human oversight until confidence in the agent is earned [2]. In analytics, the equivalent design question is what the agent does when something unexpected happens. Sous's auto-refresh is the reference pattern: it loads and rebuilds on its own when the new period's file validates, and it stops and asks when columns are renamed or measures go missing. Ask every vendor: show me what your agent does with a malformed file. The ones who never considered the question will demo the happy path.
Questions to bring to the demo
- Which syndicated providers do you enforce correct math for, and how, specifically?
- Show me the query behind this number.
- Average my % ACV across two markets. Does the system refuse, warn, or comply?
- What does the product know about my brand in month 6 that it did not know in week 1?
- If we cancel, what do we get back, in what format?
- What happens when the new period's file has a renamed column?
- Which decisions does the agent make alone, and which does it bring to a human?
Any serious vendor should enjoy this list. Evasion on any single question is signal.
The % ACV question deserves a note, because it is the sharpest 30 seconds in the whole evaluation. Averaging % ACV across markets is illegal math on syndicated data, the kind that produces numbers that look plausible and are wrong. A system with a real rulebook refuses or reframes the request. A system without one complies instantly and fluently, and you have just watched the exact failure that would otherwise have debuted in a buyer meeting. That one question settles the rulebook claim either way.
Running the evaluation
Structure the pilot around questions you already know the answers to. Pull the 10 real questions your team answered by hand last period, run them through the tool, and score three things: how many answers match your known-good numbers, whether you can trace each answer to its query, and how the system handled the ambiguity in each question (did it make the same judgment calls your analyst did, and did it surface them). Then re-run 2 of the questions a month later, after corrections, and see whether the system learned. That final step tests brand memory, the criterion no demo can fake, because a demo has no month 2.
Resist the urge to evaluate on novel, exotic questions. The exotic question flatters whichever tool guesses most confidently. The boring known-answer question is the one that measures accuracy, and accuracy on boring questions is what you are actually buying.
How the options compare
| Criterion | Dashboard | Analyst | Generic AI chatbot | Analytics agent |
|---|---|---|---|---|
| Provider rulebook coverage | Hard-coded at build | If experienced | None | Enforced per provider |
| Verification and auditability | Fixed logic, opaque to users | Ad hoc, in their head | None | Fixed-rule checks, traceable queries |
| Brand memory | None | Excellent, until they leave | Session-only | Compounding brand context |
| Data and context ownership | Vendor-shaped | In one person's head | Nowhere | Owned by the brand |
| Autonomy with guardrails | None | Full human, no leverage | Autonomy without checks | Autonomous with stop-and-ask |
The chatbot column is the one to study. It matches the agent column on interface and on nothing else, which is exactly why agent washing works on buyers who evaluate by demo feel [1]. Score the five rows, weight rulebooks and verification heaviest, and the field narrows fast.