Is AI describing your
business accurately?
Getting mentioned is the easy half. The harder question is whether the sentence next to your name is still true. Moose grades every answer that names you against what you sell today, and shows you the sentence that failed.
Sentiment: Neutral. The answer states what the product does, with no praise or criticism.
Narrative: "built for small Shopify stores" contradicts the profile, which covers small brands through enterprise retailers.
Parity: The answer denies service and helpdesk tooling. The profile lists both as shipped.
Klaviyo is a marketing automation tool best known for email, and it is built for small Shopify stores that want automated welcome and abandoned cart flows.
It does not include customer service or helpdesk tooling, so larger retailers usually pair it with a separate support platform.
Unbothered by AI search. Has never once been misdescribed, mostly because he refuses to explain himself. Consulted on this page.
A mention you would not have approved is worse than no mention
Models learned you from the internet as it was two years ago. You repositioned, shipped three products, moved upmarket. The training data did not. So the answer names you, counts as a win on every visibility dashboard, and describes the company you used to be.
Then it gets worse in a specific way: the answer says you do not do something you have done for a year, or credits a competitor's feature to you. A buyer reads that and rules you out before you ever hear about it.
Green arrow, healthy trend line. It counted your name and stopped there. It has no opinion on whether the sentence around your name was true.
Four answers put you in the old category. Three deny a capability you shipped last spring. Two hand a competitor's feature to you. Each one names the sentence and the engine that said it.
Nothing is graded against a vibe
You write down what is true about the company once, in the Brand Truth Profile. Every verdict after that is measured against your words, not a model's guess at what a good answer looks like.
No profile, no verdicts. Drift and parity are measured against your fields, so Moose leaves them blank rather than inventing a yardstick.
Narrative drift, and feature parity
Every answer that mentions you gets both, and each one comes with a single sentence of reasoning that has to name the specific claim it is about.
Is this answer describing the company you are, or one you used to be? Wrong customer type, wrong category, dead positioning.
Is the answer factually wrong about what you can do? This one grades misinformation. Leaving things out is fine.
Omission. If an answer stays factually correct but does not list every feature you have, that is still good. Nobody's answer lists everything.
Feature claims are only judged inside the scope of what the prompt asked about. A question about SMS is not marked down for skipping your analytics.
Built to not cry wolf
A monitor that flags everything is a monitor you stop opening. The scoring rules are written to make false alarms the expensive mistake rather than the safe one.
The scoring rules say plainly that false alarms make the metric useless, and that if the model cannot point to a specific in-scope sentence that is wrong or misleading, it must not flag it. Silence is the default, not the exception.
A plain feature list, or a table row stating what you do, is neutral. Negative needs real criticism or an unfavourable contrast, not merely being mentioned last.
Sentiment verdicts carry a quote copied from the answer word for word, so you can judge the call yourself instead of trusting a summary of it.
If the scoring call fails, the fields stay empty. Moose does not fabricate a verdict to fill a box, and the model is told never to invent details that are not in the response or your profile.
Feature claims are only assessed inside the scope of the prompt. An answer about SMS deliverability is not marked down for not mentioning your reviews product.
Which model is getting you wrong
Every observation is stamped with the engine that produced it, so drift, parity, sentiment, mentions and citations all break down per model. "Show me every answer where an AI got a feature wrong, and which one said it" is one filter.
For ChatGPT, Grok, Perplexity and Google AI Mode, Moose reads what the real consumer product answers from a browser session on your machine. If that fetch fails there is an API fallback, and every observation records which route produced it and whether a fallback was used. You can always tell how an answer was obtained.
One reading is trivia. The change is the story.
Runs go on your schedule: daily, weekdays, weekly or manual. What you look at afterwards is the difference between runs rather than another snapshot.
Daily, weekdays, weekly or manual. Scheduled runs can email the summary, so the finding reaches you without you opening anything.
Mention rate, citation rate and positive sentiment charted by day, week or month, rather than one fixed window somebody else chose.
Compare last month to the month before, or the two weeks either side of a launch, side by side.
Narrative alignment, feature accuracy, negative sentiment count, mention rate and citation rate, each with the change since the previous run.
Lost mention, gained mention, new citation, negative sentiment. Each one names the prompt and engine, with the quote attached.
Filter everything by narrative drift or feature parity, so "every answer where an AI got a feature wrong" is one click rather than an export and a spreadsheet.
Numbers you can act on without checking them first
Most of this category quietly counts its own failures as your losses. Three decisions here are deliberate.
If an engine could not be reached or read, that observation is excluded from every metric, delta and alert. A network failure never becomes a fake visibility drop, and a failed baseline never becomes a fake recovery.
The same prompt and engine can be sampled up to five times per run, because a single AI answer swings wildly while the mention rate across samples holds steady. Results carry a Wilson 95% confidence margin.
Two points of movement inside the margin is not a trend, and Moose will not present it as one. You get to spend your week on the changes that are real.
If the brand was never mentioned, or the engine never answered, the observation is marked not applicable without spending a scoring call on it. Repeated samples of the same prompt collapse to one verdict for lists, emails and alerts, so a heavily cited prompt does not flood the report with duplicates.
Catch it on purpose
Rather than guessing what to track, Moose writes prompts anchored on your brand name and the fields you filled in: capabilities, competitor comparisons, customer segments, and the misrepresentations you already see.
Prompts are generated from your brand name and your Brand Truth fields, so the answers they return are about you.
Diagnostic questions you write yourself are passed through unchanged rather than rewritten into something more generic.
Unbranded queries are left out deliberately. They do not return citations about you, which is the entire point of the exercise.
Review, edit and approve each suggestion before anything gets tracked. They land in their own Diagnostic Prompts category so they never muddle your core set.
Where the wrong story came from
Knowing an engine has you wrong is half a finding. The other half is the page that taught it, and that is usually something you can go and fix.
Turning a finding into a fix
A verdict on its own is a complaint. In the same window, it becomes a brief, a draft, and a change you approved.
Who gets scoring
Two caveats worth reading before you download, because we would rather you find out here.
Narrative drift and feature parity run on paid managed and BYOK workspaces that have a Brand Truth Profile filled in. Daily cadence, all eight engines, email reports, and export to PDF or CSV.
Free preview mode gets no scoring at all: no drift, no parity. It is capped at 3 prompts, weekly or manual runs, and the local-fetch engines. Enough to see whether you appear, not enough to grade how.
Questions people ask
Do I have to fill in the whole Brand Truth Profile?
Every field is optional, but the profile is the yardstick, so an empty one means no drift or parity verdicts at all. The canonical description, ideal customer and key capabilities do most of the work. Start with those three and add the rest as you see what gets flagged.
Does this flag an answer for leaving features out?
No. Feature parity grades misinformation, and leaving a feature out is not a fault. An answer that stays factually correct while mentioning only two of your six products is still good. It only flags a claim that is wrong: a capability denied, a capability invented, or a competitor's feature credited to you.
How do I know a verdict is not a hallucination?
Each verdict carries a one-sentence reason that has to name the specific claim it is about, and sentiment verdicts carry an exact quote copied from the answer. If the scoring call fails, the fields stay empty rather than being filled with a guess.
Which engines can it read?
ChatGPT, Claude, Gemini, Grok, Perplexity, Google AI Mode, Google AI Overviews and Bing Copilot. For ChatGPT, Grok, Perplexity and Google AI Mode there are routes that read what the real consumer product answers, with an API fallback if that fetch fails. Every observation records which route produced it.
Can a single bad answer swing my numbers?
Less than you would expect, because the same prompt and engine can be sampled up to five times per run and the results carry a Wilson 95% confidence margin. Repeated samples also collapse to one verdict in lists, emails and alerts, so a heavily cited prompt does not flood the report.
Is this available on the free plan?
No. Drift and parity scoring runs on paid managed and BYOK workspaces with a Brand Truth Profile filled in. Free preview mode is capped at 3 prompts, weekly or manual runs, and the local-fetch engines, and it gets no scoring at all.
More things to ask Moose
All use casesAsk whether you have written it before you brief it again.
The version of you an engine can resolve and cite.
The searches Google writes for itself before it answers.
Twelve narrower searches behind one buyer question.
The actual top ten, People Also Ask, and the AI Overview with sources.
The full list, and what Moose does step by step in each one.
Write down what is true.
Then hold every answer to it.
Fill in the Brand Truth Profile once, run a scan, and read the first nine sentences that do not match. Most teams find at least one they did not know about.