Use case Accuracy

Is AI describing your
business accurately?

Getting mentioned is the easy half. The harder question is whether the sentence next to your name is still true. Moose grades every answer that names you against what you sell today, and shows you the sentence that failed.

Watch Moose grade an answer Start with the yardstick
Chat Workflows
Klaviyohttps://klaviyo.com
New chat
Home Inbox Context Visibility Workflows Audio Chats Library Connections
Recents
Where are we described wrong?Chat · 2m Visibility scan · 11 promptsVisibility scan · 20h Brand Truth Profile updatedContext · 1d Capture ChatGPT fan-out qu…Chat · 2d
Sam SamPremium
Observe
Visibility
Klaviyo  |  paid managed visibility  |  latest run Aug 1, 10:02 PM
Prompt Manager Run now
Share of mentions
74%
Share of citations
61%
Sentiment
58% positive
Narrative drift
82% aligned
Feature parity
71% accurate
Topic: All AI Engine: All Sentiment: All Narrative: Drifted or partial Parity: Inaccuracy found
Answers that describe you wrong 9 of 51 observations
What is Klaviyo and who is it for? Gemini (local fetch) · “built for small Shopify stores” Drifted Inaccuracy
Best CRM for a consumer brand doing $40M ChatGPT (local fetch) · “no native service or helpdesk layer” Partial Inaccuracy
Klaviyo vs Mailchimp for a growing brand Perplexity (local fetch) · “both are email marketing tools” Partial Needs review
Which platforms do enterprise retailers use? Grok (local fetch) · “more of an SMB tool than an enterprise one” Drifted Good
Gemini (local fetch) response
What is Klaviyo and who is it for?
2026-08-01·Brand & category·Mentioned
Investigate in chat Close
Position #3 Sentiment Neutral Narrative Drifted Parity Inaccuracy found
Why this score

Sentiment: Neutral. The answer states what the product does, with no praise or criticism.

Narrative: "built for small Shopify stores" contradicts the profile, which covers small brands through enterprise retailers.

Parity: The answer denies service and helpdesk tooling. The profile lists both as shipped.

Response, as returned

Klaviyo is a marketing automation tool best known for email, and it is built for small Shopify stores that want automated welcome and abandoned cart flows.

It does not include customer service or helpdesk tooling, so larger retailers usually pair it with a separate support platform.

Sources cited  a-review-blog.com/klaviyo-review (2023)
Hi, Moose 0.3.281 Beta Need help?
Moose, a dog, unbothered

Unbothered by AI search. Has never once been misdescribed, mostly because he refuses to explain himself. Consulted on this page.

The problem

A mention you would not have approved is worse than no mention

Models learned you from the internet as it was two years ago. You repositioned, shipped three products, moved upmarket. The training data did not. So the answer names you, counts as a win on every visibility dashboard, and describes the company you used to be.

Then it gets worse in a specific way: the answer says you do not do something you have done for a year, or credits a competitor's feature to you. A buyer reads that and rules you out before you ever hear about it.

What a visibility tool tells you
Mentioned in 74%

Green arrow, healthy trend line. It counted your name and stopped there. It has no opinion on whether the sentence around your name was true.

What Moose tells you
Mentioned, and wrong in 9 of them

Four answers put you in the old category. Three deny a capability you shipped last spring. Two hand a competitor's feature to you. Each one names the sentence and the engine that said it.

Step one

Nothing is graded against a vibe

You write down what is true about the company once, in the Brand Truth Profile. Every verdict after that is measured against your words, not a model's guess at what a good answer looks like.

Ground · Brand Truth Profile
How should AI describe your company today?
A B2C CRM for consumer brands: email, SMS, push and reviews on one unified customer profile, with service and analytics in the same platform. Used by small brands through enterprise retailers.
Write the version you would want to see in ChatGPT, Perplexity, Gemini or Google AI answers.
What does AI often get wrong about your company?
Serves the wrong customer type Too small-business focused Too enterprise focused Wrong product category Old products or positioning Misses new products or features Says we lack features we offer Compares us to the wrong competitors Outdated company information Overstates what we do Understates what we do Confuses us with another company
Canonical description and offeringsOne paragraph on what the company is, plus each product or category with its own description. This is the sentence you are trying to see quoted back.
Ideal customer, in a list and in your wordsSegments from a fixed list, plus free text for the part a checkbox cannot hold. Most drift verdicts come from this field.
Key capabilities and featuresWhat AI should know you can do. Parity is judged against this, so a shipped capability listed here becomes a factual claim an answer can get wrong.
Outdated narratives and vocabularyThe story you have moved on from, the words you prefer, and the words you do not want near your name.
Competitors, with domains and notesWho you compete with, plus the comparisons you want watched closely. This is what catches a competitor feature being credited to you.
Sensitive claims to qualify or avoidRegulated language, unproven numbers, anything legal would rather you did not have an engine repeating on your behalf.
Authoritative and outdated sourcesURLs that tell the current story, and URLs that keep telling the old one. When an answer cites the second list, you know exactly why it drifted.
Most damaging mistakes to watch forPick the ones that would cost you a deal: wrong category, wrong customer, incorrect pricing, brand confusion. Those get weighted heaviest.

No profile, no verdicts. Drift and parity are measured against your fields, so Moose leaves them blank rather than inventing a yardstick.

The two verdicts

Narrative drift, and feature parity

Every answer that mentions you gets both, and each one comes with a single sentence of reasoning that has to name the specific claim it is about.

Narrative drift

Is this answer describing the company you are, or one you used to be? Wrong customer type, wrong category, dead positioning.

Yes The answer describes the brand in a way that meaningfully contradicts the profile. Wrong customer type, wrong category, or positioning you retired.
Partial Some claims line up and others drift. Common when an answer has the category right but the customer wrong.
No The description matches what you wrote. Nothing to do, which is most answers on most days.
Feature parity

Is the answer factually wrong about what you can do? This one grades misinformation. Leaving things out is fine.

Good Nothing in the answer is factually wrong about your capabilities, even if it did not list them all.
Needs review A claim is ambiguous or borderline enough that a person should look, rather than being flagged as an error.
Inaccuracy found A specific, quotable claim about what you do or do not do is wrong. The reason names the sentence.
N/A The prompt did not ask about capabilities, so there is nothing to grade and no scoring call spent.
What parity flags
Saying you do not have a capability you do have.
Claiming a capability, integration or limit you do not have.
Attributing a competitor's feature to you.
What it does not punish

Omission. If an answer stays factually correct but does not list every feature you have, that is still good. Nobody's answer lists everything.

Feature claims are only judged inside the scope of what the prompt asked about. A question about SMS is not marked down for skipping your analytics.

Restraint

Built to not cry wolf

A monitor that flags everything is a monitor you stop opening. The scoring rules are written to make false alarms the expensive mistake rather than the safe one.

The rule
A false alarm is treated as the failure

The scoring rules say plainly that false alarms make the metric useless, and that if the model cannot point to a specific in-scope sentence that is wrong or misleading, it must not flag it. Silence is the default, not the exception.

Sentiment
Neutral until someone is critical

A plain feature list, or a table row stating what you do, is neutral. Negative needs real criticism or an unfavourable contrast, not merely being mentioned last.

Evidence
An exact quote, copied word for word

Sentiment verdicts carry a quote copied from the answer word for word, so you can judge the call yourself instead of trusting a summary of it.

Failure
An empty field beats a guess

If the scoring call fails, the fields stay empty. Moose does not fabricate a verdict to fill a box, and the model is told never to invent details that are not in the response or your profile.

Scope
Judged on what was asked

Feature claims are only assessed inside the scope of the prompt. An answer about SMS deliverability is not marked down for not mentioning your reviews product.

Per engine

Which model is getting you wrong

Every observation is stamped with the engine that produced it, so drift, parity, sentiment, mentions and citations all break down per model. "Show me every answer where an AI got a feature wrong, and which one said it" is one filter.

ChatGPT
78%aligned
Two answers still place you in the old category.
Local fetch
Claude
96%aligned
Cites your own docs in four of five answers.
Managed / key
Gemini
61%aligned
Worst drift. Leans on a 2023 review blog.
Local fetch
Grok
74%aligned
Reads you as SMB only on enterprise prompts.
Local fetch
Perplexity
89%aligned
Accurate, but compares you to the wrong two.
Local fetch
Google AI Mode
84%aligned
Category right, customer segment vague.
Local fetch
AI Overviews
80%aligned
Short answers, so one wrong clause carries weight.
Local fetch
Bing Copilot
92%aligned
Mirrors whatever your site says most recently.
Managed / key
Read from the consumer product itself

For ChatGPT, Grok, Perplexity and Google AI Mode, Moose reads what the real consumer product answers from a browser session on your machine. If that fetch fails there is an API fallback, and every observation records which route produced it and whether a fallback was used. You can always tell how an answer was obtained.

Over time

One reading is trivia. The change is the story.

Runs go on your schedule: daily, weekdays, weekly or manual. What you look at afterwards is the difference between runs rather than another snapshot.

Runs on your schedule

Daily, weekdays, weekly or manual. Scheduled runs can email the summary, so the finding reaches you without you opening anything.

Trend, bucketed how you think

Mention rate, citation rate and positive sentiment charted by day, week or month, rather than one fixed window somebody else chose.

Any range against any range

Compare last month to the month before, or the two weeks either side of a launch, side by side.

Run over run, with deltas

Narrative alignment, feature accuracy, negative sentiment count, mention rate and citation rate, each with the change since the previous run.

Change highlights per prompt

Lost mention, gained mention, new citation, negative sentiment. Each one names the prompt and engine, with the quote attached.

Filter the whole report

Filter everything by narrative drift or feature parity, so "every answer where an AI got a feature wrong" is one click rather than an export and a spreadsheet.

The math

Numbers you can act on without checking them first

Most of this category quietly counts its own failures as your losses. Three decisions here are deliberate.

01
Failed checks are thrown out rather than counted as losses

If an engine could not be reached or read, that observation is excluded from every metric, delta and alert. A network failure never becomes a fake visibility drop, and a failed baseline never becomes a fake recovery.

02
Repeat sampling, with a margin of error

The same prompt and engine can be sampled up to five times per run, because a single AI answer swings wildly while the mention rate across samples holds steady. Results carry a Wilson 95% confidence margin.

03
A margin means you can tell noise from news

Two points of movement inside the margin is not a trend, and Moose will not present it as one. You get to spend your week on the changes that are real.

Your scoring allowance goes on answers that matter

If the brand was never mentioned, or the engine never answered, the observation is marked not applicable without spending a scoring call on it. Repeated samples of the same prompt collapse to one verdict for lists, emails and alerts, so a heavily cited prompt does not flood the report with duplicates.

Diagnostic prompts

Catch it on purpose

Rather than guessing what to track, Moose writes prompts anchored on your brand name and the fields you filled in: capabilities, competitor comparisons, customer segments, and the misrepresentations you already see.

Suggested for review 6 of 11 · you approve each one
Is Klaviyo only for small Shopify stores?Anchored on: misrepresentation: too small-business focused
Does Klaviyo include customer service or helpdesk tools?Anchored on: key capability: service and helpdesk
What is the best B2C CRM for a consumer brand?Anchored on: canonical description and category
Klaviyo vs Attentive for SMS and email togetherAnchored on: comparison to monitor closely
Which platform unifies email, SMS, push and reviews?Anchored on: key capabilities
Do enterprise retailers use Klaviyo?Anchored on: ideal customer segments
Anchored on your brand, on purpose

Prompts are generated from your brand name and your Brand Truth fields, so the answers they return are about you.

Your questions, word for word

Diagnostic questions you write yourself are passed through unchanged rather than rewritten into something more generic.

Generic queries excluded

Unbranded queries are left out deliberately. They do not return citations about you, which is the entire point of the exercise.

You approve every one

Review, edit and approve each suggestion before anything gets tracked. They land in their own Diagnostic Prompts category so they never muddle your core set.

Provenance

Where the wrong story came from

Knowing an engine has you wrong is half a finding. The other half is the page that taught it, and that is usually something you can go and fix.

01 The whole answer is keptFull response text, the snippet that mentioned you, and the sources the engine cited, including whether any of them was one of your own domains.
02 Matched against your outdated sourcesYou listed the pages still telling the old story. When a drifted answer cites one, the line from cause to symptom is right there.
03 Per-URL citation historyWhether a page is cited now, used to be, or has been replaced by a competitor. That history feeds the AEO audits, so a page gets audited knowing its standing.
04 From verdict to pageThis is how you get from "Gemini says we are SMB only" to the review post from 2023 that keeps teaching it, and then to the page you should be publishing instead.

Turning a finding into a fix

A verdict on its own is a complaint. In the same window, it becomes a brief, a draft, and a change you approved.

Why does Gemini keep saying we have no service tooling?
Three of its four answers on that prompt cite the same 2023 review post, which was written before the helpdesk shipped. Your own docs page for it is not cited by any engine yet. Competitors are named in six gap prompts where you are not.
Fix the page and watch it.
Brief drafted from the live results for that question, staged against your CMS, waiting for your approval. I have also set a workflow: if feature accuracy on service prompts drops again, it produces the brief without being asked. Export of this filtered view is ready as PDF or CSV.
Plainly stated

Who gets scoring

Two caveats worth reading before you download, because we would rather you find out here.

Paid managed and BYOK
Drift and parity scoring, all engines

Narrative drift and feature parity run on paid managed and BYOK workspaces that have a Brand Truth Profile filled in. Daily cadence, all eight engines, email reports, and export to PDF or CSV.

Free preview
Mentions and citations only

Free preview mode gets no scoring at all: no drift, no parity. It is capped at 3 prompts, weekly or manual runs, and the local-fetch engines. Enough to see whether you appear, not enough to grade how.

Questions people ask

Do I have to fill in the whole Brand Truth Profile?

Every field is optional, but the profile is the yardstick, so an empty one means no drift or parity verdicts at all. The canonical description, ideal customer and key capabilities do most of the work. Start with those three and add the rest as you see what gets flagged.

Does this flag an answer for leaving features out?

No. Feature parity grades misinformation, and leaving a feature out is not a fault. An answer that stays factually correct while mentioning only two of your six products is still good. It only flags a claim that is wrong: a capability denied, a capability invented, or a competitor's feature credited to you.

How do I know a verdict is not a hallucination?

Each verdict carries a one-sentence reason that has to name the specific claim it is about, and sentiment verdicts carry an exact quote copied from the answer. If the scoring call fails, the fields stay empty rather than being filled with a guess.

Which engines can it read?

ChatGPT, Claude, Gemini, Grok, Perplexity, Google AI Mode, Google AI Overviews and Bing Copilot. For ChatGPT, Grok, Perplexity and Google AI Mode there are routes that read what the real consumer product answers, with an API fallback if that fetch fails. Every observation records which route produced it.

Can a single bad answer swing my numbers?

Less than you would expect, because the same prompt and engine can be sampled up to five times per run and the results carry a Wilson 95% confidence margin. Repeated samples also collapse to one verdict in lists, emails and alerts, so a heavily cited prompt does not flood the report.

Is this available on the free plan?

No. Drift and parity scoring runs on paid managed and BYOK workspaces with a Brand Truth Profile filled in. Free preview mode is capped at 3 prompts, weekly or manual runs, and the local-fetch engines, and it gets no scoring at all.

More things to ask Moose

All use cases
Graded against your own words

Write down what is true.
Then hold every answer to it.

Fill in the Brand Truth Profile once, run a scan, and read the first nine sentences that do not match. Most teams find at least one they did not know about.