Build an entity graph
of your own site
Engines do not read your site page by page. They resolve it into things: people, products, organisations, topics, and the relationships between them. Ask Moose to build that graph and you see the version of you an engine can actually cite.
On it, Sam. Let's map out how Fathom's entities connect across the site.
I can build an entity graph for this project. Choose how deep the run should go, then start it below.
Hi, Moose maps the people, products, places, and topics across this project's crawled pages, plus how they connect.
Two hundred and ninety entities, three relations. The site names a lot of things and connects almost none of them. Every relation I found came from schema markup on two templates. The rest of the graph is loose nouns.
Two things stand out. Fathom Analytics resolves as two separate entities, once as a product and once as an organization, with nothing tying them together. And User Documentation sits on 246 pages with the highest salience on the site, so by volume the docs are the brand.
Want me to run the Full pass over the top 100 pages, or draft the schema that merges the two Fathom entities first?
Chief desk supervisor. Resolves as exactly one entity, related to the sofa, the garden, and lunch. A cleaner graph than most websites manage.
An engine does not see pages. It sees things, and how they relate.
Before an engine can recommend you, it has to know what you are. A product, made by an organisation, founded by a person, in a category, competing with these others. That resolution happens whether you help it or not.
If your site never states the connections, the engine either infers them from third parties or leaves you as a floating name it cannot place. That is how you end up mentioned but never recommended, or described using a competitor's category.
Four things, in one pass
Every entity carries the pages it appeared on, so you can always get from a claim back to the URL that made it.
People, organisations, products, places and topics, each with the number of pages it appears on. Types are counted, so you can see at a glance whether your site reads as a product or as a pile of articles.
A weighted score per entity, from how often and how prominently it appears. This is where sites get a surprise: the entity with the highest salience is rarely the one on the homepage.
Founder, publisher, brand, part of, operating system. Only the ones your markup or copy genuinely states. A low number here is the most common finding, and the most fixable.
Entity splits, like a brand appearing once as a product and once as an organization with nothing tying them together. Engines that see two weak entities cite neither.
Quick, then Full when it earns it
Quick reads only what your site already declares, which is the honest baseline. Full adds a model pass and finds the entities you talk about without ever marking up.
Reads the JSON-LD, microdata, headings and internal link structure already on your pages. Runs instantly, costs nothing, and shows you exactly what an engine can extract without guessing.
A model reads the body text and pulls the entities and relations you state in prose but never mark up. Takes a few minutes. Runs on your own key, or entirely on a local model.
The graph is a starting point, not a report
Once it exists, it stays in the project. Every one of these runs against it, in the same chat window.
The way this normally gets done
A validator, a crawl export, a whiteboard diagram nobody updates, and a strong opinion about schema from someone who read a blog post in 2019.
A validator checks syntax on one page. It cannot tell you that your product and your company are two disconnected entities across the site.
Spreadsheets count text. Deciding that two spellings are the same organisation, and that a third is a different one, needs a resolution step nobody does by hand consistently.
Whiteboard graphs describe how the brand wishes it were structured. Engines read what is on the pages, which is usually thinner and messier.
Without a versioned graph you cannot show that a schema change worked, so entity work stays the thing that gets cut from the sprint.
Ask once, get the graph and the fix list
The crawl is already on your machine from indexing the site, so the graph is built from pages you own, not from a third-party index of them.
No new screen to learn. Ask for an entity graph, pick Quick or Full, and it runs against the pages already crawled for this project.
The counts are the point. Which type dominates, which entity carries salience it should not, and how few relations you actually declare.
Ask for the markup that merges a split entity or states a missing relation. Moose drafts it, stages it in your CMS, and waits for you to approve.
Quick runs entirely offline against your local crawl. Full sends page text to whichever model you chose, and if that model is local it never leaves the machine either. The graph is written to local storage and versioned, so you can diff it after a schema change.
Why the numbers are low on purpose
A graph that invents relations is worse than no graph, because you would ship markup asserting things your site never said. Three real relations beats forty plausible ones.
Quick reads only what your pages state: JSON-LD, microdata, headings, internal links. That is the same material an engine gets for free, which makes it the honest baseline to fix against.
Name variants collapse into one entity when the evidence supports it. When it does not, you get two entities and a note saying they may be the same thing, rather than a silent guess.
If no page says who founded the company, the graph has no founder relation. It reports the absence, because that absence is the finding you can act on.
Position matters as much as frequency. An entity in a title and a heading outweighs one buried in a footer on every page, which is why nav links do not dominate the ranking.
Each entity carries the URLs it came from, so you can check any row yourself. Nothing in the graph exists without a page behind it.
Each run is kept. Ship a schema change, re-run Quick, and diff the counts. That is the difference between entity work and entity opinions.
From the graph to the markup
A diagram you cannot act on is a poster. This one turns into staged changes without leaving the chat.
Questions people ask
What is an entity graph, in plain terms?
A list of the things your site talks about, and the stated connections between them. Not keywords, and not pages. An engine builds one of these about you whether you look at it or not, so the useful move is to see yours and decide whether it says what you meant.
Is 290 entities and 3 relations bad?
It is normal, and it is the reason to run this. Most sites name hundreds of things and explicitly connect almost none of them, because relations only come from markup and most markup stops at the article template. The fix is usually a handful of lines on two templates.
Do I need schema markup for this to work?
Quick depends on it, which is precisely why the result is useful: it shows you what an engine can extract today. Full reads the prose too, so a site with no markup still gets a graph, and the gap between the two runs tells you what to mark up first.
Does the Full pass cost me anything?
It uses whichever model you have connected, so it is your key and your spend, capped at the top 100 pages by default. Point it at a local model and it costs nothing but time. Quick is always free and always offline.
Does any of this get uploaded?
Quick never leaves your machine. Full sends page text to the model you chose, and if that model runs locally then nothing leaves either. The graph is stored locally and versioned, so there is no Hi, Moose copy of your site structure.
How is this different from a coverage check?
Coverage answers whether you have written about a topic. The graph answers what your site says you are. You need both: one stops you writing duplicates, the other stops engines describing you as a category you do not sell in.
More things to ask Moose
All use casesAsk before you brief, and get the pages, the verdict and what is missing.
Ask why traffic moved and get an answer with the tables under it.
The searches Google writes for itself before it answers, and who it cites.
Twelve searches behind one prompt, including the ones naming brands it trusts.
One question, five real engines, about a minute. Free on every plan.
The full list, and what Moose does step by step in each one.
See the version of you
an engine can cite.
Download Hi, Moose, index your site, and build the graph in about a minute. Quick runs on the free plan.