A crawler-drafted description of your organization is only trustworthy once a person has stood behind it. That makes the confirmation the real work — and most tools skip it. Here is what building it in, rather than assuming it, actually looks like.
In a previous essay I argued that extraction is nearly solved, and that this is not the good news it sounds like. Point a competent pipeline at a website and it will produce candidate entities and relationships in minutes. The question it answers — can this fact be extracted? — is close to settled. The question that matters — should this fact be published as an authoritative statement of what the organization is? — it does not touch at all.
That gap is not specific to entity-map files. It applies to any machine-readable account of your organization, including the one an AI system assembles on its own by reading your pages at the moment someone asks about you. Whatever the format, the same question sits underneath it: who confirmed this?
Most tools never make you answer. They hand you a draft and let its completeness stand in for its truth. This essay is about the step that comes after the draft — the one that turns a machine assembled this into someone has stood behind this — and what it takes to build that step in rather than assume it.
Keep two questions apart
The first discipline is to refuse to collapse two things that look similar and are not.
The first is what a site establishes on its own — the entities, services, and relationships a crawler can find in the visible evidence, with no help from you. Call it the independent baseline. It is a measurement of what the page actually says to a machine, and it is worth having precisely because it owes nothing to your intentions.
The second is what the organization means to be known for — the associations it has decided are strategic and is prepared to stand behind. These are not the same list. A site establishes a great many things by accident: a location it mentioned once, a service named in passing, a partner who left two years ago and lingers in a bio. Discovery finds all of it. None of it is automatically a priority.
The mistake to design against is letting discovery masquerade as intent — treating everything found as everything meant. A tool that assumes the two are identical will confidently optimize a site around whatever it happened to surface, which is not the same as what the business would choose. So the two questions stay apart: what the crawler establishes, and what the organization confirms. The distance between them is the thing worth looking at.
The tool doesn’t decide what you mean
If the baseline is what a machine finds and the strategy is what a business intends, then someone has to supply the intent. The tool cannot.
This is where the confirmation step becomes concrete rather than rhetorical. When strategic associations are supplied, the analysis measures how strongly the site currently establishes them — how well the visible evidence supports the things you have said you care about. When they are not supplied, the tool can propose candidates for review, but it does not promote them on its own. A discovered entity is a suggestion until a person confirms it is strategic. The human is not decoration on the workflow; the human is the step.
That is the direct answer to the question the previous essay ended on. “Who checked?” stops being an embarrassing blank when checking is a required move in the process — when the path runs from crawler-discovered candidates, to business-confirmed priorities, to a targeted audit of those priorities, to findings, to actions, and only then to a rerun that measures whether anything changed. Confirmation is not a courtesy at the end. It is the hinge the whole sequence turns on.
“Weak” is a diagnosis, not a verdict
Once you know what a site should establish and can measure how well it does, you still have to say something useful about the gap. “Weak” is not useful. It is a grade, and a grade tells an implementer nothing about what to do.
The work is in the handoff. For anything that needs action, a technical implementer should get the current state, why the issue matters, the specific observed gap, what existing evidence to preserve, the actual instructions, the pages involved, the supporting crawler evidence, acceptance criteria for calling it done, a warning against overdoing it, and a protocol for rerunning the same scope afterward. That is the difference between a report that assigns blame and one that can be executed.
And the recommendation has to fit the real failure, which is harder than it sounds, because the reflex fix is often wrong. Suppose a service already has an explicit relationship to the company, prominent naming on its page, and valid structured data — but appears on only one crawler-visible page. The reflex is “add more schema.” More schema would do nothing; the schema is fine. The real problem is that the site asserts the service once and never reinforces it anywhere else, so a machine sees a single, unsupported mention. The fix is breadth and factual reinforcement, not markup. Misread that, and you generate busywork at best and, at worst, push a site toward the kind of over-optimization these systems increasingly discount.
A worked example
A firm offers a service that matters to its business. On the one page describing it, everything is correct: the service is named clearly, its relationship to the firm is explicit, the structured data validates.
A scan flags the service as weakly established. The reflex reading is that something is missing from the markup, and the reflex fix is to add more of it.
But nothing is missing from the markup. The service is simply asserted in exactly one place, with nothing anywhere else on the site to corroborate it. To a machine assembling a picture from repeated, connected evidence, one clean mention is a thin signal — not because it is wrong, but because it stands alone.
The problem was never the schema. It was that the site said it once and never again.
What it measures, and what it does not
It is worth being exact about what a diagnostic like this can claim, because the field is full of promises it cannot keep.
The tool measures whether a site’s crawler-visible information gives search and AI systems a clearer foundation from which to retrieve and understand the organization and its relationships. Its outputs — measures of clarity, coverage, relationship strength, connectivity, structured-data alignment, retrievability, and how well the site supports the associations a business has confirmed — are internal diagnostic measures. They describe the input a machine has to work with.
They are not ranking factors. They are not citation scores. They do not predict, prove, or guarantee that ChatGPT, Gemini, Perplexity, Google’s AI Overviews, or any other system will surface, cite, or recommend you. No honest tool can offer that, because the selection logic is private, changes constantly, and depends on far more than any one site controls. What a diagnostic like this offers is control over the input — the clarity, consistency, and retrievability of the evidence these systems find — and nothing on the output side.
That is a smaller claim than the market usually makes, and it is the true one.
The part that stays yours
A machine can now draft a description of your organization in minutes, and the draft will look finished. Deciding whether it is true — which discovered facts you will stand behind, and which you will not — is the part that has not been automated and should not be. In an environment where AI systems increasingly summarize an organization before anyone visits it, confirming what that summary is built from is no longer a technical nicety. It is how you keep authorship of your own identity.
The crawler will always tell you what it found. Only you can say what you meant.