Where the coverage analyzer asks whether the right things are mentioned, this one asks what is mentioned at all and whether the markup confirms it. It is the more diagnostic of the two: it hands you the list a machine would build from your page, which is frequently not the list you thought you had written. Seeing the difference is usually more useful than the score.

Here is where that sits. The median is 79/100, which ranks it 6 of the 25 measures in this dataset. The middle half of results falls between 38 and 79 — a band 41 points wide. That is wide: sites genuinely diverge here, which is to say getting ahead is still possible.

How the scores are distributed

An average on its own misleads. The chart below splits all 981 scans into five score bands, totalling 100%. The spread is wide and fairly even, which is what a measure looks like while it is still being adopted unevenly.

  • 0–19
  • 20–39
  • 40–59
  • 60–79
  • 80–100
Entity Extractorn = 981
Entity Extractor score distribution: how 981 scans spread across five score bands, median 79/100
Entity Extractor — score distribution, 8 August 2026 – 31 August 2026. Free to reuse under CC BY 4.0 with attribution.
Score bandEntity Extractor
0–1912% · 122
20–3913% · 128
40–5917% · 171
60–7933% · 325
80–10024% · 235
Total runs981

The grade mix

The same data as letters: A+ 95–100, A 85–94, B 75–84, C 65–74, D 50–64 and F below 50.

Where your score stands

The number reflects how many identifiable things the page names, how varied they are in type, and whether structured data pins them down. It is one of the few scores where reading the extracted list matters more than the figure — a page can score adequately while the extracted entities reveal it is about something subtly different from what was intended.

Entity Extractormedian 79
0255075100
Tool10th25thMedian75th90thMeanLowestHighest
Entity Extractor038797910062.50100

How to read it: 38 is bottom-quartile, 79 is a typical page, 79 starts the upper quartile and 100 or above is the top 10%. These thresholds apply to this tool only.

The checks that fail most

The failures are about definition rather than quantity. Enough named entities, a mix of types rather than twenty instances of one, entities that recur rather than appear once, and structured data that says which specific thing an ambiguous name refers to. That last one is the difference between a page a machine can file and one it can only guess at.

Mix of entity types59%
Structured data to define entities43%
Enough named entities29%
Salient (repeated) entities26%
The checks that fail most often on the Entity Extractor, with each failure rate; “Mix of entity types” leads at 59%
Entity Extractor — failed checks by rate, from 981 scans.
CheckDid not passWeight
Mix of entity types59%Partial
Structured data to define entities43%Partial
Enough named entities29%Hard fail
Salient (repeated) entities26%Partial

“Hard fail” means the check’s full weighting is lost; “partial” means only part of the score is. Rates are the share of scans in which the check ran and did not pass.

The weakest areas

This tool groups its checks into categories. The chart below shows how often each category was the lowest-scoring one in a scan — which area is most often the weak link. It is not the average score in that category; it is how often it came last.

Entity Diversity & Markup50%
Entity Richness50%

“Entity Diversity & Markup” was the weakest area in 50% of scans.

What to fix first

Add structured data naming the main subject of the page and linking it to an authoritative identifier where one exists. That single addition resolves ambiguity that no amount of prose can, and it converts the page from something a machine infers a topic for into something that declares its own.

The hard fails on this tool. Each of these 1 checks costs its entire weighting when it does not pass, where everything else on the list costs only part of the score. The most common is “Enough named entities”, unmet in 29% of scans.

A typical scan passes 2.4 of 4 checks and leaves 1.6 open items behind. What drags a score down is usually not one catastrophe but accumulated small omissions.

Run the Entity Extractor on your own page

Month by month

The figures on this page are replaced at every refresh. At the end of each month a permanent edition is published for that month and never edited again — which is the only reason this series can be compared at all. One edition has been published so far; when the second lands this section will set the two months side by side.

EditionMedianvs. previousScored 80+Issues/runRuns
August 202679baseline24%1.6981

Frequently asked questions

Why does entity type variety matter?

Because a subject is normally made of several kinds of thing — a person at an organisation, in a place, working on a product, under a standard. A page naming only one type usually covers only one dimension of its subject. Variety is a proxy for having described the context rather than just the thing.

Do I need Wikidata or schema.org identifiers?

They are not required, but they are the strongest available way to remove ambiguity. A sameAs pointing at an authoritative record turns "we mean this specific company" from an inference into a statement. For any name shared by more than one well-known thing, it is worth the few minutes.

How does this differ from the Entity Coverage Analyzer?

This one reports what is there; the coverage analyzer judges whether what is there is enough for the subject. Use this to check that a machine reads your page as being about what you think it is about, and the coverage analyzer to find what you left out.

What is a good score on the Entity Extractor?

In this dataset the median is 79/100, so scoring 79 puts you exactly in the middle. 79 is the start of the upper quartile and 100 or above is the top 10%; below 38 is the bottom quarter. Those thresholds are specific to this tool — the 25 measures in this dataset have very different medians, so 70 here and 70 elsewhere do not mean the same thing.

Which check fails most often on the Entity Extractor?

“Mix of entity types”, which did not pass in 59% of scans. The costliest failures, though, are the ones scored as hard fails — this tool has 1, the most common being “Enough named entities” at 29%.

Questions that govern the whole series — how the sample is collected, how sites are anonymized, how often the figures are refreshed — are answered in the methodology on the hub.

Using this data

The figures, tables and charts on this page are free to reuse under CC BY 4.0; the only condition is a link back to this page. Suggested citation: SeoMods — Entity Extractor benchmarks, 2 September 2026, https://seomods.com/seo-score-benchmarks/entity-extractor. If you need to cite by date, use the August 2026 edition instead: this page changes at every refresh and that one never does.