Where the coverage analyzer asks whether the right things are mentioned, this one asks what is mentioned at all and whether the markup confirms it. It is the more diagnostic of the two: it hands you the list a machine would build from your page, which is frequently not the list you thought you had written. Seeing the difference is usually more useful than the score.
Here is where that sits. The median is 79/100, which ranks it 6 of the 25 measures in this dataset. The middle half of results falls between 38 and 79 — a band 41 points wide. That is wide: sites genuinely diverge here, which is to say getting ahead is still possible.
How the scores are distributed
An average on its own misleads. The chart below splits all 981 scans into five score bands, totalling 100%. The spread is wide and fairly even, which is what a measure looks like while it is still being adopted unevenly.
- 0–19
- 20–39
- 40–59
- 60–79
- 80–100
| Score band | Entity Extractor |
|---|---|
| 0–19 | 12% · 122 |
| 20–39 | 13% · 128 |
| 40–59 | 17% · 171 |
| 60–79 | 33% · 325 |
| 80–100 | 24% · 235 |
| Total runs | 981 |
The grade mix
The same data as letters: A+ 95–100, A 85–94, B 75–84, C 65–74, D 50–64 and F below 50.
Where your score stands
The number reflects how many identifiable things the page names, how varied they are in type, and whether structured data pins them down. It is one of the few scores where reading the extracted list matters more than the figure — a page can score adequately while the extracted entities reveal it is about something subtly different from what was intended.
| Tool | 10th | 25th | Median | 75th | 90th | Mean | Lowest | Highest |
|---|---|---|---|---|---|---|---|---|
| Entity Extractor | 0 | 38 | 79 | 79 | 100 | 62.5 | 0 | 100 |
How to read it: 38 is bottom-quartile, 79 is a typical page, 79 starts the upper quartile and 100 or above is the top 10%. These thresholds apply to this tool only.
The checks that fail most
The failures are about definition rather than quantity. Enough named entities, a mix of types rather than twenty instances of one, entities that recur rather than appear once, and structured data that says which specific thing an ambiguous name refers to. That last one is the difference between a page a machine can file and one it can only guess at.
| Check | Did not pass | Weight |
|---|---|---|
| Mix of entity types | 59% | Partial |
| Structured data to define entities | 43% | Partial |
| Enough named entities | 29% | Hard fail |
| Salient (repeated) entities | 26% | Partial |
“Hard fail” means the check’s full weighting is lost; “partial” means only part of the score is. Rates are the share of scans in which the check ran and did not pass.
The weakest areas
This tool groups its checks into categories. The chart below shows how often each category was the lowest-scoring one in a scan — which area is most often the weak link. It is not the average score in that category; it is how often it came last.
“Entity Diversity & Markup” was the weakest area in 50% of scans.
What to fix first
Add structured data naming the main subject of the page and linking it to an authoritative identifier where one exists. That single addition resolves ambiguity that no amount of prose can, and it converts the page from something a machine infers a topic for into something that declares its own.
A typical scan passes 2.4 of 4 checks and leaves 1.6 open items behind. What drags a score down is usually not one catastrophe but accumulated small omissions.
Run the Entity Extractor on your own page
Month by month
The figures on this page are replaced at every refresh. At the end of each month a permanent edition is published for that month and never edited again — which is the only reason this series can be compared at all. One edition has been published so far; when the second lands this section will set the two months side by side.
| Edition | Median | vs. previous | Scored 80+ | Issues/run | Runs |
|---|---|---|---|---|---|
| August 2026 | 79 | baseline | 24% | 1.6 | 981 |
Frequently asked questions
Why does entity type variety matter?
Because a subject is normally made of several kinds of thing — a person at an organisation, in a place, working on a product, under a standard. A page naming only one type usually covers only one dimension of its subject. Variety is a proxy for having described the context rather than just the thing.
Do I need Wikidata or schema.org identifiers?
They are not required, but they are the strongest available way to remove ambiguity. A sameAs pointing at an authoritative record turns "we mean this specific company" from an inference into a statement. For any name shared by more than one well-known thing, it is worth the few minutes.
How does this differ from the Entity Coverage Analyzer?
This one reports what is there; the coverage analyzer judges whether what is there is enough for the subject. Use this to check that a machine reads your page as being about what you think it is about, and the coverage analyzer to find what you left out.
What is a good score on the Entity Extractor?
In this dataset the median is 79/100, so scoring 79 puts you exactly in the middle. 79 is the start of the upper quartile and 100 or above is the top 10%; below 38 is the bottom quarter. Those thresholds are specific to this tool — the 25 measures in this dataset have very different medians, so 70 here and 70 elsewhere do not mean the same thing.
Which check fails most often on the Entity Extractor?
“Mix of entity types”, which did not pass in 59% of scans. The costliest failures, though, are the ones scored as hard fails — this tool has 1, the most common being “Enough named entities” at 29%.
Questions that govern the whole series — how the sample is collected, how sites are anonymized, how often the figures are refreshed — are answered in the methodology on the hub.
Using this data
The figures, tables and charts on this page are free to reuse under CC BY 4.0; the only condition is a link back to this page. Suggested citation: SeoMods — Entity Extractor benchmarks, 2 September 2026, https://seomods.com/seo-score-benchmarks/entity-extractor. If you need to cite by date, use the August 2026 edition instead: this page changes at every refresh and that one never does.