Methodology

How we grade a site — and why our number differs

Every grader invents its own scale. Ours is a readiness score: the share of weighted, observable checks your site passes, measured on the raw HTML an AI assistant or crawler actually receives. Here is every part of it, in the open.

The 100 points

The headline grade is the sum of four categories. AI readiness is scored separately out of 100 so a strong classic-SEO site can't hide a weak AI profile inside one blended number.

Performance

30 pts

What we look at: Server response time (TTFB), page weight, how much of the page depends on JavaScript to render, whether the page is reachable without redirect chains, and internal link depth.

Why it counts: Assistants and crawlers fetch raw HTML with a short budget. A page that only paints after hydration is effectively empty to them.

SEO

30 pts

What we look at: Title and description quality, title uniqueness across crawled pages, a single H1, canonical tags, sitemap.xml, robots.txt, Open Graph/Twitter cards, heading structure, indexability and image alt text.

Why it counts: These are the classic signals that decide whether a page can be indexed and shown at all.

Mobile & readability

30 pts

What we look at: Viewport meta, declared language, semantic landmarks (main/article/nav), body content volume, answer-shaped paragraph length, and correct 404 status codes.

Why it counts: The same structure that helps a phone reader helps a model find the answer block worth quoting.

Security

10 pts

What we look at: HTTPS with a correct redirect from HTTP, HSTS header, and a published security.txt.

Why it counts: Low weight, high signal — it is cheap to fix and it is a trust marker for both browsers and buyers.

AI readiness

separate /100

llms.txt, AI crawler permissions in robots.txt, JSON-LD presence, Organization and FAQ schema, consistent entity naming, question-and-answer formatting, answer-shaped meta descriptions, cited evidence, content freshness, an about page, visible pricing signals and comparison content — plus what live assistants say when asked about your brand.

How a category score is calculated

Each check carries a weight from 1 to 5 based on how much it changes whether a machine can find, parse and repeat your content. A category score is:

category points = (sum of weights passed / sum of weights evaluated) x category max
  • • Checks we cannot evaluate are marked unknown and dropped from both sides of the fraction — they never silently cost you points.
  • • Checks are run against every page in the crawl (homepage plus up to 12 internal links from your sitemap), and site-wide checks like robots.txt are evaluated once.
  • • Anything derived from a live model answer is labelled Estimated; anything we observed directly in your HTML or headers is Measured.

Why HubSpot says 93 and we say 60

Both numbers can be correct, because they are answers to different questions. A grader's score is only meaningful against its own scale, so compare yourself over time and against rivals inside one tool — never across two.

Different question

Classic graders ask 'is this a well-built marketing site for a human visitor?'. We ask 'can a machine fetch, parse, understand and safely repeat this site?'. Passing the first does not imply passing the second.

Different check set

We run 40+ checks, and roughly a third of them — llms.txt, AI crawler rules, entity consistency, answer formatting, evidence, comparison and pricing content — simply do not exist in a 2015-era SEO grader. A site can be perfect on the older list and still fail a lot of ours.

Different scope

Many graders score one URL. We crawl the homepage plus up to 12 internal pages and average site-wide, so one immaculate homepage cannot carry a thin site.

Different generosity

Some scales award most of the points for simply existing (a page that loads, on HTTPS, with a title). Ours starts at zero and pays only for checks passed, which pushes typical sites into the 50-75 band rather than the 85-95 band.

Different rendering

We read the raw HTML response, not the hydrated DOM, because that is what most crawlers and assistants see. Heavily client-rendered sites score lower with us on purpose.

What the score is not

  • • It is not a traffic, ranking or revenue forecast.
  • • It is not a measured citation log — assistant answers are samples from a live model, and models vary between runs and update their knowledge on their own schedule.
  • • It is not a Core Web Vitals lab test. Our performance signals come from response timing and payload, not from a full browser trace with field data.
  • • It is not comparable to another vendor's score. It is comparable to your own last run, and to your rivals scanned the same way.

See your own numbers

Run the grader free, then check each failing item against the definitions above.

Newsletter

The AI visibility briefing

One email a fortnight: what changed in how assistants pick sources, the checks that moved scores most, and a template you can ship the same day.