Technical audits usually open with page speed, because speed is easy to measure and looks good in a report. It is a poor place to start. A page that is not indexed does not become more useful for loading in a second. The order of checks should follow the order in which a search engine actually reaches a page: find it, crawl it, index it, rank it. What follows is a checklist in that order.
Why order matters more than completeness
A hundred-line audit handed over as a list almost always produces nothing. A developer starts at the top, and the top is usually something cheap and harmless: image compression, cache headers. The problem keeping a third of the site out of the index sits at line seventy-four.
So an audit is more useful run as a funnel. At each step there is one question: how many pages were lost here. The step with the largest losses is the work for this month.
Step 1. What is actually in the index
The first number to get is how many of the site's pages are indexed against how many should be. The second comes from the sitemap or a CMS export; the first from the webmaster console.
The gap reads like this:
- Substantially fewer indexed than expected: a crawl or a page-quality problem. This is the main scenario, and the rest of the funnel is about it.
- Substantially more, and things are in the index that should not be: filter pages, sort orders, internal search results, utility URLs. Also a problem, just the inverse.
- If the numbers agree the site is probably technically healthy, and the answer lies in content and links.
Look separately at the exclusion reasons the console names itself. For projects working in the Russian market this matters: Yandex and Google exclude pages for different reasons, and a site can be fully indexed in one and half-indexed in the other.
Step 2. Sitemap and robots.txt
Two files that break more often than people expect, and break quietly.
For the sitemap: does it return 200, does it list only URLs that themselves return 200, does it carry lastmod, and is that field set to the build date. The last one is everywhere and it does damage: if every deploy claims every page changed, the engine stops trusting the field at all.
For robots.txt: is the sitemap path declared, are there sections still disallowed out of habit, are static files and scripts blocked. Blocked CSS and JS is an old mistake that still turns up, and it stops the engine seeing the page the way a person does.
Check both on the live URL rather than in the project source. Hosting rewrites paths and serves something other than what sits in the repository, and a gap between what was built and what is served only shows up when you request the real URL.

Step 3. Status codes and redirect chains
Crawling the site gives a distribution of status codes. Healthy looks like this: an overwhelming majority of 200s, a small share of 301s on legacy URLs, zero 5xx.
What needs attention:
- Redirect chains longer than one hop. Every link in the chain is crawl budget spent. Four-hop chains show up on sites that have migrated twice.
- Redirects listed in the sitemap. A sitemap should carry destinations, not waypoints.
- Soft 404s, meaning a page returning 200 that says nothing was found. To a search engine that is a real page, and they accumulate.
- 5xx under load. Worth testing at peak hours.
Step 4. Canonicals and duplicates
Duplicates are the most common reason an index holds twice as many pages as the site has.
The minimum set of checks: does the homepage resolve to one URL (trailing slash, www, http and https all collapsing into one), do filter and sort pages canonicalise to the base category, does the canonical point at a redirect, does it match what the sitemap says.
Pagination deserves its own note. Page two of a listing should canonicalise to itself, not to page one, or the engine loses access to products that exist only there.
Step 5. Speed and Core Web Vitals
Now that the page is found and indexed, speed is worth discussing. Google publishes three metrics and their "good" thresholds: LCP within 2.5 seconds, INP of 200 milliseconds or less, CLS of 0.1 or less (web.dev/articles/vitals, checked 19 September 2026). INP replaced FID and became a stable metric in 2024, so any guidance still talking about FID is out of date.
The practical part people skip: lab measurements and field data are different things. A lab test on a fast connection will report an excellent result for a site that takes four seconds for real users. Make decisions on field data and use the lab test to find the cause.
Of the three, CLS is usually the cheapest and most visible fix: images and ad slots without declared dimensions, so the page jumps as it loads. That is an hour's work.
Step 6. The mobile version
Indexing has run on the mobile version for years, so that is the one to check, and to check it for content rather than layout.
The question that matters: do the text, the structured data, and the set of links match between mobile and desktop. A trimmed mobile version that hides half the copy and drops the link block is the version the search engine sees.
Then the smaller things: text readable without zooming, controls big enough to hit with a thumb, no horizontal scroll, no consent banner covering the content.

Step 7. Structured data
Structured data changes how the snippet looks, and that affects clicks. It does not raise rankings on its own.
Worth checking: is the markup valid, does it describe what is actually on the page, are there duplicate blocks of the same type. The last one breaks rich results more often than a syntax error does: two FAQ blocks on a page and the engine shows neither.
And the rule that costs more than the rest: markup must not assert what is not on the page. A rating in the markup with no reviews on the page is grounds for a manual action.
Step 8. Internal links
The last step, and it explains most cases of "the page exists but is not indexed".
Look at depth: how many clicks from the homepage to a typical page. Three is fine; five or more and the page is effectively out of reach. Look for orphans: present in the sitemap, with no internal link pointing at them. Look at anchors: if every link says "read more", the engine gets no signal about what the destination is about.
Orphans and internal linking are fixed with structure rather than a list of edits: pages should group into topics, and topics should link to each other. Where those groups come from is covered in our piece on keyword clustering.
What does the work, and what it costs
Webmaster consoles are free and mandatory: without them step one cannot be done at all. Paid tools are for crawling the site and for comparing against competitors.
| Task | Tool | Through gbseo.ru | Retail |
|---|---|---|---|
| Crawl, status codes, duplicates | 499₽/mo | 10 000₽/mo | |
| Audit, internal links, orphans | 1 999₽/mo | 9 000₽/mo | |
| Parsing logs and exports | 399₽/mo | 1 800₽/mo | |
| Diagrams and the client report | 499₽/mo | 1 200₽/mo |
Prices from the gbseo.ru catalogue as of 19 September 2026; "retail" is the vendor's own price for the same subscription. The full set runs 4 999₽ a month against 36 700₽ at retail. The calculator on the homepage prices any subset.
A word on language models in an audit. They are useless where a site has to be crawled and genuinely useful where the crawl output has to be read: a forty-thousand-row export you need a pattern out of is exactly the work a person does slowly and badly.
Setting the order of work
A crawl produces hundreds of lines. One table sorts them: how many pages are affected, and what the page loses.
| Priority | What it usually is |
|---|---|
| First | Pages not indexed; 5xx; blocked sections; duplicates at scale |
| Next | Redirect chains; soft 404s; orphans; crawl depth |
| After that | Core Web Vitals; structured data; anchor text |
| Eventually | Heading cosmetics, alt text on decorative images |
An audit that does not end in an order of work is an inventory. Turning the result into something a client can read is covered in our write-up on the working stack.
FAQ
How often should a technical audit run?
A full one every six months, and always after a migration, a CMS change, or a redesign. A short indexation-and-status-code check is worth running monthly: it takes half an hour and catches most outages.
Does a small site need a paid crawler?
For fifty pages, no: the webmaster consoles are enough. A paid tool starts paying for itself somewhere in the hundreds of pages, where reading them by eye stops being possible.
What if Yandex indexes the site and Google does not?
Check whether robots.txt and status codes differ per crawler, and read the exclusion reasons in Search Console separately. A split between engines almost always means something specific to one of them rather than a general technical fault.
Worth fixing Core Web Vitals if rankings are already decent?
CLS, yes: it is cheap and users notice. LCP and INP on a site that already ranks are worth touching after indexation and duplicates are closed: the gain is smaller and the work is dearer.
How long does a first audit take?
A site of a few thousand pages takes two or three days including reading the export and setting priorities. Most of that time goes into turning a list of problems into an order of work, not into the crawl itself.
