One code cannot establish usable evidence

A server can return 200 for a challenge screen, login page, JavaScript container or custom error. A CDN can also serve stale content. Monitoring reports success while the requester receives something other than the article. Conversely, visible text after script execution does not prove that the initial HTML contained it.

Separate network response, original body, browser rendering, crawl permission, index state and observed citation. Later stages do not follow automatically from earlier ones. Resolve network errors first, empty output at the generation layer next, and canonical or indexing directives afterward.

Begin with an actual response record

Save requested and final URLs, redirect chain, status, content type, body size and time. Identify redirects to home, login or region restrictions. A HEAD response cannot establish that article text is present because it intentionally contains no response body.

Inspect initial HTML for a unique heading, substantive text, canonical URL and language. Check both robots metadata and X-Robots-Tag headers. Record real device and regional differences; do not “repair” access by serving crawlers different facts from those shown to readers. The crawler-log guide covers identity and request evidence.

Run four local counterexamples

assessment-results.json records complete HTML, an empty main element, an added noindex directive and duplicated identifiers. The checker reports the corresponding problems. It parses local strings and makes no commercial AI requests.

The main element is a convention in this fixture, not a universal requirement. Production checks should use the site's actual article container so that valid pages without main are not classified as empty. This parser does not execute JavaScript, validate CSS visibility or simulate Google's rendering. Browser and live-response checks are separate evidence.

Compare initial HTML with rendered content

Capture the original response, then use a browser to wait for the content container to stabilize. Compare title, substantive text and critical facts. If evidence appears only after interaction, establish whether the intended discovery path reaches it. Test overlays, transparent text, clipping and failed lazy loading; a DOM string is not necessarily visible.

An explicit checkpoint might require AX220, 20 L and the indoor restriction together. A visible title with a missing limitation is a factual completeness problem, not merely visual polish. A real deployment needs desktop, mobile, script-failure and cache-refresh checks; the downloadable parser makes no claim to cover those states.

Distinguish an outage from indexing delay

Empty responses, error content and security blocks require access repair. If body content, directives and canonical relationships are valid but indexing remains absent, consult the platform's crawl and indexing reports. Repeated submission is not a guarantee.

Google's HTTP status documentation explains crawl-related response handling, including the difference between successful HTTP delivery and usable content. Do not return homepage text with 200 for every missing article. Use appropriate error status or an intentional migration redirect.

Zhihe Growth's release acceptance boundary

Zhihe Growth treats status, readable text, key facts, language, canonical URLs, internal links, downloads and mobile display as separate checks. Retrieve public pages after deployment; local preview success is not production acceptance. Retain a rollback path and release record when content or assets fail.

These repairs improve access conditions, not proven citation frequency. Observe citations through the platform matrix and fixed question protocol. Protected client records and administrative areas should remain authenticated rather than opened to improve crawl statistics.

Diagnosis exercise: different failures under the same 200 status

Initial response Browser observation Candidate cause Read-only next check
Empty container Text after scripts Client rendering Data and rendering conditions
Login prompt Logged-in user sees content Accidental authentication Logged-out retrieval
Old article Still stale after refresh Server or CDN cache Compare response and package hashes
Correct body Invisible text Styling failure Computed styles and screenshot
Error message Normal page shell Soft error Confirm target existence
Missing specification Other content visible Build omission or truncation Compare critical fields

These are hypotheses, not diagnoses from appearance alone. Preserve evidence and isolate changes instead of altering cache, templates and security simultaneously. Do not default to disabling protections.

Check content type and bytes too: a PDF URL returning an HTML error with 200 is not a successful download. Validate file signatures or parsing and hashes. Matching a brand name is insufficient because an error page may retain the same navigation.

Materials and limitations

Download the HTML fixtures and checker, run assessments.py and inspect html_checks. This is an access-diagnosis lesson, not a complete crawler or security-bypass utility. Google's AI feature guidance supplies further indexing context.

Knowledge center · GEO services · Research and evidence