A practical SEO checklist

Robots.txt Checklist

Use this robots.txt checklist to review the file a crawler receives, test the rules against real URL paths and publish only the access changes you intend. Work through allowed and blocked cases, required resources and sitemap declarations. Robots.txt controls crawling; it does not secure private content or reliably keep a web page out of search.

Free to useNo account neededSaved in your browser
A rule sheet beside two page paths, one open and one stopped by a small barrier.
An original editorial illustration
Your working checklist

Work through the checklist

Click a square to mark a task as Pass. Click again to undo. Use the result menu for a problem, a blocker or a task that does not apply. You do not need to enter notes.

15 checks

Find the Live robots.txt File and Check Its Response

A robots.txt file is a plain-text crawl policy for the origin that serves it. An origin is the combination of protocol, host and port. Start with the public response before changing any rules.

Choose crawl controls that match the actual goal

Avoid using crawl rules to solve an indexing or privacy problem.

Use this for: Each new block or existing rule whose purpose is unclear.

For each path you plan to block, decide whether the goal is less crawling, exclusion from search, or private access. Use robots.txt for crawl control. For a public HTML page that must stay out of search, use a crawlable noindex directive. For private material, arrange authentication or access control. Neither robots.txt nor noindex keeps a URL secret.

Example

Illustrative choice: /search/?q=... creates unwanted crawl paths; /thank-you/ is a public page intended to stay out of search; an unpublished customer contract needs authenticated access. These require different controls.

  1. Name the intended outcome for each affected path.
  2. Choose crawl rules, a readable noindex, or access control for that outcome.
  3. Check that the proposed mechanism can actually achieve the stated goal.
More help and optional notesHigh · Website owner or developer
PassEvery proposed rule has a suitable crawl-control purpose and private/indexing needs use the appropriate mechanism.
FailA crawl block is being relied on to remove a public page from search or protect private content.

If it fails: Move the requirement to the correct page, response-header or access-control setting before editing robots.txt.

Retest: Check that the proposed mechanism can actually achieve the stated goal.

Not applicable: No robots.txt restriction is used or proposed, and there is no indexing/privacy requirement to classify in this review.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Open robots.txt on each relevant public origin

Find the file a crawler will actually request.

Use this for: Production hosts or protocols that serve content in the review.

List the public main host and any independent shop, blog or asset host. In a private browser window, open /robots.txt at the root of each relevant origin. A file at /blog/robots.txt cannot control a site rooted at /. Check the lowercase filename. If no crawl restrictions are needed, an intentionally absent file is not automatically an SEO defect.

Example

Illustrative scope: https://example.com/robots.txt does not control https://shop.example.com/ or http://example.com/. The shop needs its own review.

  1. List the origins that serve the pages or resources you need to review.
  2. Open the exact root /robots.txt URL for each origin.
  3. Confirm each active crawl policy is served from the origin it is meant to control.
More help and optional notesHigh · Website owner or developer
PassThe intended file is reachable at each correct root, or an intentional absence with no needed restrictions has been confirmed.
FailRules were published at the wrong host, protocol, port or subdirectory.

If it fails: Publish through the responsible host or CMS so the correct root URL serves the intended file.

Retest: Confirm each active crawl policy is served from the origin it is meant to control.

Not applicable: This applies to every public origin in the review; an intentionally absent file is checked and may Pass rather than skipped.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Check the file status and content without a login

Catch server errors, access challenges and misleading HTML responses.

Use this for: Every robots.txt endpoint in scope.

Open Chrome Developer Tools → Network, reload the robots.txt address, and select that request. Read Headers → Status Code, then Response. A deployed rule file should return 200 with its intended text. Google treats most 4xx responses, including 401, 403 and 404, as no robots restrictions; 429 and server/network failures need investigation. A 404 is acceptable only when absence is intentional, not when you depend on rules.

Example

Illustrative failure: a CDN returns a 403 challenge for /robots.txt. That is not a reliable block-all policy. Ask the host to restore the intended public response.

  1. Reload /robots.txt with Network recording.
  2. Check the status, final URL and response body.
  3. Reload after any fix and confirm the expected public response.
More help and optional notesCritical · Website owner or developer
PassThe response matches the intended file policy and has no unintended challenge, redirect loop or server error.
FailThe intended rules are replaced by an error, login, challenge or unrelated HTML.

If it fails: Fix the endpoint or CDN/host configuration; keep required restrictions in a valid file.

Retest: Reload after any fix and confirm the expected public response.

Not applicable: This applies to every robots.txt endpoint in the review, including checking an intentionally absent file’s response.

More detail (optional)

Google initially pauses crawling when it cannot fetch robots.txt because of server errors, then may use a cached version. Check its current documentation and fetch history rather than assuming a permanent block or immediate refresh.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Save a simple UTF-8 file with one rule per line

Make the intended rules readable by a robots parser.

Use this for: Sites that publish a robots.txt file.

Use your host’s robots editor or a plain-text editor. Put a User-agent line before its Allow and Disallow rules, write each directive on a separate line and save as UTF-8. Start paths with /. A # begins a comment. Google ignores unsupported fields such as noindex and crawl-delay; move indexing instructions to HTML or response headers. Keep large generated files within Google’s 500 KiB parsing limit.

Example

Illustrative lines: User-agent: * followed on a new line by Disallow: /search/. An empty Disallow: does not block the entire site; Disallow: / does.

  1. Open the source that generates the live file.
  2. Check encoding, separate lines, supported fields and path syntax.
  3. Inspect the published text again to confirm no editor markup was added.
More help and optional notesHigh · Website owner or developer
PassThe served file is readable UTF-8 with valid intended rules and no critical rule beyond the parser limit.
FailMarkup, unsupported instructions or broken syntax changes the intended crawl policy.

If it fails: Correct the generating source and simplify redundant rules before republishing.

Retest: Inspect the published text again to confirm no editor markup was added.

Not applicable: No robots.txt file is intentionally published or needed on this origin.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Test Crawler Groups and URL Matching

A rule only matters inside the group that applies to the crawler. Resolve that group first, then compare matching paths. Google uses specificity, not the position of an Allow line, to resolve rule conflicts.

Resolve the effective group for each important crawler

Find rules that apply after agent selection.

Use this for: Files with explicit crawler groups or more than one group.

Compare the crawler token with every User-agent group. For Google, the most specific matching group applies; repeated equally specific groups combine. The generic * group is not added to a specific Googlebot group. User-agent names are case-insensitive. Test Googlebot and any explicitly named crawler your site depends on; verify another engine’s own behavior before generalizing.

Example

Illustrative trap: User-agent: * with Disallow: /search/ plus a separate User-agent: Googlebot with Allow: / leaves /search/ crawlable to Googlebot. Repeat the needed restriction in its effective group or remove the unnecessary separate group.

  1. Identify the intended crawler token for each test.
  2. Collect only the rules in its applicable group or combined equal-specificity groups.
  3. Verify the effective rules against one intended allowed and one intended blocked URL.
More help and optional notesHigh · Website owner or developer
PassEach sampled crawler receives the intended complete rule set.
FailA named group silently drops a needed generic restriction or blocks an unintended path.

If it fails: Simplify the groups or add the intended rules to the correct applicable group.

Retest: Verify the effective rules against one intended allowed and one intended blocked URL.

Not applicable: No robots.txt file or crawler rule group is used on this origin.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Check the longest matching rule and Allow exceptions

Avoid broad rules overriding a more specific intention.

Use this for: URLs matching both Allow and Disallow rules.

Within the effective group, find every path rule that matches the URL. Google selects the most specific matching path rule; an equally specific Allow/Disallow conflict resolves to Allow. Moving a rule to the bottom does not make it win. Use a narrow Allow exception only when the permitted child path really needs to be fetched.

Example

Illustrative rules: Disallow: /catalog/ and Allow: /catalog/public/ permit /catalog/public/manual.html while blocking /catalog/draft.html. Equal Allow: /same/ and Disallow: /same/ permit that path for Google.

  1. Collect the matching rules for a sample URL.
  2. Compare their specificity, including any exact conflict.
  3. Verify both the intended exception and a neighboring blocked path.
More help and optional notesHigh · Website owner or developer
PassThe selected rule gives the intended outcome for exception and neighboring URLs.
FailA broad block or unexpected Allow changes access outside its intended scope.

If it fails: Narrow the rule or its exception and repeat the URL comparisons.

Retest: Verify both the intended exception and a neighboring blocked path.

Not applicable: No sampled URL matches competing Allow and Disallow rules.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Test path case, directory boundaries and exact URLs

Catch differences that visual scanning misses.

Use this for: Every path restriction, especially short prefixes.

Copy paths from actual public URLs. Path matching is case-sensitive even though directive names and user-agent tokens are not. A prefix such as /sale also matches /sales-team/. Use /sale/ for that directory, and separately decide how the slashless /sale URL should behave. Check encoded or non-ASCII URLs with a compatible parser instead of guessing from how a browser displays them.

Example

Illustrative cases: Disallow: /draft/ blocks /draft/report/ but does not match /Draft/report/ or /draft. Those alternatives need their own policy if they exist.

  1. Copy the exact paths and capitalization from representative URLs.
  2. Add slashless, neighboring-prefix and case variants that actually exist.
  3. Confirm matching results agree with the policy for every variant.
More help and optional notesHigh · Website owner or developer
PassThe rule matches intended paths and leaves intended neighboring paths accessible.
FailCase or prefix differences allow unwanted paths or block useful pages.

If it fails: Correct the path pattern and preserve any legitimate variants that need access.

Retest: Confirm matching results agree with the policy for every variant.

Not applicable: The file contains no path restriction or exception to test.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Check wildcard and end-anchor edge cases

Keep broad patterns from blocking useful URL variations.

Use this for: Rules containing * or $.

For supported Google path patterns, * matches zero or more characters and $ anchors the end of the URL. Test complete path-and-query examples. A pattern ending in .pdf$ does not match .pdf?download=1. Do not block every parameter URL simply because it has a question mark; the parameter may identify a useful page or a necessary resource.

Example

Illustrative test: Disallow: /exports/*.csv$ matches /exports/october.csv but not /exports/october.csv?download=1. This is a matching example, not a recommendation to hide files with robots.txt.

  1. Identify each wildcard or end-anchor rule.
  2. Add matching, nonmatching and query-string variants to the test cases.
  3. Check the expected outcomes with the intended crawler parser.
More help and optional notesHigh · Website owner or developer
PassTested wildcard behavior matches the intended scope, including relevant query variants.
FailThe pattern is broader or narrower than the actual requirement.

If it fails: Replace an overbroad pattern with a narrower path rule or explicit tested exceptions.

Retest: Check the expected outcomes with the intended crawler parser.

Not applicable: No path rule contains a wildcard or end anchor.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Keep Needed Pages, Resources and Sitemaps Accessible

A valid file can still implement the wrong policy. Check useful pages, their rendering resources and the sitemap destinations against the effective rules.

Allow the resources needed to render public pages

Let crawlers fetch the CSS, JavaScript and images that explain the page.

Use this for: Pages that depend on separate rendering resources.

Open an important page with Chrome Network recording and identify its stylesheet, script and meaningful image URLs. Test those exact URLs against robots.txt on their own hosts. A browser can load a resource that a compliant crawler is told not to fetch. If you block an asset directory, create a justified narrow exception or remove the unnecessary block.

Example

Illustrative failure: /assets/app.js builds the product information, but Disallow: /assets/ prevents Googlebot from fetching it. Allowing the page alone does not fix that resource block.

  1. Identify the resources needed for one important page template.
  2. Check each resource against the correct origin’s effective rules.
  3. After a fix, verify that the page and required resources are allowed together.
More help and optional notesHigh · Website owner or developer
PassRequired sampled resources are crawlable and their public responses work.
FailA required script, style or meaningful image remains blocked or unavailable.

If it fails: Remove the unintended resource restriction or restore its response, then check the dependent page.

Retest: After a fix, verify that the page and required resources are allowed together.

Not applicable: The sampled page needs no separate rendering resources; unavailable inspection access is Blocked.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Leave crawl access open where noindex must be read

Avoid preventing a crawler from seeing an exclusion directive.

Use this for: Public URLs intentionally excluded through noindex.

For each public page or PDF that relies on noindex, test whether its effective robots rules allow a fetch. If the resource is blocked, the crawler cannot retrieve the new directive. Remove an accidental crawl block only after the correct noindex is in place. Keep confidential material protected by authentication; do not expose it just to make this test pass.

Example

Illustrative conflict: /downloads/price-draft.pdf sends X-Robots-Tag: noindex, but /downloads/ is disallowed. The crawl restriction prevents the header from being read.

  1. Identify public resources whose search exclusion depends on noindex.
  2. Compare them with the applicable robots rules and verify the directive exists.
  3. Retest the corrected resource’s crawl access and its noindex output.
More help and optional notesHigh · Website owner or developer
PassEach sampled public noindex resource is fetchable by its intended crawler.
FailA robots block prevents reading a directive that the exclusion plan depends on.

If it fails: Correct the accidental crawl rule while retaining the intended noindex or private access control.

Retest: Retest the corrected resource’s crawl access and its noindex output.

Not applicable: No public resource in this scope relies on noindex for exclusion.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Use the real absolute URL in each Sitemap line

Point crawlers to the sitemap your site actually publishes.

Use this for: Sites using robots.txt for sitemap discovery.

Find the sitemap URL in the CMS or deployment output. Add a Sitemap: line with the complete HTTPS address, including the real host and filename. A sitemap index can be listed instead of every child sitemap. The line is independent of crawler groups, so it does not need to be repeated under each agent. Do not guess /sitemap.xml if the actual file has another address.

Example

Illustrative declaration: Sitemap: https://example.com/sitemap_index.xml. Sitemap: /sitemap.xml is not a complete sitemap URL.

  1. Find the published sitemap or sitemap-index address.
  2. Compare it exactly with every Sitemap line.
  3. Open each declared absolute address and confirm it points to the intended sitemap.
More help and optional notesHigh · Website owner or developer
PassEvery declaration uses the actual complete sitemap address.
FailA line is relative, points to staging, or names a nonexistent or obsolete destination.

If it fails: Replace the declaration with the actual public sitemap address.

Retest: Open each declared absolute address and confirm it points to the intended sitemap.

Not applicable: The site intentionally uses another sitemap-discovery method and publishes no Sitemap declaration.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Check the declared sitemap and its listed page access

Find contradictions between discovery and crawl restrictions.

Use this for: Sitemaps named in the file.

Open each declared sitemap and check its final response in Network. It should serve the expected sitemap content without a login or error. Test the sitemap URL against robots rules, then sample important preferred URLs from it. A useful canonical page listed for discovery should not be accidentally disallowed. Resolve the policy mismatch at the rule or sitemap source.

Example

Illustrative mismatch: the sitemap lists /guides/installation/ while Disallow: /guides/ blocks it. If the guide belongs in search, remove that accidental restriction. Listing it in a sitemap does not override the block.

  1. Fetch the declared sitemap or index and inspect its content.
  2. Check crawl access to the sitemap and a relevant sample of listed pages.
  3. Retest the sitemap and sampled page rules after resolving any mismatch.
More help and optional notesHigh · Website owner or developer
PassThe declared sitemap is accessible and sampled preferred pages have the intended crawl access.
FailThe sitemap is unavailable or robots rules contradict the listed pages’ intended use.

If it fails: Restore the sitemap response and align its membership with the intended page policy.

Retest: Retest the sitemap and sampled page rules after resolving any mismatch.

Not applicable: The robots file contains no Sitemap declaration to verify.

Keep in mind: A readable sitemap and allowed URLs do not guarantee crawling or indexing.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Test and Release the Smallest Safe Change

Keep a copy of the existing file, test intended allows and blocks, and verify the public response after publication. A local parser result proves matching behavior for that file, not that Google has fetched it.

Build an allow-and-block test set before publishing

Make rule changes reviewable before they affect the live site.

Use this for: Every proposed robots.txt change.

Copy the current file into a working text file and keep the original as a rollback copy. Build a small table of crawler, exact URL, expected allow/block and matching reason. Include the homepage, key template URLs, blocked examples, allowed exceptions and required assets. A developer can run the draft through Google’s open-source robots parser; a checker should explicitly support the intended engine’s matching behavior. Keep unrun tests Untested, or use Blocked when the needed tool or access is unavailable.

Example

Illustrative case: Googlebot, /catalog/public/manual.html, expected Allow because /catalog/public/ is the narrower exception. The worked table below shows additional boundary cases; it is not a real site audit.

  1. Keep the existing file and identify the proposed changed lines.
  2. Write representative expected outcomes and run the draft through a suitable parser.
  3. Resolve every unexplained difference before approving the candidate file.
More help and optional notesHigh · Developer
PassThe proposed file produces the expected outcomes for all chosen regression cases.
FailA tested case produces an unexpected allow or block outcome.

If it fails: Fix the smallest responsible rule and rerun the full chosen set.

Retest: Resolve every unexplained difference before approving the candidate file.

Not applicable: No robots.txt change is proposed in this review.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Publish through the source that serves production

Avoid edits being overwritten by a CMS, build or CDN.

Use this for: An authorized change with passing draft test cases.

Find whether the file is generated by the CMS, hosting configuration or deployment repository. Change that source rather than only a temporary copy. Keep the prior version available. Publish the reviewed change, refresh a stale site/CDN cache if your host requires it, and reload the public /robots.txt response. Never copy a staging-wide Disallow: / into an indexable production site.

Example

Illustrative release: the repository file is corrected, but the deployed response still shows yesterday’s block. Compare the served body with the approved candidate before marking the change complete.

  1. Confirm the production source and rollback copy.
  2. Publish the reviewed file through the normal site release process.
  3. Fetch the public file and rerun the same URL test set against its served contents.
More help and optional notesCritical · Developer
PassThe live response contains the approved version and passes the selected rule cases.
FailThe served file differs, remains stale or changes an intended allow/block outcome.

If it fails: Correct the generating source or stale response; restore the prior file if the release introduced a critical block.

Retest: Fetch the public file and rerun the same URL test set against its served contents.

Not applicable: No robots.txt change needs publication.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Check Google’s fetched version and the affected page

Separate a successful release from a completed crawler refresh.

Use this for: Sites with an applicable Search Console property.

In Search Console, open the robots.txt report and select the file. Review its last fetch, errors and Versions against your release. After an urgent unblock or fetch repair, use the file’s menu → Request a recrawl if appropriate. Then inspect an affected page with URL Inspection → Test live URL. A current browser response, Google’s cached robots file and the indexed page record can show different moments.

Example

Illustrative result: your public file is fixed, but the report’s last fetch predates the release. Keep the crawler-refresh check pending instead of claiming that Google has already processed the change.

  1. Compare the report’s fetched version and time with the public file.
  2. Request a robots recrawl only when a critical change warrants it.
  3. Check affected page access with a live test and revisit the indexed record after recrawling.
More help and optional notesHigh · Website owner or developer
PassGoogle’s reported fetched file reflects the change and the affected live URL shows the intended crawl access.
FailA fresh fetched version or live test still shows an unintended block or fetch error.

If it fails: Investigate the report’s current failure; keep unrefreshed evidence pending rather than repeatedly making new rule changes.

Retest: Check affected page access with a live test and revisit the indexed record after recrawling.

Not applicable: Search Console verification is outside this review. If it is required but property access is missing, choose Blocked.

Keep in mind: Recrawl requests do not guarantee immediate page crawling, indexing or ranking.

Optional: sample URL, expected result, observation date or a reminder. Keep confidential data out of shared exports.

No result recorded yet.

Changing this result updates your checkmark and progress immediately. Notes are optional.

Reset this checklist?

This removes saved statuses, evidence and result dates for the checks on this page. Other checklists keep their records. Export a CSV first if you need a copy.

A useful result, every time

Check a task. Choose a result.

This is a manual SEO workflow. It helps you decide what to investigate and keep a usable record. It does not crawl your website or verify your answers automatically.

  1. Choose your scope. Pick the relevant task group and representative pages you are authorized to inspect.
  2. Run the stated check. Follow the steps, then choose your result. Click the square for Pass, or use the result menu for another status. You do not need to type anything.
  3. Assign the next action. Fix a problem when you can, or ask the right person for help. Optional notes can remind you what to do next. Check the result again after a fix.

What your progress means

Review progress counts passes and failures against applicable checks. Failed work still counts as reviewed. Blocked and untested checks remain unfinished. This is not an SEO score.

Untested
You have not checked this yet.
Pass
You checked it and it works as described.
Fail
You checked it and found something to fix.
Blocked
You need access, information or someone’s help.
Not applicable
This task does not apply to your website.

Export CSV downloads the checks currently shown and any optional notes. Print expands the instructions for that view. Saved results stay in this browser; there is no account sync.

Apply the workflow

How do these robots.txt checks work on example URLs?

Apply the completed checks to these original teaching examples, then choose the next workflow from the problem that remains.

Robots.txt Testing Checklist: Work Through Rule Cases

This illustrative file and table show expected matching outcomes, not tests performed on a real website. Its duplicate Googlebot group is deliberate for teaching; a simple shared policy can usually use one generic group.

Remove the Googlebot group’s /search/ restriction in a draft and Googlebot can crawl /search/ because its specific group does not inherit the generic rules. Keep both the original and changed expectation in your regression set.

Treat an unexpected Allow as a policy failure even when the syntax is valid. In the case and query examples below, add or change a rule only if those URLs actually exist and need a different outcome.

Illustrative expected results, assuming no other rules apply
CrawlerPath and queryExpected crawl resultReason
Googlebot/catalog/draft.htmlBlock/catalog/ matches; no narrower Allow.
Googlebot/catalog/public/manual.htmlAllow/catalog/public/ is more specific.
Googlebot/Catalog/draft.htmlAllowThe path’s capital C does not match /catalog/.
Googlebot/catalogAllowThe slashless URL does not match /catalog/.
Googlebot/exports/month.csvBlockThe path ends with .csv.
Googlebot/exports/month.csv?download=1AllowThe query means the URL does not end with .csv.
Bingbot/search/?q=lampBlockThe generic group supplies the /search/ restriction.

Illustrative robots.txt fixture. Adapt it to the actual site; this is not a tested production policy.

User-agent: *
Disallow: /search/
Disallow: /catalog/
Allow: /catalog/public/
Disallow: /exports/*.csv$

User-agent: Googlebot
Disallow: /search/
Disallow: /catalog/
Allow: /catalog/public/
Disallow: /exports/*.csv$

Sitemap: https://example.com/sitemap.xml

Robots.txt Sitemap Checklist: Follow the Declaration

A Sitemap line points to a discovery file. It does not grant permission to crawl a blocked URL, choose a canonical for the engine or guarantee indexing. Keep the declaration, sitemap response and intended page access consistent.

For an illustrative sitemap index at https://example.com/sitemap_index.xml, open that address, follow a child sitemap and test a preferred page from it. A wrong host in the declaration needs a declaration fix; a blocked preferred page needs a crawl-policy decision.

Choose the Next Check from the Remaining Problem

Use this robots.txt audit checklist to finish a tested rule set. If the intended crawl policy is correct but a page is absent from search, move to its indexing directives and wider eligibility checks instead of changing working rules repeatedly.

A clear process, with honest limits.

Follow the steps, choose the result that matches what you found, and check again after changes. Notes are optional. A completed check does not guarantee indexing, traffic or rankings. How the checklists work.