Guide 11 of 149 minIntermediate

Configuration, Scale and Reporting

Indexation & Crawler Settings: Control What Google Sees

Configure how the crawler treats robots.txt, noindex tags, canonical signals and crawl budget so audit results reflect what search engines actually see rather than raw HTML output.

Want to earn a Rankar Academy certificate?

This guide is free to read in full, with nothing to sign up for. Join the Academy to track your learning, complete courses and sit the certification assessment.

Join the Academy →

Performance model

RankAudit · Configuration, Scale and Reporting

Stage 1

Indexability checks

Stage 2

Crawler behaviour settings

Stage 3

Sitemap versus crawl comparison

Indexation & Crawler Settings: Control What Google Sees: Indexability checks to Crawler behaviour settings to Sitemap versus crawl comparison.

Indexability checks

The Indexation report cross-references robots.txt rules, meta robots tags, and X-Robots-Tag headers to classify every page as indexable, blocked, or noindexed, then flags cases where these signals contradict each other, such as a page in the sitemap that is also marked noindex.

  • Cross-checks robots.txt, meta robots, X-Robots-Tag
  • Flags contradictory signals

Crawler behaviour settings

In project settings under Crawler, you can set the user agent RankAudit identifies as, adjust JavaScript rendering on or off for sites relying on client-side rendering, and set a custom crawl-delay to match your server's tolerance.

  • Choose crawler user agent
  • Toggle JavaScript rendering
  • Set custom crawl-delay

Sitemap versus crawl comparison

A comparison view lists pages present in the sitemap but not found during crawling, and pages found during crawling but absent from the sitemap, which often reveals sitemap generation bugs or forgotten content sections.

Crawl budget insights

For larger sites, the crawl budget panel estimates how much of the site search engines are likely to crawl regularly based on internal link depth and update frequency, helping prioritise which sections most need better internal linking.

Why this matters

Failing to accurately configure RankAudit's crawler behaviour can lead to misleading audit reports and flawed strategic decisions. For instance, if your site uses JavaScript rendering for critical content, but you've disabled JavaScript execution in RankAudit, your audit will flag 'empty' pages or missing content that Google, with its modern rendering capabilities, is perfectly capable of seeing and indexing. This generates false positives, wasting time investigating non-existent issues, and potentially overlooking genuine problems masked by the noise.

Conversely, if you permit the crawler to follow `noindex` directives or `robots.txt` exclusions when you explicitly want to audit pages Google is actively ignoring, you will miss crucial diagnostic data. Imagine an internal staging environment accidentally left with `noindex` in production. If RankAudit respects this, it won't report these pages, leaving a critical indexation vulnerability undiscovered until Google drops them from search results. Accurate configuration ensures your audit reflects the reality of Google's current interaction with your site, not a simplified or misaligned interpretation.

Prioritising Crawl Paths: Balancing Depth and Relevance

When auditing large sites or those with complex architectures, strategically prioritising crawl paths is essential to optimise audit efficiency and relevance. Default 'discover all' crawls can consume excessive resources and time, diluting focus from critical sections. Consider the primary conversion funnels or revenue-generating sections of your site first. If a specific product category or blog section is underperforming, configuring RankAudit to deep-crawl these areas, whilst potentially performing a shallower crawl on less critical archives, will yield more actionable insights faster.

This prioritisation also helps manage the 'crawl budget' RankAudit expends on your site, mirroring Google's own behaviour. Identify high-value templates, content types, or URL patterns that are central to your SEO strategy. Ensure these are either explicitly included in your crawl scope or given higher priority in your depth settings. Conversely, set limits on crawling highly paginated archives or deeply nested utility pages unless a specific issue mandates their full exploration. This ensures audit findings are concentrated where they have the greatest impact.

Key considerations for prioritisation:

• Identify high-converting URL segments.

• Set maximum crawl depth for different subdomains/paths.

• Exclude known low-value or purely administrative sections.

• Utilise 'seed URLs' to initiate crawls from specific entry points.

Do it now

Immediately open your active project in RankAudit and review the 'Crawler Configuration' settings under the 'Project Settings' menu. Focus specifically on the 'Robots.txt Directives' and 'JavaScript Rendering' options. Ensure these align with how Google currently interacts with your site, particularly concerning critical JavaScript-rendered content and intentional `noindex` or `nofollow` directives you wish to override for diagnostic purposes. This ensures your next audit run provides a truly representative view.

• Navigate to Project Settings > Crawler Configuration.

• Verify 'Respect robots.txt' status for your audit objective.

• Check 'Enable JavaScript Rendering' for dynamic sites.

• Confirm 'Follow noindex' settings for diagnostic override.

Key takeaways

  • Resolve contradictory indexability signals as a priority
  • Enable JavaScript rendering for client-side rendered sites to get accurate results
  • Compare sitemap and crawl output to catch sitemap errors
  • Use crawl budget insights to focus internal linking effort

Do it now

Tune crawler behaviour, manage many sites, and package findings for clients or leadership. Crawl a site and triage the issue list.