Getting a Crawl Running
Crawl Setup: Launch Your First Site Audit
Create your first RankAudit project, point it at a domain or sitemap, choose crawl depth and speed, and launch the initial crawl that populates every other report in the tool.
Want to earn a Rankar Academy certificate?
This guide is free to read in full, with nothing to sign up for. Join the Academy to track your learning, complete courses and sit the certification assessment.
Join the Academy →Performance model
RankAudit · Getting a Crawl Running
Stage 1
Creating a new audit project
Stage 2
Choosing crawl scope and limits
Stage 3
Robots, authentication and staging sites
Creating a new audit project
From the RankAudit dashboard, click "New Project" and enter the root domain you want to audit. You will be asked to confirm the preferred protocol (http or https) and whether to include or exclude subdomains, which matters for sites that split content across www and regional subdomains.
Give the project a clear name, especially if you manage several sites, since this name appears in every report export and white-label PDF generated later.
- Dashboard → New Project → enter domain
- Choose subdomain inclusion rules
- Name the project for easy identification later
Choosing crawl scope and limits
The setup screen lets you cap the crawl by page count, crawl depth (number of clicks from the homepage), or by uploading a sitemap URL to seed the crawler directly. Smaller sites can usually crawl unrestricted, while larger sites benefit from a depth or page cap to control processing time and API usage.
You can also set crawl speed, which throttles requests per second to avoid overloading a shared host during business hours.
- Set max pages or max depth
- Optionally seed from an XML sitemap
- Adjust request rate to protect the host
Robots, authentication and staging sites
Before launching, review the robots.txt handling toggle: by default RankAudit respects robots.txt, but you can override this for staging environments that block all crawlers. If the site sits behind basic authentication, enter the credentials in the advanced settings panel so the crawler can log in.
- Toggle robots.txt compliance for staging sites
- Add basic auth credentials if required
Launching and monitoring the first crawl
Click "Start Crawl" and you will land on a live progress screen showing pages discovered, pages crawled, and errors encountered in real time. Crawl time depends on site size and speed settings; you can navigate away and RankAudit will notify you when the crawl completes.
Why this matters
Neglecting the initial crawl setup, particularly the scope and limits, directly impacts the utility of all subsequent RankAudit reports. An improperly configured crawl might miss critical sections of a website, such as a category tree hidden behind faceted navigation or canonicalised variants of pages. This can lead to skewed insights, as the tool won't have a complete picture of the site's architecture or indexability, causing SEOs to act on incomplete data regarding technical issues or content opportunities.
Conversely, a meticulously configured crawl ensures comprehensive data collection, providing a reliable foundation for all analyses. Consider a large e-commerce site: if the crawl depth is too shallow, product pages beyond the second click might be entirely missed. An SEO relying on such a limited crawl could conclude, erroneously, that product schema is universally absent, or that internal linking is sufficient, when in reality, the crawl simply failed to reach the majority of relevant pages to report on these issues effectively.
Prioritising Crawl Source: Domain vs. Sitemap
The choice between crawling primarily from the domain (following internal links) or a sitemap is fundamental and dictates the initial discovery process. Crawling the domain simulates a search engine bot's natural discovery path, prioritising internally linked pages and identifying orphaned content. This method is crucial for auditing internal linking structures, detecting broken links, and understanding the crawl budget implications of your site's architecture, as it precisely follows the links a bot would encounter in a natural traversal.
Conversely, initiating the crawl from an XML sitemap provides an explicit list of pages you intend for search engines to discover. This is highly effective for ensuring all canonical URLs are included in your audit, especially if they are poorly linked internally or are recently added. However, relying solely on a sitemap crawl can mask critical internal linking issues or reveal pages that are discoverable via sitemaps but entirely unlinked within the site's navigation, leading to potential crawlability and indexability problems. A hybrid approach, often involving a domain crawl supplemented by sitemap discovery, usually yields the most comprehensive and actionable data.
When deciding your primary crawl source, consider these points:
Use domain crawl first to assess internal link equity flow and discoverability.
Use sitemap crawl to confirm all intended indexable pages are included.
Combine sources for large or complex sites to ensure full coverage.
Prioritise domain-only if internal linking is a primary audit focus.
Do it now
Access your newly created RankAudit project. Navigate to the 'Crawl Settings' tab and locate the 'Crawl Source' options. Experiment with selecting 'Start from Domain' and then 'Start from Sitemap', observing how the 'Included URLs' preview changes based on your input. While you won't launch a full crawl now, this step familiarises you with where these crucial configurations reside and how they visibly influence the projected crawl scope.
- Open your RankAudit project.
- Click on the 'Crawl Settings' tab.
- Locate the 'Crawl Source' section.
- Toggle between 'Start from Domain' and 'Start from Sitemap' to see the effect.
Key takeaways
- ✓Set subdomain and protocol rules before the first crawl to avoid duplicate project data
- ✓Cap page count or depth on large sites to keep crawl times manageable
- ✓Add authentication credentials for staging sites before launching
- ✓Use the live progress screen to catch early crawl errors
Do it now
Set up your first site audit and learn to read what it produces. Crawl a site and triage the issue list.
