Getting a Crawl Running · Lesson 1 of 14 · 9 min
Crawl Setup: Launch Your First Site Audit
Create your first RankAudit project, point it at a domain or sitemap, choose crawl depth and speed, and launch the initial crawl that populates every other report in the tool.
Video walkthrough coming soon
This lesson is fully written below. The recorded walkthrough for “Crawl Setup: Launch Your First Site Audit” will appear here once it is published.
Performance model
RankAudit course Getting a Crawl Running
Stage 1
Creating a new audit project
Stage 2
Choosing crawl scope and limits
Stage 3
Robots, authentication and staging sites
Reading time · unlocks the next lesson
0:00 / 5:00 · pausedThe timer counts only while this tab is open and resumes where you left off.
Creating a new audit project
From the RankAudit dashboard, click "New Project" and enter the root domain you want to audit. You will be asked to confirm the preferred protocol (http or https) and whether to include or exclude subdomains, which matters for sites that split content across www and regional subdomains.
Give the project a clear name, especially if you manage several sites, since this name appears in every report export and white-label PDF generated later.
- Dashboard → New Project → enter domain
- Choose subdomain inclusion rules
- Name the project for easy identification later
Choosing crawl scope and limits
The setup screen lets you cap the crawl by page count, crawl depth (number of clicks from the homepage), or by uploading a sitemap URL to seed the crawler directly. Smaller sites can usually crawl unrestricted, while larger sites benefit from a depth or page cap to control processing time and API usage.
You can also set crawl speed, which throttles requests per second to avoid overloading a shared host during business hours.
- Set max pages or max depth
- Optionally seed from an XML sitemap
- Adjust request rate to protect the host
Robots, authentication and staging sites
Before launching, review the robots.txt handling toggle: by default RankAudit respects robots.txt, but you can override this for staging environments that block all crawlers. If the site sits behind basic authentication, enter the credentials in the advanced settings panel so the crawler can log in.
- Toggle robots.txt compliance for staging sites
- Add basic auth credentials if required
Launching and monitoring the first crawl
Click "Start Crawl" and you will land on a live progress screen showing pages discovered, pages crawled, and errors encountered in real time. Crawl time depends on site size and speed settings; you can navigate away and RankAudit will notify you when the crawl completes.
Key takeaways
- Set subdomain and protocol rules before the first crawl to avoid duplicate project data
- Cap page count or depth on large sites to keep crawl times manageable
- Add authentication credentials for staging sites before launching
- Use the live progress screen to catch early crawl errors
Practise it now in RankAudit
Crawl a site and triage the issue list
1h 59m of material in this course · 14 lessons
