Crawlability is the ability of Google's robots to access, navigate and read your site's pages — an essential prerequisite for indexing and ranking.
Crawlability describes how easily Googlebot — Google's indexing robot — can access and navigate your site. Factors that affect it: a restrictive robots.txt, missing or broken internal links, redirect chains, server errors (5xx), complex pagination structures, excessive JavaScript that hides content, and an exhausted crawl budget (on very large sites). Crawlability is the first condition for indexing — what cannot be crawled cannot be indexed, and what is not indexed cannot rank.
Googlebot has a limited crawl budget per site. If you waste crawl time on worthless pages (duplicate filter pages, parameter URLs), Google does not reach the important pages. Large sites can lose tens of thousands of pages from the index due to crawlability issues.
Source: Google Search Central — CrawlingTechnical SEO covers the optimizations of your site's infrastructure — crawlability, indexing, loading speed, HTTPS and data structure — that let Google find and understand your pages.
An XML sitemap is a file that lists all the important URLs of your site, helping Google discover and index them faster and more completely.
Robots.txt is a text file placed at the root of a site that tells crawlers (Google, Bing, GPTBot) which pages they may or may not access and index.
The canonical tag (rel="canonical") tells Google which is the main version of a page when similar or duplicate URLs exist, preventing duplicate-content penalties.
Explore how to apply this concept to your industry and city.