9 Reasons Why Your Pages Aren’t Indexed - ZipTie.ai - AI Search Intelligence

9 Reasons Why Your Pages Aren’t Indexed

ZipTie Team

15 min read

Published: September 20, 2022

Indexing has been a popular topic in the SEO industry for a while. Seeing so many posts, articles, and forums discussions about indexing had me realize that many aspects of indexing SEO are still confusing for webmasters. This made me think about ways to help them understand their indexing issues and look for solutions. This

Indexing has been a popular topic in the SEO industry for a while. Seeing so many posts, articles, and forums discussions about indexing had me realize that many aspects of indexing SEO are still confusing for webmasters. This made me think about ways to help them understand their indexing issues and look for solutions.

This article is a list of the most common reasons your pages aren’t indexed.

If you struggle to get your valuable pages indexed and don’t know which site optimization aspects to focus on or where to start, this article is for you.

You will learn how to identify the issues causing your pages to not be indexed, why they occur, and what recommendations I have for fixing them.

Let’s start with the basics.

1. Your pages are non-indexable.

Google won’t index a given page if you clearly instruct it that a page shouldn’t be indexed. There are many ways to do that, some providing stronger signals to Google than others.

One way to make a page non-indexable is by adding a “noindex” meta tag to it – it would look like this:

Google won’t index it. Period.

Unfortunately, it’s common for webmasters to add “noindex” tags by mistake.

To make sure this isn’t the case for you, check the list of all pages with a “noindex” tag to ensure the tags are only placed on pages that really shouldn’t be indexed.

Use a crawler like OnCrawl or Screaming Frog. After crawling the site, you will be able to see any “noindex” directives added to your URLs. You can export the crawl data and go through the URLs with “noindex” to see if they were mistakenly added to any valuable pages.

But there are other signals that tell Google your pages shouldn’t be indexed. However, these signals aren’t definite and, in some situations, Google could still index such pages.

Your pages may not be indexed if they are:

Note that it may have been your intention to make these pages non-indexable – but if some of your pages aren’t indexed, and it feels like a mistake, ensure it’s not because of the issues mentioned above.

Look at the Excluded section of the Index Coverage report. Pay attention to URLs appearing with the following statuses indicating that the specified pages can’t be indexed:

2. There is a JavaScript SEO issue on your page.

As Bartosz Góralewicz showed, Google had tremendous issues with rendering JavaScript in the past.

The process of downloading, parsing, and executing JavaScript is time-consuming and resource-heavy for Google.

Over the years, Google did an excellent job improving its rendering, but there is still a risk that Google won’t index your JavaScript content.

Here is when Google might not index your JavaScript-based content:

What exactly could happen if JavaScript is not rendered and crucial content on a page relies on it?

Here is an example of Angular.io: if JavaScript isn’t rendered, the only content Google will see is: “This website requires JavaScript.”

Another example of a site that is hurt by its implementation of JavaScript SEO is disqus.com. Disqus uses dynamic rendering in the form of prerendering, which would present Google with a static page version.

This solution is generally recommended by Google but, in this case, the page doesn’t get rendered correctly, likely due to faulty implementation.

The result? Googlebot is getting an empty page:

To mitigate JavaScript SEO issues, ensure your essential content can be accessed by Googlebot with JavaScript enabled and disabled. If Google has issues with JavaScript on your site and JavaScript is used to generate your key content, your JavaScript-heavy pages may not be indexed.

Usually, URLs with a JavaScript-related problem will be classified by Google’s Index Coverage as:

Be sure to familiarize yourself with Google’s JavaScript SEO best practices.

3. Page is classified by Google as soft 404.

Google uses many tools to ensure the web pages it shows in search results are of the highest quality and provide a positive user experience.

One of the tools that Google utilizes is a soft 404 detector. If a page is detected as soft 404, it won’t get into Google’s index.

Soft 404s are not official response codes on websites. A 404 page returns a correct 200 status code, but its contents make it look like an error page, e.g. because it’s empty or contains thin content – or so Google thinks.

As you can see, a soft 404 detector, like every mechanism, is prone to false positives. It means that your pages may be wrongly classified.

Google could wrongly classify your pages in a few cases:

  1. Google can’t properly render your JavaScript content. Ensure you’re not blocking JavaScript in robots.txt and that Googlebot can render your crucial resources.
  2. Google found some words it typically associates with soft 404 pages, such as: “page not found” or “product unavailable.” In this case, adjust your copy. Depending on the situation, you may want to redirect such pages or make them 404s.
  3. The page should be a 404 page that mistakenly responds with a 200 status code. This could be the case if you decide to create a custom 404 page. You need to configure your server to respond with a 404 status code.
  4. A redirect has been implemented, but the target page isn’t thematically connected to the origin page. Redirect it to the closest matching alternative.

Google’s selection of pages as soft 404s was impacted by its Caffeine update – here is how Gary Illyes explained it:

“Basically, we have very large corpora of error pages, and then we try to match text against those. This can also lead to very funny bugs, I would say, where, for example, you are writing an article about error pages in general, and you can’t, for your life, get it indexed. And that’s sometimes because our error page handling systems misdetect your article, based on the keywords that you use, as a soft error page. And, basically, it prompts Caffeine to stop processing those pages.”

4. Your page is of low quality.

One of the most important ranking signals for Google is content quality.

Over the years, Google introduced many algorithm changes to highlight how crucial it is for pages to create content that is:

That’s why we shouldn’t expect Google to index content that doesn’t follow these guidelines.

Moreover, if Google sees some of your low-quality content, it may view the whole website as low-quality and, subsequently, limit its crawling and indexing.

Usually, a page with low-quality content will be classified as:

There are a few ways to tackle low-quality content issues on your site – consider:

5. The page has duplicate content.

This is related to the previous point about low-quality content, but this issue refers to multiple pages containing the same or very similar content.

A page with duplicate content likely won’t be indexed in Google.

The main dangers of having a lot of duplicate content on your site include:

Some examples of duplicate content include:

Duplicate content is a common indexing issue on eCommerce or other large websites, and it’s particularly severe for them.

Usually, duplicate content will be classified by Google as:

There are two most common solutions to tackle duplicate content issues:

6. Your pages are slow.

Having a slow website can negatively impact user experience, but it could also lead to indexing issues.

Let me elaborate:

The critical aspect of improving your site’s performance with Google’s crawling and indexing processes in mind is optimizing your server.

If your website is visibly slow for users who interact with it – for example, it fails the Core Web Vitals assessment – it’s still a problem that requires your attention.

But what you should focus on is whether your server can handle Google’s crawl requests. For example, when you add new content and Google’s crawling increases, you may find that this content isn’t indexed because the server slowed down.

You need to make sure your website can handle traffic spikes from Google to be crawled and indexed at high rates.

7. There is an indexing bug on Google’s side.

Google is probably one of the most advanced systems in the world, and it has been actively (and successfully) maintained for over 20 years now.

However, every software has bugs. And some bugs on Google’s end can cause your pages to not be indexed or to be reported as such.

An example of a widely noticed Google indexing bug happened in October 2020:

We are currently working to resolve two separate indexing issues that have impacted some URLs. One is with mobile-indexing. The other is with canonicalization, how we detect and handle duplicate content. In either case, pages might not be indexed....

It took Google 2 weeks to fix the bug and similar bugs happen from time to time.

8. Your page or website is too new.

No content is indexed immediately. In many situations, your pages will end up being indexed, but it will take some time.

As John Mueller stated:

“When a new page is published on a website, it can take anywhere from several hours to several weeks for it to be indexed. In practice, I suspect most good content is picked up and indexed within about a week.”

Two factors cause such indexing delays:

  1. It takes time for Google to discover a new page.
  2. It takes time for a page to get to the top of Google’s crawling queue.

Usually, a URL in Google’s crawling queue will be classified as Discovered, currently not indexed.

You may also experience delays in crawling and indexing content if you publish it on a new website.

9. Google refused to visit the page.

Google sometimes refuses to visit a page because it thinks it’s not worth crawling and indexing it.

This may be the result of two things:

  1. Google is not convinced to visit a specific page because the page lacks relevant signals. For instance, if no links point to a given page, Google likely won’t visit and index your page. Another signal would be if Google couldn’t find the page in your sitemap.
  2. Google is not convinced to visit those URLs because they fall into a specific URL Pattern. For example, Google recognizes a given page pattern as related to some previously visited pages. It could be pages with duplicate content or, for example, author or user profiles. If Google sees other pages that appear to follow this pattern, it doesn’t need to waste time and resources crawling them.

Your next steps here revolve around solutions that I mentioned in other chapters:

Wrapping up

You can now see that some indexing issues may have little to do with your website and more with Google’s limited resources and bugs or errors.

However, in most cases, your pages may be lacking quality or sufficient signals to get indexed. It’s also possible that you are preventing Googlebot from accessing some pages that should be indexed.

Always remember to: