Skip to main content

How Search Engines Find and Index Small Business Sites

How search works · 9 min read ·

How a small site gets from nothing to a search result: discovery, crawling, indexing and serving, and what you can influence at each step.

Illustration: A minimal diagram of four boxes in a row: found, crawled, indexed, shown, with thin arrows and a single search bar above

You publish a new page about your business. You search for its name the next morning and it is not there. A week later, still nothing. It is tempting to assume that something is broken, or that search engines are ignoring you. Usually, neither is true. The page is somewhere in a pipeline that takes time, with steps that can stall for ordinary reasons.

Understanding the pipeline turns a mystery into a checklist. This article walks through the stages a page passes through on its way to a search result, in plain language, and points out what a small business can influence at each one.

The short version

Google describes search in three stages: crawling, indexing and serving. Other search engines work in similar ways.

  1. Crawling. Automated programs, often called crawlers, find pages and download their text, images and other content.
  2. Indexing. The content is analysed and stored in a large database called the index.
  3. Serving. When someone searches, the engine looks in the index for relevant pages and shows results.

A page can only appear in results if it passes all three. Each can fail separately.

Stage one: discovery and crawling

Before a page can be crawled, a search engine has to know it exists.

How pages are discovered

  • Links from known pages. Crawlers follow links. If a page they already know links to yours, they find it. This is the most common route, and the reason that links from other sites matter.
  • Sitemaps. A sitemap is a file that lists the pages you want known. Google explains that it helps crawlers find content more efficiently, and that it is especially useful for sites that are new with few external links, large or rich in media.
  • Direct submission through the tools search engines provide to site owners.
  • Links within your own site. If your home page links to your services page, and your services page links to a sub-page, crawlers can follow the chain.

A page that nothing links to, and that is not in a sitemap, may never be found. These are sometimes called orphan pages.

How crawlers decide what to fetch

Crawlers cannot visit everything all the time. They decide which sites to visit, how often and how many pages to fetch, based on signals such as how important the site seems, how often it changes and how quickly it responds. A small, new site gets modest attention. That is normal.

What can stop crawling

  • A robots file that blocks access. The robots file is mainly meant to manage crawler traffic. It is not a way to hide pages from search, and a mistake in it can accidentally block your whole site.
  • Server errors or slow responses. If your site is down or very slow when the crawler visits, it may try later, or less often.
  • Links that crawlers cannot follow. Google advises using proper HTML link elements with real addresses. Links created only by scripts, or that use non-standard forms, may not be followed.
  • Login walls. Crawlers cannot sign in.
  • Content hidden behind interactions. Text that appears only after a click may not be seen.

Stage two: indexing

After crawling, the engine tries to understand the page. Wikipedia describes indexing as the collecting, parsing and storing of data to enable fast and accurate retrieval. In practice, the engine breaks the text into words, notes what the page seems to be about, looks at headings, links and images, and stores a representation in a vast database. A common structure for this is an inverted index, which records, for each word, which documents contain it, so that a query can find matching documents in milliseconds instead of scanning everything.

Duplicates and canonical versions

Google explains that during indexing it identifies duplicate content and chooses the most representative version to show. If you have the same text on several addresses, such as with and without "www" or with tracking parameters, the engine picks one and may ignore the rest. You can help by making sure each page has one main address.

Quality and usefulness

Not every crawled page is indexed. Google notes that indexing is not guaranteed for every page it processes. Pages may be left out because they are very thin, near-duplicates of others, blocked by an instruction, of low value or technically broken.

What can stop indexing

  • A noindex instruction. A tag or header that tells search engines not to index the page. This is useful when you want it, and disastrous when it is left on by accident, for example after a site is built in a staging environment.
  • A robots block combined with noindex. Google explains that for a noindex rule to work, the page must not be blocked from crawling, because the crawler has to see the rule.
  • Redirect loops and errors.
  • Thin or duplicate content.
  • Poor mobile versions. Google indexes the mobile version of a page, so content that is missing on mobile may not be indexed.

Stage three: serving

When someone types a query, the engine searches its index for pages that match, and ranks them. Google says that results depend on factors such as the user's location, language and device, and that not every indexed page appears for every search. Ranking uses a large number of signals that nobody outside the company knows in full.

What you can say with confidence is that the engine tries to show pages that are relevant to the query and helpful to the person asking. Google's guidance on helpful content says its systems are designed to reward content created primarily for people, with clear information about who is behind it, first-hand experience where relevant and accurate sourcing.

Where a small business can help

You cannot control the engine, but you can remove obstacles and make your pages easier to find and understand.

  1. Link your pages together with ordinary, descriptive links. Make sure every page you want found can be reached from the home page in a few clicks.
  2. Publish a sitemap and tell the search engines about it through their webmaster tools.
  3. Check that nothing blocks you. Look at your robots file and for any noindex instructions on important pages.
  4. Make the site work on a phone, since the mobile version is what is indexed.
  5. Use clear page titles and headings that say what each page is about.
  6. Write useful, original content, with your own words, details and experience.
  7. Make the site fast and reliable.
  8. Get legitimate links. Listings in reputable directories, mentions from local organisations and links from partners help crawlers find you and indicate that you exist in the real world.
  9. Use the free tools. Google describes Search Console as a service that helps you check that Google can find and crawl your site, fix indexing problems, see search performance and receive alerts about issues.

What to expect and when

A brand-new site with few links may take days or weeks to appear for its own name, and longer for more general searches. Pages on established sites may appear within hours or days. No one can promise a date. If a page is still absent after a few weeks, use Search Console to inspect it and see whether it has been crawled and indexed, and what the engine reports about it.

A worked example

A small firm of accountants launches a new website in March. By April, searching for the firm's name shows nothing. The owner checks Search Console, which says that pages are discovered but not indexed. She looks at the pages and finds that the developer left a noindex tag on every page from the testing phase. She removes it, submits the sitemap and requests inspection of the home page.

Within a week, the home page appears for the firm's name. Other pages follow over the next few weeks. She adds descriptive links from the home page to each service page, and registers the firm in two reputable directories, which link to the site. Over the following months, she notices the site appearing for searches that combine the town and the services, and she keeps an eye on Search Console for errors.

Common mistakes

  1. Leaving noindex on after launch.
  2. Blocking crawlers in the robots file by accident.
  3. Pages that nothing links to.
  4. Menus that only work with scripts.
  5. Different content on mobile and desktop.
  6. Expecting instant results.
  7. Duplicate versions of the same page.

A checklist

  • Home page links to every important page
  • Sitemap published and submitted
  • Robots file and noindex checked
  • Mobile version complete
  • Clear titles, headings and descriptive links
  • Original, helpful content
  • A few legitimate links from real sources
  • Search Console set up and monitored

The search page here shows a simple search in action, the categories page shows how listings are grouped and the submit page lets a business add a listing that links to its site.

Frequently asked questions

Do I have to submit my site to search engines? Not always, but a sitemap and webmaster tools speed up discovery and show problems.

Why is my page crawled but not indexed? It may be a duplicate, thin, blocked or considered low value. Inspect it for details.

Will a directory listing help? A link from a reputable page can help crawlers find you, and a listing helps people find you directly.

Questions and answers

How does a search engine discover a new site?
Mainly by following links from pages it already knows, and from sitemaps and submissions made through webmaster tools.
Does being crawled mean being indexed?
No. A page can be crawled and still not be stored in the index, for instance if it is a duplicate, thin or blocked.
How long does it take?
It varies from days to weeks or more. New sites with few links and little history tend to be slower.
What can a small business do to help?
Make pages easy to reach through ordinary links, publish a sitemap, avoid blocking crawlers by accident and write pages that are clearly useful.

Sources

Get help