Daily current affairs on DailyCA
DailyCA NotesStudy notes for competitive exams My revisionsRevisions

How Search Engines Work

Basic 4 min read Updated

Saved in this browser only. See all revisions due

Key points

  • Crawling then indexing then ranking is the fixed order
  • A search engine searches its own index, not the live internet
  • Organic results are earned; paid results are marked Ad or Sponsored
  • SEO improves organic rank; SEM is paid advertising
  • PageRank was developed by Larry Page and Sergey Brin
  • robots.txt tells crawlers what not to crawl
  • The deep web is unindexed and mostly lawful; the dark web is a small part of it
  • Shodhganga is maintained by INFLIBNET; NDL was developed by IIT Kharagpur
On this page
  1. What a search engine is
  2. The three stages
  3. The results page
  4. PageRank and robots.txt
  5. Surface, deep and dark web
  6. Academic and official sources

What a search engine is

A search engine is a system that finds web pages matching a query, such as Google, Bing, DuckDuckGo or Yandex. A search engine is a service or website; the browser is the software that opens it. The two are not the same.

The three stages

  1. Crawling: an automated program called a crawler, spider or bot follows links and visits pages. The Google crawler is Googlebot.
  2. Indexing: the collected pages are analysed and stored in a huge index.
  3. Ranking or retrieval: algorithms order the matching results and display them.

Key point

When you search, the engine does not search the whole internet at that moment. It searches its own pre-built index. That is why a page published minutes ago may not appear until it has been crawled and indexed.

The results page

The SERP, Search Engine Results Page, carries two kinds of result:

  • Organic results: natural results ranked by relevance.
  • Paid results: advertisements marked Ad or Sponsored. In official work these must be recognised and set aside.

SEO, Search Engine Optimisation, is the effort to rank higher in organic results. Paying for advertisements is SEM, Search Engine Marketing.

PageRank and robots.txt

PageRank, developed by Larry Page and Sergey Brin at Stanford, judges the importance of a page from the number and quality of links pointing to it. Google was founded in 1998.

robots.txt is a file in the root folder of a website that tells crawlers which parts must not be crawled.

Surface, deep and dark web

LayerMeaningExample
Surface webIndexed by search enginesPublic websites
Deep webNot indexed, mostly lawfulE-mail inbox, bank account, institute intranet, database results
Dark webA small part of the deep web reachable only with special softwareSites reached through Tor

Example

Your own e-mail inbox is part of the deep web, because a search engine cannot index it. The deep web is not the same as the dark web.

Academic and official sources

  • Google Scholar: academic articles and citations.
  • Shodhganga: repository of Indian doctoral theses, maintained by INFLIBNET.
  • e-ShodhSindhu: consortium for electronic resources.
  • NDL, the National Digital Library of India, developed by IIT Kharagpur.
  • SWAYAM: platform for online courses.
  • For rules and orders, go to the source: DoPT, the Department of Expenditure, India Code, the e-Gazette, CPPP and GeM.

Exam tip

A metasearch engine such as Dogpile has no index of its own and only combines results from other engines. That distinction is a frequent question.

Practice questions

Answer all, then check. Explanations appear after checking.

1The automated program that visits web pages and follows links to collect information for a search engine is called a:
2What is the correct order of the three stages of a search engine?
3The part of the web that is not indexed by search engines, such as e-mail inboxes and password-protected databases, is known as the:
4Shodhganga, the repository of Indian doctoral theses, is maintained by:
5When a query is typed, a search engine actually searches:

Finished this topic? Tick it off.

Saved in this browser only. See all revisions due