A STUDY OF RANKING ALGORITHM USED BY VARIOUS SEARCH ENGINE

Size: px

Start display at page:

Download "A STUDY OF RANKING ALGORITHM USED BY VARIOUS SEARCH ENGINE"

Patrick Gaines
5 years ago
Views:

A STUDY OF RANKING ALGORITHM USED BY VARIOUS SEARCH ENGINE Bohar Singh 1, Gursewak Singh 2 1, 2 Computer Science and Application, Govt College Sri Muktsar sahib Abstract The World Wide Web is a

1 A STUDY OF RANKING ALGORITHM USED BY VARIOUS SEARCH ENGINE Bohar Singh 1, Gursewak Singh 2 1, 2 Computer Science and Application, Govt College Sri Muktsar sahib Abstract The World Wide Web is a popular and interactive medium to disseminate information today. It is a system of interlinked hypertext documents accessed via the Internet. We use Search Engines to search for information across the Internet. It is very difficult for a user to find the high quality information. When we search any information on the web, the number of URL s has been opened. User wants to show the relevant on the top of the list. So that Page Ranking algorithm is needed which provide the higher ranking to the important pages. In this paper, we discuss the PageRank algorithm used by Google search engine to provide the higher ranking to important pages and then studied Trust rank algorithm used by Yahoo search engine and HITS algorithm used by Twitter and Ask search engine. Keywords - PageRank, Hits, Search Engine, Trust Rank, Backlinks, Damping. I. INTRODUCTION With the increasing number of Web pages and users on the Web, the number of queries submitted to the search engines are also increasing rapidly. Therefore, the search engines needs to be more efficient in its process. The search engines become very successful and popular if they use efficient Ranking mechanism. Google search engine is very successful because of its PageRank algorithm[1]. Page ranking algorithms are used by the search engines to present the search results by considering the relevance, importance and content score and web mining techniques to order them according to the user interest. Some ranking algorithms depend only on the link structure of the documents i.e. their popularity scores, whereas others look for the actual content in the documents, while some use a combination of both i.e. they use content of the document as well as the link structure to assign a rank value for a given document. If the search results are not displayed according to the user interest then the search engine will lose its popularity. So the ranking algorithms become very important. Page Rank is a probability distribution used to represent the likelihood that a person randomly clicking on links will arrive at any particular page. The main disadvantage is that it favors older pages, because a new page, even a very good one, will not have many links unless it is part of an existing site. Trust Rank is a major factor that now replaces PageRank as the flagship of parameter groups in the Google algorithm. It is of key importance for calculating ranking positions and the crawling frequency of web sites. There are various Ranking algorithm used by different search engines like PageRank algorithm used by Google search engine, Trust Rank algorithm (basically combination of PageRank and Trust Rank) used by Yahoo search engine and HITS algorithm used by Twitter to suggest user accounts to follow and also by ASK search engine.[2] II. PAGERANK ALGORITHM Sergey Brin and Larry Page developed PageRank algorithm during their Ph. D. at Stanford University based on the citation analysis. PageRank algorithm is used by the famous search engine, Google. PageRank is a numeric value that represents how important a page is on the web. Google figures that when one page links to another page, it is effectively casting a vote for the other page. The more votes that are cast for a page, the more important the page must be. Also, the importance of DOI: /IJRTER P5OAE 201

2 the page that is casting the vote determines how important the vote itself is. Therefore, PageRank provides a more advanced way to compute the importance or relevance of a Web page than simply counting the number of pages that are linking to it (called as backlinks ). If a backlink comes from an important page, then that backlink is given a higher weighting than those backlinks comes from non-important pages[1]. Figure 1: Linking of web pages 2.1 Actual Algorithm Let us suppose page A has pages P1...Pn which point to it. The damping factor is denoted by d which can be taken between 1 and 0. C(A) is used to denote the number of outgoing links of page A. The PageRank of a page A PR(A) is given as follows: PR(A) = (1-d) + d (PR(P1)/C(P1) PR(Pn)/C(Pn)) It is noted that the Page Ranks form a probability distribution over web pages, so that sum of Page Ranks of all web pages will be one. Page Rank or PR(A) can be calculated using a simple iterative algorithm, and corresponds to the principal eigenvector of the normalized link matrix of the web [3].Different terms used in Page Rank are : PR(Pn) - PR(P1) represent Page Rank for the first page in the web all the way up to PR(Pn) for the last page. C(Pn) - The count, or number, of outgoing links for page 1 is represented by C(P1), for page n is represented by C(Pn) and so on for all pages. PR(Pn)/C(Pn) this term represents the share of the vote page A, if our page (page A) has a backlink from page n d - All these fractions of votes are added together but, to stop the other pages having too much influence, this total vote is damped down by multiplying it by the factor d. It is generally assumed that the damping factor will be set around How the PageRank calculated? Calculation of PageRank is some bit tricky. The PageRank of one page depends on the PageRank of all other pages pointing to it. We could not calculate PR of all those pages until the pages pointing to them have their PR calculated and further so on. But from survey of literature paper it is derived that PR(A) can be calculated using a simple iterative algorithm, and corresponds to the principal Eigen vector of the normalized link matrix of the web. This means that we can calculate a page rank of pages without knowing the final value of the PR of the other pages pointing to them. That seems to be unusual but, basically, each time we do iteration we re finding a closer estimate of the final All Rights Reserved 202

3 So we calculate each value and repeat this process number of times until the numbers stop changing much. Consider the following example of two pages, each pointing to the other. Figure 2: Linking of two web pages There are two page, Page A and Page B, both have one outgoing link i.e. C (A) = 1 and C(B) = 1 Consider the following Guess to understand how PageRank PR is calculated: a) Guess 1 In first case we don t know what their PR should be to start with, so let s take a guess at 1.0 and do some calculations: d= 0.85 PR(A) = (1 d) + d(pr(b)/1) PR(B) = (1 d) + d(pr(a)/1) i.e. PR(A) = * 1 = 1 PR(B) = * 1 = 1 Here the numbers are not changing. Let s take another guess. b) Guess 2 In second case let s take the guess at 0 and do same calculation: PR(A) = * 0 = 0.15 PR(B) = * 0.15 = And do iteration PR(A) = * = PR(B) = * = And again: PR(A) = * = PR(B) = * = and so on. The numbers just keep going up. c) Guess 3 In third case let s start the guess at 20 each and do iteration for final value: PR (A) = 20 PR (B) = 20 For First calculation PR(A) = * 20 = PR(B) = * = And All Rights Reserved 203

4 PR(A) = * = PR(B) = * = Calculated values are going down and It will get to 1.0 and stop. So it is clear that you have to start with your guess, and do iteration until the average PageRank for all pages will be The Random Surfer Model In their publications, Lawrence Page and Sergey Brin give a very simple intuitive justification for the PageRank algorithm. They consider PageRank as a model of user behavior, where a surfer clicks on links at random with no regard towards content. The random surfer visits a web page with a certain probability which derives from the page's PageRank. The probability that the random surfer clicks on one link is solely given by the number of links on that page. This is why one page's PageRank is not completely passed on to a page it links to, but is divided by the number of links on the page. So, the probability for the random surfer reaching one page is the sum of probabilities for the random surfer following links to this page. This probability is reduced by the damping factor d. The justification within the Random Surfer Model, therefore, is that the surfer does not click on an infinite number of links, but gets bored sometimes and jumps to another page at random[3]. 2.4 The damping factor d The probability for the random surfer not stopping to click on links is given by the damping factor d, which depends on probability therefore, is set between 0 and 1. The higher d is, the more likely will the random surfer keep clicking links. Since the surfer jumps to another page at random after he stopped clicking links, the probability therefore is implemented as a constant (1-d) into the algorithm. Regardless of inbound links, the probability for the random surfer jumping to a page is always (1-d), so a page has always a minimum PageRank [1]. 2.5 A Different Notation of the PageRank Algorithm The modified version of PageRank algorithm is: PR(A) = (1-d) / N + d (PR(P1)/C(P1) +...+PR(Pn)/C(Pn)) Where N is the total number of all pages on the web. Regarding the Random Surfer Model, the second version's PageRank of a page is the actual probability for a surfer reaching that page after clicking on many links. The Page Ranks then form a probability distribution over web pages, so the sum of all pages' Page Ranks will be one[5]. 2.6 PageRank indicator on google toolbar PageRank is also displayed on the toolbar of your browser if you ve installed the Google toolbar on your browser ( But the Toolbar PageRank only goes from 0 10 and seems to be something like a logarithmic scale: Toolbar PageRank (log base 10) Real PageRank , ,000-10, , ,000 4 and so All Rights Reserved 204

5 Figure 6: PageRank indicator on google toolbar The Google Toolbar's PageRank feature displays a visited page's PageRank as a whole number between 0 and 10. The most popular websites have a PageRank of 10. The least have a PageRank of 0. Google has not disclosed the precise method for determining a Toolbar PageRank value III. TRUST RANK ALGORITHM Search engine optimization for Trust Rank is same as for PageRank algorithm. Additionally, one just has to ensure that pages are not considered as spam. Trust Rank is a link analysis technique described in paper by Stanford University and Yahoo! researchers for semi-automatically separating useful WebPages from spam. Yahoo uses both the terms PageRank and TrustRank for this purpose. According to Yahoo, PageRank is a family of well-known algorithms for assigning numerical weights to hyperlinked documents (or web pages or web sites) indexed by a search engine. PageRank uses link information to assign global importance scores to documents on the web. The PageRank of a document is a measure of the link-based popularity of a document on the Web. Many Web spam pages are created only with the intention of misleading search engines. These pages, mainly created for commercial reasons, use various techniques to achieve higher-thandeserved rankings on the search engines' result pages. While human experts can easily identify spam, it is too expensive to manually evaluate a large number of pages [1]. According to Yahoo, A spam farm, is an artificially created set of pages that point to a spam target page to boost its significance. Trust-ranking ( TrustRank ) is a form of PageRank with a special teleportation (i.e., jumps) to a subset of high-quality pages. Using the predefined techniques, a search engine can automatically find bad pages (web spam pages) and more specifically, find those web spam pages created to boost their significance through the creation of artificial spam farms (collections of referencing pages). In specific embodiments, a PageRank process with uniform teleportation and a trust-ranking process are carried out and their results are compared as part of a test of the spam-ness of a page or a collection of pages. The main premise of the Trust Rank algorithm is that good pages will usually link to other good pages, unless they have been deceived. And Bad pages can definitely link to good pages as an attempt to look good. Therefore, the basic assumption is that the Trust Rank will mostly transfer to good pages as good and bad pages vote for them. The second premise is that pages that contain many outbound links pay less attention to the sites they link to. Therefore, similarly to the PageRank, a page's vote splits between all of its outbound links according to the new algorithm. The third assumption is that the farther you are from the initial safe sites set, you are more likely to encounter pages that are less trustworthy. Therefore, similarly to the PageRank algorithm, the new algorithm also has a damping element that weakens every vote as it get farther from the initial safe All Rights Reserved 205

IV. HITS ALGORITHM The HITS algorithm stands for Hypertext Induced Topic Selection and is used for rating and ranking websites based on the link information when identifying topic areas.

6 IV. HITS ALGORITHM The HITS algorithm stands for Hypertext Induced Topic Selection and is used for rating and ranking websites based on the link information when identifying topic areas. HITS algorithm used by Twitter to suggest user accounts to follow and also by ASK search engine[11]. Kleinberg's hypertext-induced topic selection (HITS) algorithm is a very popular and effective algorithm to rank documents based on the link information among a set of documents. The algorithm presumes that a good hub is a document that points to many others, and a good authority is a document that many documents point to. Hubs and authorities exhibit a mutually reinforcing relationship i.e. a better hub points to many good authorities, and a better authority is pointed to by many good hubs. To run the algorithm, we need to collect a base set, including a root set and its neighborhood, the in-and out-links of a document in the root set. The HITS algorithm treats web page as a directed graph G (V, E), where V is a set of Vertices representing pages and E is a set of edges that correspond to links. HITS calculate hub and authority scores per query for the focused sub graph of the web. A good authority must be pointed to by several good hubs while a good hub must point to several goods authorities [11]. Hubs Authorities Figure 7: Hub and Authority pages User queries are generally divided into two types. The specific query where the user requires exact matches and narrow information, secondly broad-topic query for user who look for narrow answers and information relation to the broad topic. HITS concentrates on the latter type and aims to find the most authoritative and informative pages for the topic of the query. HITS algorithm, can be stated as follows: Using existing system, get the root set for the given query. Add all the pages linking to and linked from pages in the root set, giving an extended root set or base set. Run iterative eigenvector based computation over a matrix derived from the adjacency Report the top establishment and hubs. The first step in the HITS algorithm shows that the root set for a given query is taken from a search engine. The second step basically expands the root set by one link neighborhood to form the base set. The hub and authority value of page can be calculated in the following way: v = At.u u = A.v Where A be the adjacency matrix of the graph, At is the transpose of matrix A and the authority weight vector denoted by by v and the hub weight vector denoted by All Rights Reserved 206

7 HITS algorithm is in the same spirit as PageRank. They both make use of the link structure of the Web graph in order to decide the relevance of the pages. The difference is that unlike the PageRank algorithm, HITS only operates on a small sub graph from the web graph. This sub graph is query dependent; whenever we search with a different query phrase, the seed changes as well. HITS ranks the seed nodes according to their authority and hub weights. The highest ranking pages are displayed to the user by the query engine. V. CONCLUSION On the basis of analysis of different ranking algorithm we conclude that all different link analysis algorithms that employ different models to calculate web page rank. In this paper it is discussed what the PageRank algorithm is, how it is important for ranking the web pages, what are the various aspects in calculating PageRank and the description of actual working of PageRank algorithm. In this paper two more ranking algorithm Trust Rank and HITS algorithm used by Yahoo and Ask search engine respectively are discussed. So Page Ranking is very important for now days for searching any efficient and correct information from internet. REFERENCES I. S. S. Mridula Batra, comparative study of page rank Algorithm with different ranking algorithms Adopted by search engine for website ranking, vol. Vol 4 (1), pp II. D. S. R. P. A.M. Sote, Application of Page Ranking Algorithm in Web Mining, IOSR Journal of Computer Science, pp , III. PageRank Explained, [Online]. Available: IV. PageRank Algorithm - The Mathematics of Google Search, [Online]. Available: V. Ricardo Baeza-Yates and Emilio Davis, Web page ranking using link attributes, In proceedings of the 13th international World Wide Web conference on Alternate track papers & posters, PP , VI. Aallan borodin, Link Analysis Ranking: Algorithms, Theory, and Experiments, University of Toronto VII. L. Page, S. Brin, R. Motwani, and T. Winograd, The PageRank Citation Ranking: Bringing Order to the Web, Technical Report, Stanford Digital Libraries SIDL-WP , VIII. Neelam Duhan,A.K.Sharma and Komal Kumar Bhatia, Page Ranking Algorithms : A Survey, In proceedings of the IEEE International Advanced Computing Conference (IACC),2009 IX. Dilip Kumar Sharma, A Comparative Analysis of Web Page Ranking Algorithms, (IJCSE) International Journal on Computer Science and Engineering Vol. 02, No. 08, 2010, X. Sung Jin Kim and Sang Ho Lee, An Improved Computation of the PageRank Algorithm, In proceedings of the European Conference on Information Retrieval (ECIR), XI. L. Li, Y. Shang, and W. Zhang, Improvement of HITS-based algorithms on web documents, in Proceedings of the Eleventh International Conference on the World Wide Web, May All Rights Reserved 207

Web Structure Mining using Link Analysis Algorithms

Web Structure Mining using Link Analysis Algorithms Ronak Jain Aditya Chavan Sindhu Nair Assistant Professor Abstract- The World Wide Web is a huge repository of data which includes audio, text and video.