Professional Documents
Culture Documents
Nodes: Webpages
Edges: Hyperlinks TODAY 2.00Pm there will be
A demo on writing a scraper
In F-106 by Mr. Amit
I teach a
class on IR.
CS F469:
Classes are
in the
LTC
CSIS Dept
at BITS
HYD
BITS PILANI
HYD
CAMPUS
2
Web as a directed graph:
Nodes: Webpages
Edges: Hyperlinks
I teach a
class on IR.
CS F469:
Classes are
in the
LTC building
CSIS DEPT
BITS HYD
BITS PILANI
HYD
CAMPUS
3
How to organize the Web?
First try: Human curated
Web directories
Yahoo, DMOZ, LookSmart
Second try: Web Search
Information Retrieval investigates:
Find relevant docs in a small
and trusted set
Newspaper articles, Patents, etc.
But: Web is huge, full of untrusted documents,
random things, web spam, etc.
2 challenges of web search:
(1) Web contains many sources of information
Who to trust?
Trick: Trustworthy pages may point to each other!
(2) What is the best answer to query
newspaper?
No single right answer
Trick: Pages that actually know about newspapers
might all be pointing to many newspapers
All web pages are not equally important
http://universe.bits-
pilani.ac.in/hyderabad/arunamalapati/Profile
vs.
http://www.bits-pilani.ac.in/Hyderabad/index.aspx
D
E F
3.9
8.1 3.9
1.6
1.6 1.6 1.6
1.6
Each links vote is proportional to the
importance of its source page
j rj/3
rj = ri/3+rk/4
rj/3 rj/3
A vote from an important The web in 1839
i j di
ry = ry /2 + ra /2
ra = ry /2 + rm
rm = ra /2
out-degree of node
Flow equations:
3 equations, 3 unknowns, ry = ry /2 + ra /2
no constraints ra = ry /2 + rm
rm = ra /2
No unique solution
All solutions equivalent modulo the scale factor
Additional constraint forces uniqueness:
+ + =
Solution: = , = , =
Gaussian elimination method works for
small examples, but we need a better
method for large web-size graphs
We need a new formulation!
Stochastic adjacency matrix
Let page has out-links
1
If , then = else = 0
is a column stochastic matrix
Columns sum to 1
Rank vector : vector with an entry per page
is the importance score of page
= 1
ri
The flow equations can be written rj
i j di
=
ri
Remember the flow equation: rj
d
Flow equation in the matrix form i j i
=
Suppose page i links to 3 pages, including j
i
j rj
. =
ri
1/3
M . r = r
NOTE: x is an
eigenvector with
The flow equations can be written the corresponding
eigenvalue if:
= =
r = Mr
ry = ry /2 + ra /2 y 0 y
ra = ry /2 + rm a = 0 1 a
rm = ra /2 m 0 0 m
Given a web graph with n nodes, where the
nodes are pages and edges are hyperlinks
Power iteration: a simple iterative scheme
Suppose there are N web pages (t )
ri
( t 1)
Initialize: r(0) = [1/N,.,1/N]T rj
i j di
Iterate: r(t+1) = M r(t) di . out-degree of node i
Stop when |r(t+1) r(t)|1 <
|x|1 = 1iN|xi| is the L1 norm
Can use any other vector norm, e.g., Euclidean
y a m
Power Iteration: y
y 0
Set = 1/N a 0 1
a m m 0 0
1: =
ry = ry /2 + ra /2
2: = ra = ry /2 + rm
Goto 1 rm = ra /2
Example:
ry 1/3 1/3 5/12 9/24 6/15
ra = 1/3 3/6 1/3 11/24 6/15
rm 1/3 1/6 3/12 1/6 3/15
Iteration 0, 1, 2,
y a m
Power Iteration: y
y 0
Set = 1/N a 0 1
a m m 0 0
1: =
ry = ry /2 + ra /2
2: = ra = ry /2 + rm
Goto 1 rm = ra /2
Example:
ry 1/3 1/3 5/12 9/24 6/15
ra = 1/3 3/6 1/3 11/24 6/15
rm 1/3 1/6 3/12 1/6 3/15
Iteration 0, 1, 2,
i1 i2 i3
Imagine a random web surfer:
At any time , surfer is on some page
At time + , the surfer follows an j
ri
out-link from uniformly at random rj
i j d out (i)
Ends up on some page linked from
Process repeats indefinitely
Let:
() vector whose th coordinate is the
prob. that the surfer is at page at time
So, () is a probability distribution over pages
i1 i2 i3
Where is the surfer at time t+1?
Follows a link uniformly at random
j
+ = () p(t 1) M p(t )
Suppose the random walk reaches a steady
state + = () = ()
then () is stationary distribution of a random walk
Our original rank vector satisfies =
So, is a stationary distribution for
the random walk
A central result from the theory of random
walks (a.k.a. Markov processes):
Example:
ra 1 0 1 0
=
rb 0 1 0 1
Iteration 0, 1, 2,
(t )
ri
( t 1)
a b rj
i j di
Example:
ra 1 0 0 0
=
rb 0 1 0 0
Iteration 0, 1, 2,
Dead end
2 problems:
(1) Some pages are
dead ends (have no out-links)
Random walk has nowhere to go to
Such pages cause importance to leak out
a m a m
y a m
Power Iteration: y
y 0
Set = 1 a 0 0
a m m 0 0
=
ry = ry /2 + ra /2
And iterate
ra = ry /2
rm = ra /2
Example:
ry 1/3 2/6 3/12 5/24 0
ra = 1/3 1/6 2/12 3/24 0
rm 1/3 1/6 1/12 2/24 0
Iteration 0, 1, 2,
Here the PageRank leaks out since the matrix is not stochastic.
Teleports: Follow random teleport links with
probability 1.0 from dead-ends
Adjust matrix accordingly
y y
a m a m
y a m y a m
y 0 y
a 0 0 a 0
m 0 0 m 0
Googles solution that does it all:
At each step, random surfer has two options:
With probability , follow a link at random
With probability 1-, jump to some random page
This formulation assumes that has no dead ends. We can either
preprocess matrix to remove all dead ends or explicitly follow random
teleport links with probability 1.0 from dead-ends.
PageRank equation [Brin-Page, 98]
1
= + (1 )
[1/N]NxNN by N matrix
The Google Matrix A: where all entries are 1/N
1
=+ 1
We have a recursive problem: =
And the Power method still works!
What is ?
In practice =0.8,0.9 (make 5 steps on avg., jump)
M [1/N]NxN
7/15
y 1/2 1/2 0 1/3 1/3 1/3
0.8 1/2 0 0 + 0.2 1/3 1/3 1/3
0 1/2 1 1/3 1/3 1/3