stationary distributions and Page's Jr. (star!)
If today were the last day of my life (see previously cited practice), would I still be thinking about how to fairly sample stationary (cough PageRank) distributions?
Yes, in that I feel that understanding how to be "fair" and "representative" in a technical graph theoretic sense could help guide our technology and technologists to be fair by some magic transitivity.
No, in that most of the people closest to me and most of the people in the world wouldn't directly benefit from clear thinking in this matter.
They would, however, indirectly benefit from fairer ideologies and pluralities (to use Mr. Brokaw's term), and from commodity online services which think about fairness and representation from the get go.
And so here it is, a brief thought on sampling stationary (cough PageRank) distributions, clarified from my coarse insight from yesterday's meeting. I won't talk about what I like to call (as of this morning :) PageRank Junior (or more specifically Random Surfer II, the second that is) -- haha -- since that directly builds on a fellow student's current/unpublished work, but let me talk about Tom Cruise, my unpublished Master's thesis, and Mr. Page's theoretical son (oops, I'm not supposed to talk about him).
My inchoate insight yesterday was that thinking about the problem of picking the most representative vertices of some social-or-whatnot network is sometimes confused by the seemingly obvious reduction to combinatorial sorting algorithms. To be more concrete, if I consider networks whose degree (resonant salience rather?) distribution follows a Pareto distribution, movie collaboration graphs let's (hypothetically) say, then in this sick (not to say mentally ill) world you might expect Tom Cruise to show up somewhat often when you non-randomly cruise through the graph of movies and their actors, biased by what and who people talk about.
Okay, you say, so Tom Cruise seems like some kind of representative actor in that network, by some yearbook (facebook?) metric of popularity. But you don't want to hear about Tom Cruise and his clones all the time, and I hear you, non-random surfer, I hear you. And you'd be right.
As I think about it now, one of the fallacies I fell into in my previous work was the act of too easily equating the set of "important" vertices to a "representative" set of vertices. More concisely and ambiguously, a set of (individually) "representative" vertices is not a (collectively) "representative" set of vertices (in the sense that they (a top-k set of representative vertices) don't fairly represent the whole graph's stationary distribution, and thus they don't really maximize some fairness-in-representation functional). Duh?! *faceplant*
In other words, even though Tom Cruise and his hypothetical Hollywood clone army (of actors as popular as he) are individually representative as movie stars, only picking these most buzzworthy actors (sampling the hot body of the Pareto distribution?) to be collectively representative of the space of movies and actors seems like a mistake.
What to do instead (cough stationary-distribution-biased sampling for constructing DK-type hierarchies?) is left as an exercise to you, gentle reader, and to me, a not so gentle reader.
Yes, in that I feel that understanding how to be "fair" and "representative" in a technical graph theoretic sense could help guide our technology and technologists to be fair by some magic transitivity.
No, in that most of the people closest to me and most of the people in the world wouldn't directly benefit from clear thinking in this matter.
They would, however, indirectly benefit from fairer ideologies and pluralities (to use Mr. Brokaw's term), and from commodity online services which think about fairness and representation from the get go.
And so here it is, a brief thought on sampling stationary (cough PageRank) distributions, clarified from my coarse insight from yesterday's meeting. I won't talk about what I like to call (as of this morning :) PageRank Junior (or more specifically Random Surfer II, the second that is) -- haha -- since that directly builds on a fellow student's current/unpublished work, but let me talk about Tom Cruise, my unpublished Master's thesis, and Mr. Page's theoretical son (oops, I'm not supposed to talk about him).
My inchoate insight yesterday was that thinking about the problem of picking the most representative vertices of some social-or-whatnot network is sometimes confused by the seemingly obvious reduction to combinatorial sorting algorithms. To be more concrete, if I consider networks whose degree (resonant salience rather?) distribution follows a Pareto distribution, movie collaboration graphs let's (hypothetically) say, then in this sick (not to say mentally ill) world you might expect Tom Cruise to show up somewhat often when you non-randomly cruise through the graph of movies and their actors, biased by what and who people talk about.
Okay, you say, so Tom Cruise seems like some kind of representative actor in that network, by some yearbook (facebook?) metric of popularity. But you don't want to hear about Tom Cruise and his clones all the time, and I hear you, non-random surfer, I hear you. And you'd be right.
As I think about it now, one of the fallacies I fell into in my previous work was the act of too easily equating the set of "important" vertices to a "representative" set of vertices. More concisely and ambiguously, a set of (individually) "representative" vertices is not a (collectively) "representative" set of vertices (in the sense that they (a top-k set of representative vertices) don't fairly represent the whole graph's stationary distribution, and thus they don't really maximize some fairness-in-representation functional). Duh?! *faceplant*
In other words, even though Tom Cruise and his hypothetical Hollywood clone army (of actors as popular as he) are individually representative as movie stars, only picking these most buzzworthy actors (sampling the hot body of the Pareto distribution?) to be collectively representative of the space of movies and actors seems like a mistake.
What to do instead (cough stationary-distribution-biased sampling for constructing DK-type hierarchies?) is left as an exercise to you, gentle reader, and to me, a not so gentle reader.
0 Comments:
Post a Comment
<< Home