Showing posts with label wikipedia. Show all posts
Showing posts with label wikipedia. Show all posts

Friday, September 3, 2010

Revamping WikiDashboard

I released WikiDashboard almost three years ago. Believe it or not, the server for WikiDashboard has been running under my desk for three full years (the photo shows the actual server). It was launched in a rush to meet a deadline for an academic paper that we published at a conference (ACM SIGCHI 2008) and limited maintenance has been done so far.

The old Power Mac (http://en.wikipedia.org/wiki/Power_Mac_G5 ) has been pretty reliable but it is becoming increasingly untrustworthy lately. Frustrated with frequent crashes, hangs, and sluggishness, I finally decided to do something. As I’m migrating the tool out of the old machine, I’ve added a few new features. I hope you find it useful.


Faster and scalable infrastructure
The server is now running on Google App Engine. WikiDashboard is hosted as a web app on the same systems that power Google applications. WikDashboard should provide faster, reliable, and scalable service to you. I plan to keep the old server running for a bit but it will eventually forward the traffic to the new server.

Support ten more languages
Thank you to everyone who showed interest in having WikiDashboard in your own language version!

Bongwon Suh
http://www.parc.com/suh
@billsuh http://twitter.com/billsuh

Monday, March 8, 2010

Wikipedia's People-Ware Problem

Last week, we hosted a visit from the Wikimedia Foundation on issues relating to our work on community analytics, and what it tells us about Wikipedia's problems and possible solutions. Naoko Komura (pictured at right) of the Wikimedia Usability Initiative, as well as Eric Zachte, the staff data analyst (also pictured at right), spoke very eloquently about how we can create social tools to direct the best social attentions to the needed parts of Wikipedia.



Fundamentally, Wikipedia has always had a "people-ware" problem: the distribution of the expertise that is freely donated to the right places.  It has been and always will remain its greatest challenge. The amazing thing about Wikipedia is that it managed to do this for so long, such that a valuable knowledge repository can be built up as a result.  At first, people simply came because it was the place to be.  Now, we have to work a little harder.

We spent a lot of time talking about the best way to model this people-ware problem, either using biological metaphors (evolutionary systems with various forces), or economic models (see last post here).  However, one thing to be aware of is the danger of "analysis paralysis", where you spend so much time analyzing the problem, and forget that there are already many ideas that have been generated for moving the great experiment forward.

For example, there are many places in Wikipedia that are not well populated. It's well-known that many scientific and math concept articles, for example, could use an expert-eye to catch the errors and explain the concepts better. How can we build an expertise finder that would actually invite people to fix problems that we know exists in Wikipedia?

Another idea might be to have the whole system be more social. Chris Grams blogs about a part of this idea here. We suggested some time ago to have a system like WikiDashboard, where you actually show the readers what the social dynamics have been for a particular article.

Wikipedia was created in 2001, when social web was still in its infancy. During the ensuing 9 years, it has changed very little, and I would argue Wikipedia have not kept up with the times. Lots of "Social Web" systems and new cultural norms have been built up already.  For example, I suspect that many of us would not mind at all to reveal our identities on Wikipedia, and we might like to login with our OpenIDs and even have verified email addresses so that the system can send me verification/clarification/notification messages. The system perhaps should connect with Facebook, so that my activities (editing an article on "Windburn") is automatically sent to my stream there. My friends, upon seeing that I have been editing that article, might even join in.

I think that Wikipedia is about to change, and it is going to become a much more socially-aware place. I certainly hope that they will tackle the People-Ware (instead of the Tool-Ware) problems, and we will see it become an exciting place again.

Tuesday, September 22, 2009

PART 3: Population Shifts in Wikipedia

The research done at ASC continues to get more press, including Time magazine, NYTimes, Repubblica [Italian Newspaper]. We have been busy trying to put together a bunch more academic papers on Web2.0 (particularly some Twitter research we have been doing), so we haven't updated this blog in a while. I figure today I'd take some time and blog a bit more about our results.

To investigate which factors affected the slowdown in edit growth, we examine the evolution of the population of active editors. The stalled growth of edit activities that we have described might be partially explained by changes in the editor population. We use the same editor classification as previous posts to count the number of active editors in each month. The figures below show three views of the evolution of the population of the five editor classes.



Monthly active editors by editor class. (This is a breakdown of the total editor population depicted earlier)

The Figure above shows the monthly frequencies of active editors by class. As expected from the power law distribution, the distribution of editors is very skewed: most of the editors contribute very few edits and very few editors contribute most of the edits. In fact, the two most prolific classes of editors (100-999 and 1000+) account for only about 1% of the population, but they contribute about 55% of edits (33% and 23% respectively).


Monthly active editors by user class. The vertical axis uses a logarithmic scale.


The Figure above uses a logarithmic scale to show the consistent slowdown of the growth among all editor classes over time, which is not clear in the first figure for editors in 100-999 and 1000+ classes. The monthly population of active editors stops growing after March 2007: a surprisingly abrupt change in the evolution of the Wikipedia population for all the editor classes. This change is consistent with the slowdown of the editing activity shown in Part 1.

[Interesting enough, even though we see that that the number of 1000+ class of editors plateaued, we know from Part 2 that this class of users have been increasing their contribution rate. Their average monthly edits per editor for the years 2005 to 2008 were 1740, 1859, 1869, and 2095, respectively.]


Percentages of monthly active editors by their class. Note that the graph is truncated to highlight the declining population of 10-99 editor classes [shown in purple]. (Sorry that the coloring of the editor classes is not consistent from the earlier plots.)

The last Figure shows the percentage of monthly active editors among the five classes. Note that the Y-axis is truncated: it omits the bottom 50% which represents the very long tail of once-monthly-editors. Notice how the 10-99 editor class [shown in purple] is being squeezed and becoming a small portion of the overall population. The 10-99 editor class went from 9% in 2005 to 6% in 2008.

A healthy community requires that people can move from novice contributors to occasional contributors to elite contributors. In other words, the upward mobility of the contributors is important for a healthy community. The trend here suggest that there are some resistance in moving beyond the 10-99 edits/month barrier. Could this be evidence of the Wiki-lawyering barriers?

One theory that I might suggest is that we want a well-balanced pyramid structure in the community population. Not too top heavy, and not too bottom heavy, and with a healthy middle class. How can we design the mechanisms [incentives and appropriate barriers] on the site so that we have this structure?

Friday, August 7, 2009

PART 2: More details of changing editor resistance in Wikipedia

In the last week, we have received interesting press coverage in New Scientist (as well as Fast Company, Business Insider, and syndicated elsewhere), on the work done in our team on Wikipedia growth rate, and how it has plateaued, changing from an exponential growth model to one that look more linear. Even though this wasn't necessarily new finding, but it was really a teaser for some other observations we have found in the Wikipedia data that is about to be published in WikiSym2009 conference in October.

In the figure below, we see how the slowdown in growth of Wikipedia activity, specifically around different editor classes is different. For each month, we first partition the editors into different classes based on their monthly editing frequency. We then compare the total edit activities among the different editor classes over time.


Monthly edits by user class (in thousands).


[Consistently with the power law, we classified users using an exponential scale: we defined the classes of editors using powers of 10, e.g. 10^0, 10^1, 10^2. This resulted in five classes of users for each month: editors contributing 1 edit (i.e., 10^0), 2 to 9 edits (2-9 class), 10 to 99 (10-99 class), 100 to 999 (100-999 class), and more that 1000 edits (1000+ class).] Note that the classification of the editors was recalculated for each month.

Since the beginning of 2007, the trends of four classes slightly decrease their monthly edits. In contrast, only the highest-frequency class of editors (1000+ edits, dark blue line) shows an increase in their monthly edits.

Another way to look at this data is to analyze the relative amount of activities for each editor class by transforming the data into percentages of the total edits. The figure below complements the information in the figure above by showing the percentage of the volume of edits that each class contributes in relation to the total.


Monthly percentage of edits by each user class.

The two highest frequency classes of editors account for more than half of the total monthly edits (56% from 01/2005 to 08/2008). Furthermore, since 2005 the proportion of contributions by the highest-frequency editor class has increased slightly. In fact, the editors in 1000+ class have kept producing at an increasing rate over the past four years (their average monthly edits per editor for the years 2005 to 2008 were 1740, 1859, 1869, and 2095, respectively).

We now focus on specific evidence about what might have contributed to such slowdown. Revert is the action of deleting a prior edit. The following figure shows the percentage of edits that were reverted (reverted edits) monthly for each editor class. Note that edits related to vandalism and edits performed by robots are excluded.


Monthly ratio of reverted edits by editor class

This illustrates two indicators of a growing resistance from the Wikipedia community to new content.

First, the figure shows that the total percentage of edits reverted increased steadily over the years. The total percentage of monthly reverted edits (see dashed black line) has steadily increased over the years for the all classes of editors (e.g. 2.9, 4.2, 4.9, and 5.8 percent of all edits for 2005 through 2008 as shown by the dash line).

Second, more interestingly, low-frequency or occasional editors experience a visibly greater resistance compared to high-frequency editors [see the top two reddish lines, as compared to other lines]. The disparity of treatment of new edits from editors of different classes has been widening steadily over the years at the expense of low-frequency editors.

We consider this as evidence of growing resistance from the Wikipedia community to new content, especially when the edits come from occasional editors.

Wednesday, July 22, 2009

PART 1: The slowing growth of Wikipedia: some data, models, and explanations

In September of 2008, we blogged about a curious change in Wikipedia that we didn't know how to explain that we had known for a while, and the ASC group has been looking into understanding this change in the last 6-9 months or so. The change that we were curious about was that the growth rates of Wikipedia have slowed. We were not the only ones wondering about this change. The Economist (archived here), for example, wrote about it.

We are about to publish a paper in WikiSym 2009 on this topic, and I thought we should start to blog about what we found.


Monthly edits and identified revert activity

The conventional wisdom about many Web-related growth processes is that they're fundamentally exponential in nature. That is, if you want some fixed amount of time, the content size and number of participants will double. Indeed, prior research on Wikipedia has characterized the growth in content and editors as being fundamentally exponential in nature. Some have claimed that Wikipedia article growth is exponential because there is an exponential growth in the number of editors contributing to Wikipedia [1]. Current research show that Wikipedia growth rate has slowed, and has in fact plateaued (See figure at right). Since about March of 2007, the growth pattern is clearly not exponential. What has changed, and how should we modify our thinking about how Wikipedia works? Prior research had assumed Wikipedia works on a "edit begets edit" model (That is, a preferential attachment model where the more an article gets edits, the more likely it would receive more edits, and thus resulting in exponential growth [2].) Such a model does not preclude some ultimate limitation to growth, although at the time it was presented [2] there was an apparent trend of unconstrained article growth.


Monthly active editor - number of users who have edited at least once in that month


The number of active editors show exactly the same pattern. The 2nd figure on the right shows how since its peak in March 2007 (820,532), the number of monthly active editors in Wikipedia has been fluctuating between 650,000 and 810,000. This finding suggests that the conclusion in [1][2] may not be valid anymore. We have a different process going on in Wikipedia now.


Article growth per month in Wikipedia. Smoothed curves are growth rate predicted by logistic growth bounded at a maximum of 3, 3.5, and 4 million articles.

Some Wikipedians have modeled the recent data, and believe that a logistic model is a much better way to think about content growth. Figure here shows that article growth reached a peak in 2007-2008 and has been on the decline since then. This result is consistent with a growth processes that hits a constraint – for instance, due to resource limitations in systems. For example, microbes grown in culture will eventually stop duplicating when nutrients run out. Rather than exponential growth, such systems display logistic growth.

We will continue to blog about what we believe might be happening in the next few weeks, as we find time to summarize the results.

[1] Almeida, R.B.m, Mozafari, B., and Cho, J., On the evolution of Wikipedia. ICWSM 2007, Boulder, Co., 2007.
[2] Spinellis, D., and Panagiotis, L. The collaborative organizations of knowledge. Communications of the ACM, 51(8), 68-73, 2008.

Monday, July 20, 2009

Social attention and interactions are key to learning processes


I just finished reading a long article in the journal Science on how social factors are increasing recognized as extremely important in a new science on learning [1].

Learning is fundamentally a social activity, the article partially argued. "Social cues highlight what and when to learn." Meltzoff et al. summarize a whole slew of recent research that showed how young infants learn by imitation and copying others actions, and they build abstractions and models of others' behaviors. In fact,
"Children do not slavishly duplicate what they see but reenact a person’s goals and intentions. For example, suppose an adult tries to pull apart an object but his hand slips off the ends. Even at 18 months of age, infants can use the pattern of unsuccessful attempts to infer the unseen goal of another. They produce the goal that the adult was striving to achieve, not the unsuccessful attempts."


One point made in the article is how much the greater environment outside of school is becoming an important part of the ecology of learning.
"Elementary and secondary school educators are attempting to harness the intellectual curiosity and avid learning that occurs during natural social interaction. The emerging field of informal learning is based on the idea that informal settings are venues for a significant amount of childhood learning. Children spend nearly 80% of their waking hours outside of school. They learn at home; in community centers; in clubs; through the Internet; at museums, zoos, and aquariums; and through digital media and gaming."


Social learning, of course, is a major part of the social web. Wikipedia was designed to be an easy-to-use and freely available reference, and all of the social interactions offered by various online forums are rapidly becoming a part of the educational experience for secondary school pupils. I would argue, for example, that Wikipedia has done more for continuing education for all adult learners than any educational institution could have done by itself. ASC's research have purposefully been focused on learning and information access, instead of entertainment, because of our recognition of the importance of social factors in various kinds of learning.

As an example, social learning was explicitly part of the design of our SparTag.us prototype, which is now just being offered in limited beta software to Firefox users, was announced at the recent CHI2009 conference. It streams the annotations you make as you browse the web. The stream is collected into your notebook, and by default this stream of annotation is made available to anyone interested in it. This makes it possible to aggregate social attention later.


[1] Foundations for a New Science of Learning. A. N. Meltzoff, P. K. Kuhl, J. Movellan and T. J. Sejnowski. Science, 325 (5938), 284-288. [DOI: 10.1126/science.1175626].

[2] Photo: Alan Decker and the Machine Perception Lab, UC San Diego.

Monday, June 29, 2009

Live data again: WikiDashboard visualizes the editing patterns of 'David Rohde' case...

Yesterday, NYTimes finally broke the silence on the kidnapping of David S. Rohde by the Taliban. Turns out, Rohde had escaped, and that the news media finally reported the kidnapping since the publicity on the case would no longer be a bargaining chip for his captors. The NYTimes article showed how keeping this news off of Wikipedia was nearly impossible if it weren't for the coordinated effort of several administrators and Jimbo Wales himself.

WikiDashboard visualized this editing pattern directly. In the figure below, I've highlighted the various edit wars between the anonymous editors (97.106.51.95; 97.106.45.230; and 97.106.52.36, which are believed to be the same person) and some of the administrators such as Rjd0060 and MBisanz and the involvement of a robot XLinkBot. You can also see the huge attention on this article in the last week or so in the visualization.


Check out the editing history and the edit war in detail by reading the edit history.

All of this makes for a great way for us to announce that WikiDashboard now works on the live Wikipedia data again; Thanks to the heroic efforts of Bongwon Suh in my group. He figured out how to execute his SQL query in a quick way on the new DB server.

Thursday, April 16, 2009

Mapping the Contents in Wikipedia

Having just returned from CHI2009 conference on Human-Computer Interaction, many of the topics there focused on where and how people obtain their information, and how they make sense of it all. A recent research topic in our group is understanding how people are using Wikipedia for their information needs. One question that had constantly come up in our discussion around Wikipedia is what is exactly in it. We have so far done most of our analyses around edit patterns, but not so much analysis have gone into what do people write about? What topics are the most well-represented? Where topic areas have the most conflict?

In one of our recent CHI2009 papers, we explored this issue. Turns out that Wikipedia have these things called Categories, which people use to organize the content into a pseudo-hierarchy of topics. We devised a simple path-based algorithm for assigning articles to large top-level categories in an attempt to understand what topic areas are the most well-represented. The top level categories are:



Using our algorithm, the page "Albert Einstein" can be assigned to these top-level categories:


This mapping makes some intuitive sense. You can see that the impact Albert Einstein has made in various areas of our society such as science, philosophy, history, and religion. Using the same ideas and algorithm, we can now do this mapping for all of the pages in Wikipedia, and find out what top level categories have received the most representation. In other words, we can figure out the coverage of topic areas in Wikipedia.


(You may have to click on the graphic here to see it in more detail.)

We can see that the highest coverage has gone toward the top-level category of "culture and the arts" at 30%, followed by "people" 15%, "geography" 14%, "society and social science" 12%, and history at 11%. What's perhaps more interesting is understanding which ones of these categories have generated the most conflicts! We used the previously developed concept called Conflict Revision Count (CRC) in our CHI2007 paper, and showed which top level categories have the most conflicts:



In this figure, the categories are listed in order of the total amount of conflicts clockwise from "People". This means that People did receive the most amount of conflict, followed by Society and Social Sciences, etc. However, the percentages in each topic is normalized by the number of article-assignments in that topic. So the metric developed here can be interpreted as the amount of conflict in each topic that has been normalized by the size of the topic, which can be interpreted as the amount of contentious in articles of the topic.

"Religion" and "Philosophy" stand out as highly contentious despite having relatively few articles.
Turns out that "philosophy" and "religion" have generated 28% of the conflicts contentious-ness each. This is despite the fact that they were only 1% and 2%, respectively, of the total distribution of topics as shown above.

Digging into religion more closely, we see that "Atheism" have generated the most conflict, followed by "Prem Rawat" -- the controversial Guru and religious leader, "Islam" and "Falun Gong".



Wikipedia is the 8th ranked website in the world, so it is clear that a lot of people get their information from Wikipedia. The surprising thing about Wikipedia is that it succeeded at all. Common sense would suggest that an encyclopedia in which anyone can edit anything they want would result in utter nonsense. What happened is exactly the opposite: Many users and groups have gotten together to make sense of complex topics and debate with each other about what information is the most relevant and interesting to be included. This helps with us keeping sane in this information world, because we now have a cheap and always accessible content on some of the most obscure content you might be interested in. At lunch today, we were all just wondering what countries have the lowest birth rate. Well, surprise!! Of course, there is a page for that, which we found using our iPhones.

The techniques we have developed here enable us to understand what content is available in Wikipedia and how various top level categories are covered, as well as the amount of controversy in each category.

There are of course many risks in using online content. However, we have been researching tools that might alleviate these concerns. For example, WikiDashboard is a tool that visualizes the social dynamics behind how an wiki article came into its current state. It shows the top editors of any Wikipedia page, and how much they have edited. It also can show the top articles that a user is interested in.

We are considering adding this capability to WikiDashboard, and would welcome your comments on the analysis and ideas here.

All web users can guide the content in Wikipedia by participating in it. If we realized that the existence of our society depends on the healthy discourse between different segments of the population, then we will see it not just as a source of conflict, but a source of healthy discussion that needs to occur in our world. By having these discussions in the open (with full social transparency), we can ensure all points of view are represented in this shared resource. Our responsibility is to ensure that the discussion and conflicts are healthy and productive.



Reference:
Kittur, A., Chi, E. H., and Suh, B. 2009. What's in Wikipedia?: Mapping Topics and Conflict using Socially Annotated Category Structure. In Proceedings of the 27th international Conference on Human Factors in Computing Systems (Boston, MA, USA, April 04 - 09, 2009). CHI '09. ACM, New York, NY, 1509-1512.

Monday, January 26, 2009

Governing and authorship models at Wikipedia and Britannica

Elsewhere, we have spoke about the complex and interesting governing and authorship model in Wikipedia. How counter-intuitive is it that a model like "anyone can edit anything they want" could produce a useful information resource?!

We have conducted some characterizations of the social dynamics within this community, and tracked its changes over time. Interestingly, in the last few days, both Wikipedia and Britannica have been in the news for debates on their stance of the authorship and editorial model.

First, on Jan 24th, we learned from the BBC that the president of Britannica wrote a blog entry in which he outlined a new plan at Britannica to enable readers as well as more experts and editors to help expand and maintain the articles. While not naming Wikipedia by name, it was a clear nod toward a more collaborative relationship Britannica will have with its readers. Specifically, in the blog entry, Jorge Cauz says that "We believe that the creation and documentation of knowledge is a collaborative process but not a democratic one." Most would agree that, in the past, the collaborative process that Britannica had was much more restrictive, and now they seem to have decided to open the door wider to include more people in the editorial process.


Then, today, we learned, also from BBC, that Jimmy Wales have caused a huge stir at Wikipedia for suggested a more restrictive approach to the editing process. He now believes that Wikipedia should follow a model in which edits from anonymous users have to be vetted by one of the site's editors before becoming live.

Apparently, the heated debate is now spreading, and is being mentioned as a big news item on the Yahoo! front page after being written up by AFP. So here we have a system that has been extremely liberal with its editorial policy moving toward a more restrictive authorship model.

So what gives? Is there a right way or wrong way to constructing and compiling knowledge resources? As designers of social systems, what should be the governance model for these systems?

For one thing, we still know awfully little about the social dynamics in these large social systems. We have been quoted in the past that our characterization models of editors show that the top 1% of the editors in Wikipedia generates 50% of the edits. While that is true, the other 50% is being generated by the other 99% of the editors. This other 50% is just as important as the first 50%!

We have been recently conducting some additional research to understand class structures in Wikipedia. We already know that the distribution of editors and their frequency of edits in Wikipedia is a classic power law curve. In order to understand editors through out this distribution, we first ranked editors by their edit frequency, and then divided all of the edits into four quarters, according to this sort.

For one month worth of edit data, there are about 220 editors that are at the very top of the pyramid. These top (most frequent) editors produce the first quarter (25%) of the edits. The next 25% of the edits come from about 1000 editors. While the 3rd quarter of edits come from about 4000 editors, and the last quarter comes from about 15000 editors.

So now the research question is whether you want to design your editing policy to favor the upper class (top editors and administrators), the middle class (the 5000-6000 editors who contribute the middle 50% of all edits), or the lower class (the 15000 editors who contribute the last 25%).

One way to think about this problem is to study the amount of resistance each of these four classes of editors experience on Wikipedia. A metric that we used is the reverts-to-edits ratio. That is, on average, what percentage of edits were reverted, as experienced by each of these four classes of editors? Turns out that the reverts-to-edits ratio for each of these 4 classes of editors were 1.3%, 1.4%, 1.5%, and 4.7%, respectively. Meaning that the lower class of editors clearly experience greater resistance, such that, on average, 1 out of every 20 edits they contribute are reverted. Moreover, the resistance they have experienced have generally increased over time (from about 3% in early 2006 to 5-6% in 2007-2008, and back down to around 5% in late 2008).

So, even without the "flagged revision" mechanism such as the ones suggested by Jimmy Wales, it has already been getting harder for the lowest class of occasional editors to produce edits that remain as contribution in Wikipedia.

The AFP article points to the fact that the debate over the policy came about because of vandalism on Ted Kennedy's page, which had falsely suggested he died after suffering a collapse at a lunchon during Obama's inauguration. But apparently this was corrected within minutes, suggesting that the current system is still correcting most mistakes quite rapidly. Moreover, after I did some sleuthing in the editing history, it appears that the original vandalism edit was done by a registered user named "Gfdjklsdgiojksdkf", and not an anonymous user.

So, it is unclear to me that the current system is not working. Are we fixing something that isn't broken (at least not yet)?

Thursday, October 16, 2008

A new live-data version of WikiDashboard for Wikipedia



The ASC group (and Bongwon Suh in particular) is pleased to announce a new version of WikiDashboard for Wikipedia. In this new version, we have:

* Live Information!
WikiDashboard now uses the live feed of English Wikipedia powered by MediaWiki Toolserver. The dashboard will show any changes made on each page almost instantly. Note that the earlier version has been showing information as of April 2008. For example, you can see who's been active in pages such as: Sarah Palin's page or the US President Election page.

Notice in particular how Sarah Palin's edits really only picked up in the last 6-8 weeks, but User Ferrylodge had edited her page around July 1, before the Aug. 29th nomination.

Unfortunately, because the Toolserver is not very reliable on our queries, we are not always able to serve up live edit data in our dashboard. If you don't get a live dashboard, you can either get at the data from April 2008 that is on our own private database server, or you can wait a while and try again.

* Browse Through Time
Now, you can click on the bars in the dashboard. Clicking on an bar will bring you to the wiki historical context when the edits were made. For Article Dashboard, the system will show all the edits made on the page around the time point you choose. For User Dashboard, WikiDashboard will provide a list of edits that the user made around the time you clicked.

Please let us know if you find any problem or have any feedback. Thanks!

Thursday, August 14, 2008

Anonymous / Pseudonym edits in Wikipedia: a good idea still?


In the last few months, a somewhat sticky issue around the use of pseudonym occurred on this blog. A writer for a newspaper called SF Weekly, was being attacked online for writing an article about editor wars, in which she focused on an Wikipedian named "Griot". We blogged about this article, and a bunch of both anonymous comments as well as pseudo-anonymous comments ensued. I was hesitating about stepping in to censor the comments, since our research very much believed in "social transparency". This means presenting all of the information for everyone to see, and letting the social process sort out the truth.

Yesterday, the Electronic Frontier Foundation helped Wikipedia win an important lawsuit, which "found that federal law immunizes the Wikimedia Foundation from liability for statements made by its users." An interesting question is whether this includes _all_ statements, or just some of it. What if someone pretends to be someone else (which happened in the comments section of our blog post)? If I obtained a handle (pseudonym) of BillGates or BarackObama, and pretended to be him, can I really say anything I want? What about libel, slander, and defamation?

How far does anonymity gets us in eliciting all of the material that needs to be said? And how damaging is it to have it as part of Wikipedia? What about the use of pseudonyms? These are interesting research questions. Giant experiments like Citizendium are trying to answer some of these questions. What about different degrees of pseudonym like non-disposable pseudonym vs. disposable pseudonym, or pseudonyms that resolve to a real person and a real name under court order? (Disposable pseudonyms are handles that you can throw away easily and simply obtain a new one; blogger.com here has this option in the commenting feature, for example.)

In the spirit of "social transparency", I believe that disposable pseudonym can be quite destructive to an online community. When accountability is not maintained, quality of the material is suspect. "Social transparency" means an increase in accountability. It's a form of a reputation system. Some researchers are suggesting that online accountable pseudonyms is the way to deal with these identity problems. SezWho, Disqus are examples of how to deal with these reputation and identity problems in the blog comment space. I think it is inevitable that we will need better reputation and identity systems on the web.

As linked above, a good discussion about pseudonyms can be found in:
An Offline Foundation for Online Accountable Pseudonyms. by Bryan Ford and Jacob Strauss. In the Proceedings of the First International Workshop on Social Network Systems (SocialNets 2008), Glasgow, Scotland, April 2008.

Monday, May 5, 2008

Announcing a new release of WikiDashboard with updated dataset

Reputation systems are deeply important to social websites. For example, many users use Facebook or bookmarking systems to insert themselves in the middle of information flow, thus gaining positions as information brokers.

A recent Scientific American article highlighted recent research on the effects of reputation in the brain. The fMRI studies cited showed that "money and social values are processed in the same brain region". Thanks goes to Bob Vasaly for pointing this research out to me.

Indeed, one of the intended uses of WikiDashboard was the ability for readers and editors alike to assess the reputation and behaviors of editors in the system. For example, we can take a look at the actual behavior of a controversial editor named Griot that was at the center of a controversy in the SF Weekly, and make decisions on our own about the actual patterns of edits depicted there. Or take as another example of Jonathan Schilling, who "protects Hillary's online self from the public's hatred. He estimates that he spends up to 15 hours per week editing Wikipedia under the name "Wasted Time R"--much of it, these days, standing watch over Hillary's page."

Our goal here is not to make decisions for you, but to make the social and editing patterns available to the community so that you can make decisions on your own. In an effort to do that and in preparation for the CHI2008 conference, Bongwon recently updated the Wikipedia database and we now have fresh data to share with the community. The new database now consist of nearly 3.5 terabytes of raw revision data that we process.

The new interface also has a connection to reddit.com so that users can submit interesting WikiDashboard views that they have found interesting.

Let us know what you all think!


Bongwon Suh, Ed H. Chi, Aniket Kittur, Bryan A. Pendleton. Lifting the Veil: Improving Accountability and Social Transparency in Wikipedia with WikiDashboard. In Proceedings of the ACM Conference on Human-factors in Computing Systems (CHI2008). (to appear). ACM Press, 2008. Florence, Italy.

Tuesday, February 12, 2008

Wikipedia users: We’d like to talk with you!

We're conducting ethnographic interviews on Wikipedia use, to help us create better tools for both readers and editors. Share your experiences and give us your opinions! The interview takes about an hour, can be done remotely or at PARC, and we can schedule at your convenience. We'll even give you an Amazon gift certificate as a token of our appreciation. Please contact dschiano@parc.com.
Thanks!

Thursday, January 17, 2008

Risks in using Wikipedia?

In our research on Wikipedia, we have been using a broad framework published as an one-page article in CACM (Communication of the ACM) in 2005 by Denning et al. on the perceived risks in using Wikipedia contents. The question before researchers is how to mitigate these risks while enabling a vibrant social community who wants to get together to build a encyclopedia and help each other obtain knowledge. As Denning's article mentions: "But will this process actually yield a reliable, authoritative reference encompassing the entire range of human knowledge?"

In the framework in thinking about answers to this question, Denning's article suggests that numerous risks that we should consider:

  • Accuracy: how can you be sure that the information in the article was actually accurate and not some misrepresentation of the fact?

  • Motives: how can you be sure of the motive of the editors were to present the facts and only the facts, and not opinions? For example, as we have discussed before, how can we be sure that editors to political candidate pages are not there to simply push their political agenda? For example, what does User:Jasper23's editing history tell you about his potential political positions?

  • Uncertain Expertise: How do we determine the expertise levels of the people who are editing in Wikipedia? Some appears to really know what they're talking--for example, User:BillCl was the top editor of the NASA page, and appears to have done a bunch of edits around aviation related topics. On the other hand, the top editor (Wasted Time R) of the Hillary Clinton page appears to be also to be big fans of Bruce Springsteen, Elton John, Dixie Chicks, Tony_Bennett, and other musicians. So it is less clear this user is an expert on political positions of Hillary Clinton as compared to other candidates.

  • Volatility: If the topic was just in the middle of a huge debate, then the content of the article could be really unsettled.

  • Coverage: Coverage of certain topics appears to be better in Wikipedia than others. What are the inclusion and exclusion standards is less than clear.

  • Sources: Many articles do not cite authoritative sources, so it is hard to trace and find out if the information is actually accurate.


The article then goes on to say that "[WP] cannot attain the status of a true encyclopedia without more formal content-inclusion and expert review procedures." Our WikiDashboard tool is precisely designed to help with collaborative review of Wikipedia editing history and patterns.

Wednesday, December 12, 2007

WikiDashboard search engine plugin for Firefox/Mozilla Browsers

So my good friend, Jeff Heer, wrote and asked if there is a search engine plugin for the Firefox browser for WikiDashboard after I posted the WikiDashboard Bookmarklet last night. Having 15 minutes to kill, I looked up the documentation and wrote up this little plug-in for exactly this purpose here.

Simply click on the link below to install the WikiDashboard search engine plug-in for your Firefox browsers. Use at your own risk!

Click here to install the WikiDashboard Search Engine Plugin

Friday, October 5, 2007

Social transparency and the quality of co-created contents

How do you measure the accuracy and quality of what people are collectively creating? For example, on Yahoo! Answers, people post questions and tons of people respond. How would you measure the quality of the content?

What’s amazing about this as a research area is that it starts to touch on deep classic philosophic questions like: What do we know about authority? What does it mean? Where does authority come from? What makes someone trust you? When you ask a question about the quality of any information, you have to answer these questions. Who is the person who wrote it? Why should I trust that person? Just because Encyclopedia Britannica hires a bunch of experts to write for them, why should I believe them? What makes them an authoritative figure on how bees build their beehives? What is it about their authority, just because they’re attached to some higher education institution, that makes you want to believe them more than someone else?

When the Augmented Social Cognition research group tried to answer these questions, we ended up with an internal debate about what we mean by “quality.” And I think we come up with a model for understanding quality. We realized that, in academia, much of authority and the assignment of trust actually comes from transparency. Why should I believe in calculus? Well, because the mathematics is built on a foundation of axioms and rule sets that you can follow, which you can look up and examine. You trust calculus because there is a transparency built into the system. You can come to your own conclusion about the quality of the information based upon an examination of the facts. This is the scientific method!

What’s interesting is that exactly the same argument is being applied to Wikipedia. It says to you: you should believe in the quality of the information in Wikipedia because it’s transparent. Anyone can look at the editing history and see who has edited an entry, whether they chose to sign their name after it, and what kind of edits they made in other parts of Wikipedia. Everything is transparent and completely traceable; you can examine Wikipedia back to the first word that was written. And Wikipedia is relying on the fact that it’s completely transparent to gain authority. There is nothing opaque about it. I think that’s why Wikipedia has become so successful. It’s because they stumbled upon some of these fundamental design principles and paradigms that makes this work. They could have made the design decision where one can only examine the last 50 edits. Wikipedia could have come up with many other design choices that would not make the system completely transparent. Is it an accident that they ended up with a system that can be traced back to the first edits? I think not.

However, (and that's a big however!), some people are still having trouble with the quality of information on Wikipedia even though it’s transparent. Why? One possiblity is that they have an all-or-nothing attitude. Well, if one article could be way-off, why should I trust another article? They don't, and probably don't want to, examine the history of individual articles before deciding on their individual trustworthiness, perhaps because it's too hard and too time-consuming.

So one hypothesis is that readers don't have the right tools to easily examine and trace back the editing history. That's why the idea of the WikiDashboard might be a really powerful way for fixing these problems. Social dashboards of these kinds are visualizations or graphical depictions of editing histories that will make it much easier for people to look at the history of an article and make up their own minds about its trustworthiness. The tool will enable us to do fundamental research on testing the hypothesis that transparency is what enables trust.

One thing we have done is to actually ran some experiments to understand if people are more willing to believe in information if you make the editing histories and activities more transparent. More on that on the next post.

Monday, September 10, 2007

WikiDashboard: Providing social transparency to Wikipedia


WikiDashboard Tool (alpha-release)

We are pleased to announce the release of our first research prototype of a social dynamic analysis tool for Wikipedia called WikiDashboard. This is a quick guide to our social dynamic analysis tool for Wikipedia

Motivation

The idea is that if we provide social transparency and enable attribution of work to individual workers in Wikipedia, then this will eventually result in increased credibility and trust in the page content, and therefore higher levels of trust in Wikipedia.

You might ask "Why would increasing social transparency result in higher quality articles and increase trust?"

Indeed, the quality of the articles in Wikipedia has been debated heavily in the press [here, here, here, here, and let's not forget the Nature magazine debacle].

Wikipedia itself keeps track of these studies and openly discusses them here, which is a form of social transparency itself. However, even Wales himself have has been quoted as saying that "while Wikipedia is useful for many things, he would like to make it known that he does not recommend it to college students for serious research." Indeed, the standard complaint I often hear about Wikipedia is that because of its editorial policy (anyone can edit anything), it is an unreliable source of information.

The opposite point of view, however, has not been debated or expressed nearly as much: Precisely because anyone can edit anything and that anyone can examine the edit history and see who has made them, it will (or has already) become a reliable source of information. I think Michael Scott, the character on the popular TV show "The Office", puts it succinctly: "Wikipedia is the best thing ever. Anyone in the world, can write anything they want about any subject. So you know you are getting the best possible information."

While tongue-in-cheek, it brings up a valid point. Because the information is out there for anyone to examine and to question, incorrect information can be fixed and two disputed points of view can be examined side-by-side. In fact, this is precisely the academic process for ascertaining the truth. Scholars publish papers so that theories can be put forth and debated, facts can be examined, and ideas challenged. Without publication and without social transparency of attribution of ideas and facts to individual researchers, there would be no scientific progress. Therefore, it seems somewhat ironic that the History Department at the Middlebury College have banned its students from citing Wikipedia sources .

Related Work

Indeed, just very recently WikiScanner has brought the issue and idea of social transparency to the forefront. It helps people find out the organizations where anonymous edits in Wikipedia are coming from. A week or two later, WikiRage helps identify the hottest trends in Wikipedia.

From academic works, we have seen interesting work from IBM called History Flow that visualizes the edits to article pages in Wikipedia, and the UCSC Wiki Trust Coloring Demo that demonstrated how trust could be visualized line-by-line. These are all examples of how being able to better understand editing history and editing patterns at a glance could dramatically help users uncover problems and the trustworthiness of contents on Wikipedia.

These tools and other discussions [NYTimes , blogs, and slashdot discussion] are noticing that accountability and transparency appears to be at the heart of the process that helps generate quality articles.

Guide to our tool

The tool can be used just as if you're on the Wikipedia site itself. All of the functions (such as the article search function, and the edit and history tabs) work just as before. The site provides the dashboard for each page in Wikipedia, while proxying the rest of the content from Wikipedia.

Note that we only currently have edit data up until 2007/07/16, so more recent edits are not included in the charts. We're working to fix this.

See our guide for help on understanding the visualizations in the WikiDashboard.

Some Interesting Examples

We will use the 2008 presidential election as an example. In the figure below, we see that the activities on this page has been heating up lately:
2008 US Presidential election
http://wikidashboard.parc.com/wiki/2008_presidential_election



Here are some notable Democractic Party candidates:
Hillary Clinton
http://wikidashboard.parc.com/wiki/Hillary_Rodham_Clinton



John Edwards
http://wikidashboard.parc.com/wiki/John_Edwards



Barack Obama
http://wikidashboard.parc.com/wiki/Barack_Obama



Here are some notable Republican candidates:

Rudy Giuliani
http://wikidashboard.parc.com/wiki/Rudy_Giuliani



John McCain
http://wikidashboard.parc.com/wiki/John_McCain



Ron Paul
http://wikidashboard.parc.com/wiki/Ron_Paul



Summary

We're curious of how the Web community will use this tool to surface social dynamics and editing patterns that might otherwise be difficult to find and analyze in Wikipedia. We are also interested in applying this tool to Enterprise Wikis. Please let us know by leaving a comment on this blog post on patterns you find or questions for us. Alternatively, (if you wish to contact us in private), email us at:
wikidashboard [at] parc [dot] com

Thanks,

Bongwon Suh
Ed H. Chi

Palo Alto Research Center

(joint work with our ex-colleagues Bryan Pendleton, Niki Kittur, now both at CMU)

A Quick Guide to WikiDashboard: Providing Social Transparency to Wikipedia


This post provides a quick guide to the WikiDashboard tool for Wikipedia:

Article WikiDashboard:





User WikiDashboard:




Detailed Edit Log:



Settings Panel:

Thursday, August 23, 2007

Wisdom of the Crowd, Collective Intelligence, and Collaborative Co-creation

“Wisdom of the crowd” is a great phrase, but we’ve had difficulty really understanding what it means. As mentioned by Ross Mayfield, one way to think about it is to break it up into two halves: “collective intelligence” and “collaborative intelligence”. Voting-style systems exhibit collective intelligence. Google’s page link algorithms involve pages voting for other pages; there are authors behind these pages, so implicitly there are people voting for people or people voting on content. Those are aspects of what you might call “collective intelligence.” It involves the averaging of opinions. I think a less buzzy term for it is "collective averaging".

At the other end, we have "collaborative intelligence", in which we see content production being produced in a kind of divide-and-conquer environments. Ross Mayfield said on his blog that the Wiki style of wisdom of the crowd was more “collaborative intelligence” than collective intelligence. For example, the group of people who are experts on World War II tanks will write that part of Wikipedia; the group of people who are experts on politics in Eastern Europe at the end of World War II will write those articles. So there is an implicit self-organization according to interest and intention. It’s not everybody voting on the same thing—it’s everybody collaborating on different areas to result in something, so that the sum of the parts is greater than the parts themselves. That seems to be at the spirit of this kind of collaborative intelligence.

I don’t really like the term “collaborative intelligence”—it sounds too buzzy—so we tend to call it “collaborative co-creation” instead. It is a very interesting production method. There is a lot of research now on, for example, the open source movement—how it’s a collaborative co-creation mechanism, how successful it is, what’s wrong with it, etc.

Wikipedia probably the most interesting collaborative co-creation system right now, and it is unique in the sense that it is all-encompassing; its net has been cast very wide and it has been able to succeed because of that. There is a little bit of a success-breeds-success phenomenon going on there with the feedback cycle.

This feedback cycle is the part we’re really interested in understanding, because coordination is at the heart of collaborative creation. We want to understand how people are coordinating with one another through either self-organizing mechanisms or through explicit organizing mechanisms; we want to understand the principles by which those things happen in these environments but not in other environments.

Tuesday, May 22, 2007

Controversy Visualization

Alas, not our visualization of revert relationships, but someone got slashdotted for doing a visualization of power struggle in Wikipedia:
Slashdot article.

"todd450 pointed us to a nifty visualization of Wikipedia and controversial articles in it. The image started with a network of 650,000 articles color coded to indicate activity. The original image is apparently 5' square, but the sample image they have is still pretty neat."

The original blog post was here.