Showing posts with label social tagging. Show all posts
Showing posts with label social tagging. Show all posts

Friday, April 24, 2009

Social bookmarks as traces left behind as navigational signposts

Social tagging arose out of the need to organize found content that is worth revisiting. It is natural therefore to think of social tagging and bookmarking as navigational signposts for interesting content. The collective behavior of users who tagged contents seems to offer a good basis for exploratory search interfaces, even for users who are not using social bookmarking sites.

In Boston at the CHI2009 conference, we presented a paper that showed how our tag-based search interface called MrTaggy can be used as learning tools for people to find content relating to a particular topic. We have already announced its availability on this blog, and also touched upon the way in which it is implemented. Here we will briefly blog about an evaluation study we did on this system in order to understand its learning effects.

Short Story:

The tag-based search system allows users to utilize relevance feedback on tags to indicate their interest in various topics, enabling rapid exploration of the topic space. It turns out that the experiment shows that the system seems to provide a kind of scaffold for users to learn new topics.

Long Story:

We recently completed a 30-subject study of MrTaggy [see reference below for full detail]. We compared the full exploratory MrTaggy interface to a baseline version of MrTaggy that only supported traditional query-based search.



We tested participants’ performance in three different topic domains.





The results show:

(1) Subjects using the MrTaggy full exploratory interface took advantage of the additional features provided by relevance feedback, without giving up their usual manual query typing behavior.



(2) For learning outcomes, subjects using the full exploratory system generally wrote summaries of higher quality compared to baseline system users.



(3) To also gauge learning outcomes, we asked subjects to generate keywords and input as many keywords as possible that were relevant to the topic domain in a certain time limit. Subjects using the exploratory system were generally able to generate more reasonable keywords than the baseline system users.

(4) Finally, other convergent measures show that they also spent more time on the learning tasks, and had a higher cognitive load. Taken together with the higher learning measure outcomes, the users appear to be more engaged in exploration than the participants using the baseline system.

Our findings regarding the use of our exploratory tag search system are promising. The empirical results show that subjects can effectively use data generated by social tagging as “navigational advice” in the learning domain.

The experimental results suggest that users’ explorations in unfamiliar topic areas are supported by the domain keyword recommendations presented in the related tags list and the opportunity for relevance feedback.

Since social search engines that depend on social cues rely on data quality and increasing coverage of the explorable web space, we expect that the constantly increasing popularity of social bookmarking services will improve social search browsers like MrTaggy. The results of this project point to the promise of social search to fulfill a need in providing navigational signposts to the best contents.


Reference:

Kammerer, Y., Nairn, R., Pirolli, P., and Chi, E. H. 2009. Signpost from the masses: learning effects in an exploratory social tag search browser. In Proceedings of the 27th international Conference on Human Factors in Computing Systems (Boston, MA, USA, April 04 - 09, 2009). CHI '09. ACM, New York, NY, 625-634.

ACM Link

Talk Slides

Monday, March 23, 2009

How MrTaggy is implemented...

A short time ago, we announced the MrTaggy browsing and searching engine for social bookmarks here. One of the neat features of this system is its relevance feedback mechanism which enables users to click on keywords to navigate toward the information that they are interested in.

The overall system uses a sophisticated MapReduce computation in the backend, and the implementation is non-trivial. Here is how it works. The diagram below was recently published in an IEEE Computer Magazine article, and it roughly describes how the data flows thru the whole system. (Click on it to enlarge it.)



First, a crawling module goes out to the web and crawls social tagging sites, looking for tuples of the form . Tuples are stored in a MySQL database. In our current system, we have roughly 150 million tuples.

A MapReduce system based on Bayesian inference and spreading activation then computes the probability of each URL or tag being relevant given a particular combination of other tags and URLs. Here we first construct a bigraph between URLs and tags based on the tuples and then precompute spreading activation patterns across the graph.

To do this backend computation in massively parallel way, we used the MapReduce framework provided by Hadoop (hadoop.apache org). The results of this computation are stored in a Lucene index so that we can make the retrieval of spreading activation patterns as fast as possible.

Finally, a web server serves up the search results through an interactive frontend. The frontend responds to user interaction with relevance feedback arrows by communicating with the web server using AJAX techniques and animating the interface to an updated state.

Reference:
Ed H. Chi, "Information Seeking Can Be Social," IEEE Computer, vol. 42, no. 3, pp. 42-46, March, 2009.

Tuesday, February 24, 2009

Announcing MrTaggy.com: a Tag-based Exploration and Search System


I'm pleased to announce MrTaggy.com, a tag-based exploration and search system for bookmarked content on the Web. The tagline for the project is "An interactive guide to what's useful on the Web", since all of the content has been socially vetted (i.e. someone found it useful enough to bookmark it.)

MrTaggy is an experiment in web search and exploration built on top of a PARC algorithm called TagSearch. Think of MrTaggy as a cross between a search engine and a recommendation engine: it’s a web browsing guide constructed from social tagging data. We have collected about 150 million bookmarks from around the Web.

Unlike most search engines, MrTaggy doesn’t index the text on a web page. Instead, it leverages the knowledge contained in the tags that people add to web pages when using social bookmarking services. Tags describe both the content and context of a web page, and we use that information to deliver relevant contents.

The problem with using social tags is that they contain a lot of noise, because people often use different words to mean the same thing or the same words to mean different things. The TagSearch algorithm is part of our ongoing research to reduce the noise while amplifying the information signal from social tags.

We also designed a novel search UI to explore the tag space. The Related Tags sidebar outlines the content landscape to help you understand the space. The relevance feedback capabilities enable you to tell the system both positive and negative cues about directions where you want to go. Try clicking on the Thumbs Up and Down to give feedback to MrTaggy about the tags or results that you liked, and see how your rating changes the result set on-the-fly. At the top of the result set, we have also provided top search results from Yahoo's search engine when we think the results there might help you.

Enterprise Use

In addition to exploring TagSearch in the consumer space, we have also explored the use of TagSearch in the enterprise social tagging and intranet search systems. Surprisingly, the algorithm worked well even with a small amount of data (<50,000 bookmarks). For enterprise licensing of the underlying technology and API, contact Lawrence Lee, Director of Business Development, at lawrence.lee [at] parc [dot] com.

We would appreciate your feedback (comment on the blog here), or send them to mrtaggy [at] parc [dot] com, or submit at mrtaggy.uservoice.com.

Click here to try MrTaggy.com

Monday, November 12, 2007

How social tagging appears to affect human memory?

Three weeks ago, I was at the ASIST 2007 Annual conference in Milwaukee, which had a special theme on Social Computing and Information Science. During one of the panels on Social Tagging, a question was raised on how tagging really affects memory and retrieval. I mentioned that the ASC group here at PARC has been doing some experiments on this, and briefly talked about the results, and many attendees at the conference (over 10 people) had asked for the pre-print, so here I'm blogging about it.

Raluca Budiu, who is a post-doc working in our group, has conducted some very interesting research with us on how tagging appears to affect human information processing. She studied two techniques for producing tags: (1) the traditional type-to-tag interface of typing keywords into a free-form textbox after reading a passage or article; (2) a PARC-developed click2tag interface that allows users to click on keywords in the paragraph to tag the content.

The experiment consisted of 20 subjects and 24 passages in a within-subject design. Participants had to first study passages and tag them, and then they performed memory tests on what they had actually read and tagged. The memory tasks were that, after tagging the content, they have to either (a) freely recall and type as many facts from the passages as possible; or (b) answer 6 true/false sentences in a recognition task.

As reported in the paper, the results suggest that:

  • In the type-to-tag condition, users appears to elaborate what they have just read, and re-encoded the knowledge with keywords that might be helpful for later use. This appears to help the free-recall task (a) above. In other words, users seem to end up with a top-down process and induces them to schematize what they have learned.


  • While in the click2tag condition, users appears to re-read the passages to pick out keywords from the sentences, and this appears to help them in their recognition tasks (b) above. In other words, users seem to use a bottom-up process that simply picked out the most important keywords from the passage.


Click here to download the technical report and pre-print (the highlights in the paper are mine).

Monday, October 29, 2007

Differences between Social Tagging and Collaborative Tagging

I'm here at the InfoVis conference in Sacramento and a conversation with Marti Hearst over at UCBerkeley just reminded me why I have been bothered by the 'confusion' between the phrases "social tagging" and "collaborative tagging" for quite some time. In fact, Wikipedia has a redirection of "Social Tagging" to "Collaborative Tagging" (see http://en.wikipedia.org/w/index.php?title=Social_tagging&redirect=no). This, I would argue, is wrong. Why?

'Collaborate', according to the American Heritage Dictionary, is "to work together, especially in a joint intellectual effort." The problem is that tagging features in many of the popular Web2.0 tools such as Flickr and YouTube are not really 'collaborative', since users aren't really working together per se. In YouTube, for example, only the uploader of the original video clip can specify and edit the tags for an video. Most of the time, in Flickr, one only tag their own photos. However, Flickr is somewhat more collaborative than YouTube because the default setting for any account is to allow contacts such as friends and families to also tag the photos.

Both of these two systems don't seem that 'collaborative', because, to me, collaboration implies shared artifact, shared workspace, and shared work. On the other hand, 'social' is "living or disposed to live in companionship with others or in a community, rather than in isolation". In other words, simply existing and having some relation to others in a community. So for example, I would argue that in YouTube, we have social tagging but not collaborative tagging, because while users tag their uploaded videos in the context of a online social community, and they do not collaborate to converge on a set of tags appropriate for that video.

The use of the term 'collaborative' in past Computer-Supported Cooperative Work (CSCW) field has especially come to imply
a shared workspace. With shared workspaces, often there are some elements of coordination and conflicts involved as well (and hopefully conflict resolution as well). So in contrast to YouTube, the most 'collaborative' tagging system I know is the category tagging system in Wikipedia. Anyone can edit the category tags for an article. They can remove, add, discuss, and revert the use of any tag. In this case, the category tags are shared artifacts that anyone can edit inside a shared workspace. The work of tagging all 2 Million+ articles in Wikipedia is shared work among the community.

It's perhaps interesting to note that somewhere in between YouTube and Wikipedia tagging is perhaps the bookmarking system del.icio.us. In del.icio.us, there is a shared artifact (the tagged sites or URLs), and there is shared work of tagging all of the websites and pages out there on the Web. However, there is less of a notion of a shared workspace. My tags for an URL could be and probably is different from someone else's tags for the same URL. I also have the capability of searching within just my own del.icio.us space. So from least collaborative to the most collaborative, we have YouTube, then del.icio.us, and then finally the category tagging system in Wikipedia.

A simple way to explain this is that one must be social in order to collaborate, but one need not be collaborative to be social. So in summary, I would argue that social tagging is a superset of collaborative tagging. But a social tagging system may not necessarily be a collaborative tagging system. We should change the definitions in Wikipedia to distinguish between these two types of systems.