Showing posts with label HCI. Show all posts
Showing posts with label HCI. Show all posts

Monday, June 14, 2010

Model-Driven Research in Social Computing

I'm in Toronto attending the Hypertext 2010 conference, where I gave the keynote talk at the First Workshop on Modeling Social Media yesterday. I want to document a little bit of the points I made in the talk here.

The reason we seek to construct and derive models is to predict and explain what might be happening in social computing systems. For social media, we seek to understand how these systems evolve over time. Constructing these models should also enable us to generate new ideas and systems.

As an example, many have proposed a theory of influentials that identifying a small group of individuals who are connected to the larger social network just in the right way, we can infect or reach the rest of the people in the network. This idea is probably most well-known in the press by the popular book Tipping Point by Gladwell. This model of how information diffuse in social networks is very attractive, not just due to its simplicity, but also the potential of applying this idea in areas such as marketing.

Models such as this are meant to be challenged and debated. They are always strawman proposals. Duncan Watts' simulation on networks have shown that the validity of this theory is somewhat suspect. Indeed, recently, Eric Sun and Cameron Marlow's work, published in ICWSM2009, showed that this theory of influentials might be wrong. They suggest that "diffusion chains are typically started by a substantial number of users. Large clusters emerge when hundreds or even thousands of short diffusion chains merge together."

Most, if not all, models are wrong. Some models are just more wrong than others. But models still serve important roles. They might be divided into several categories:

  1. Descriptive Models describe what is going on within the data. This might help us spot trends, such as the growth of number of contributors, or trending topics in a community.
  2. Explanatory Models help us explain what might be the mechanisms underlying processes in the system. For example, we might be able to explain why certain groups of people contribute more content than another group.
  3. Predictive Models help us engineer systems by predicting what users and groups might want, or how they might act in systems. Here we might build probabilistic models of whether a user will use a particular tag on a particular item in a social tagging system.
  4. Prescriptive Models are set of design rules or a process that helps practitioners generate useful or practical systems. For example, Yahoo's Social Design Patterns Library on Reputation is a very good example of a prescriptive model.
  5. "Generative Models" actually have two meanings depending on who you're talking to. In statistical circles, "generative models" are models that help generate data that look like real user data and are often probabilistic models. Information Theory is a good example of this approach, in fact. Generative Models could also mean that they are models that help us generate ideas, novel techniques and systems. My work with Brynn Evans on building a social search model is an example of this approach.
In the talk, I illustrated how we have modeled the dynamics in the popular social bookmarking system, Delicious, using Information Theory. I also showed how using equations from Evolutionary Dynamics we were better able to explain what might be happening to Wikipedia’s contribution patterns. Talk Title: Model-driven Research for Augmenting Social Cognition

Friday, February 13, 2009

WikiDashboard and the Living Laboratory


Our work on WikiDashboard was slashdotted last weekend. It caused our server to fail and crash repeatedly, and we tried our best to keep it running. We received thousands of hits, and got many comments. Interestingly, this occurred because of an MIT TechReview article on the system, which was in turn caused by the reporter coming to my talk at MIT last Tuesday (video here).

The whole experience is a very good example of the concept of the Living Laboratory. We were interested in engaging the real world in doing social computing research, and found Wikipedia to be a great way to get into the research, while benefiting the discourse around how knowledge bases should be built.

We had argued that Human-Computer Interaction (HCI) research have long moved beyond the evaluation setting of a single user sitting in front of a single desktop computer, yet many of our fundamentally held viewpoints about evaluation continues to be ruled by outdated biases derived from this legacy. We believe that we need to engage with real users in 'Living Laboratories', in which researchers either adopt or create functioning systems that are used in real settings. These new experimental platforms will greatly enable researchers to conduct evaluations that span many users, places, time, location, and social factors in ways that are unimaginable before.

Outdated Evaluative Assumptions

Indeed, the world has changed. Trends in social computing as well as ubiquitous computing had pushed us to consider research methodologies that are very different from the past. In many cases, we can no longer assume:

Only a single display: Users will pay attention to only one display and one computer. Much of fundamental HCI research methodology assumes the singular occupation of the user is the display in front of them. Of course, this is no longer true. Not only do many users already use multiple displays, they also use tiny displays on cell phones and iPods and peripheral displays. Matthews et al. studied the use of peripheral displays, focusing particularly on glance-ability, for example. Traditional HCI and psychological experiments typically force users to attend to only one display at a time, often neglecting the purpose of peripheral display designs.

Only knowledge work: Users are performing the task as part of some knowledge work. The problem with this assumption is that non-information oriented work, such as entertainment applications, social networking systems, are often done without explicit goals in mind. With the rise of Web2.0 applications and systems, users are often on social systems to kill time, learn the current status of friends, and to serendipitously discover what might capture their interests.

Isolated worker: Users performing some task by themselves. Much of knowledge work turn out to be quite collaborative, perhaps more so than first imagined. Traditional view of HCI assumed the construction of a single report by a single individual that is needed by a hierarchically organized firm. Generally speaking, we have come to view such assumption with contempt. Information work, especially work done by highly paid analysts, is highly collaborative. Only the highly automated tasks that are routine and mundane are done in relative isolation. Information workers excel at exception handling, which often require the collaboration of many departments in different parts of the organizational chart.

Stationary worker: User location placement is stationary, and the computing device is stationary. A mega-trend in information work is the speed and mobility in which work is done. Workers are geographically dispersed, making collaboration across geographical boundaries and time-zone critical. As part of this trend, work is often done on the move, in the air while disconnected. Moreover, situation awareness is often accomplished via email clients such as Blackberries and iPhones. Many estimates now suggest that already more people access the internet on their mobile phone than on desktop computers. This certainly has been the trend in Japan, a bellwether of mobile information needs.

Task duration is short: Users are engaged with applications in time scales measures in seconds and minutes. While information work can be divided and be composed of many slices of smaller chunks of subgoals that can be analyzed separately, we now realize that many user needs and work goals stretch over for long period of time. User interests in topics as diverse as from news on the latest technological gadgets to snow reports for snowboarding need to be supported over periods of days, weeks, months and even years. User engagement with web applications are often measured in much longer periods of time as compared to more traditional psychological experiments that geared toward understanding of hand-eye coordination in single desktop application performance. For example, Rowan and Mynatt studied peripheral family portraits in the digital home over a year-long period and discovered that behavior changed with the seasons (Rowan and Mynatt, 2005).

The above discussion point to how, as a field, HCI researchers have slowly broken out of the mold in which we were constrained. Increasingly, evaluations are often done in situations in which there are just too many uncontrolled conditions and variables. Artificially created environments such as in-lab studies are only capable of telling us behaviors in constrained situations. In order to understand how users behave in varied time and place, contexts and other situations, we need to systematically re-evaluate our research methodologies.

Time has come to do a great more deal of experimentation in the real world, using real and living laboratories.

Wednesday, November 5, 2008

'Living Laboratories': Rethinking Ecological Designs and Experimentation in Human-Computer Interaction

During the formation of the HCI field, the need to establish HCI as a science had pushed us to adopt methods from psychology, both because it was convenient as well as the methods fit the needs. Real HCI problems have long moved beyond the evaluation setting of a single user sitting in front of a single desktop computer, yet many of our fundamentally held viewpoints about evaluation continues to be ruled by outdated biases derived from this legacy.

Trends in social computing as well as ubiquitous computing had pushed us to consider research methodologies that are very different from the past. In many cases, we can no longer assume only a single display, only knowledge work, isolated worker, location stationary with short task durations. HCI researchers have slowly broken out of the mold in which we were constrained. Increasingly, evaluations are often done in situations in which there are just too many uncontrolled conditions and variables. Artificially created environments such as in-lab studies are only capable of telling us behaviors in constrained situations. In order to understand how users behave in varied time and place, contexts and other situations, we need to systematically re-evaluate our research methodologies.

The Augmented Social Cognition group have been a proponent of the idea of 'Living Labratory' within PARC. The idea (born out of a series of conversation between myself, Peter Pirolli, Stuart Card, and Mark Stefik) is that in order to bridge the gulf between academic models of science and practical research, we need to conduct research within laboratories that are situated in the real world. Many of these living laboratories are real platforms and services that researchers would build and maintain, and just like Google Labs or beta software, would remain somewhat unreliable and experimental, but yet useful and real. The idea is to engage real users in ecological valid situations, while gathering data and building models of social behavior.

Looking at two different dimensions in which HCI researchers could conduct evaluations, one dimension is whether the system is under the control of the researcher or not. Typically, computing scientists build systems and want them evaluated for effectiveness. The other dimension is whether the study is conducted in the laboratory or in the wild. These two dimensions interact to form four different ways of conducting evaluations:

  1. Building a system, and studying it in the laboratory. This is the most traditional approach in HCI research and the one that is typically favored by CHI conference paper reviewers. The problem with this approach is that it is (1) extremely time-consuming, and (2) experiments are not always ecologically valid. As mentioned before, it is extremely difficult, if not impossible, to design experiments for many social and mobile applications that are ecologically valid in the laboratory.

  2. Not building a system (but adopt one), and still study it in the laboratory. For example, this is possible by taking existing systems, such as Microsoft Word and iWorks Pages and comparing the features of these two systems.

  3. Adopting an existing system, and studying it in the wild. The advantage here is to study real applications that are being used in ecologically valid situations. The disadvantage is that findings are often not comparable, since factors are harder to isolate. On the other hand, the advantages are that real findings can be immediately applied to the live system. Impact of the research is real, since adoption issues are already removed. We have studied Wikipedia usage in detail using this method by releasing WikiDashboard.

  4. Building a system, releasing it, and studying it in the wild. A well-publicized use of this approach is Google's A/B testing approach . Apparently, according to Marissa Mayer at Google, A/B testing allowed them to finely tune the Search Engine Result Pages (SERPs). For example, how many search results should the page contain was studied carefully by varying the number between a great number of users. Because the subject pool is large, Google can say with some certainty which design is better on their running system. A major disadvantage of this approach is the effort and resource requirement it takes to study such systems. However, for economically interesting applications such as Web search engines, the tight integration between system and usage actually shorten the time to innovate between product versions.

Of these variations, (3) and (4) are what we consider to be 'Living Laboratory' studies. This was the reason why we released WikiDashboard into the wild. We will be releasing a new social search engine called MrTaggy in the near future. The idea is the same: to test some social search systems in the wild to see how they perform with real users.

Friday, September 26, 2008

The Social Web: an academic research fad?

One enduring core value in Human-Computer Interaction (HCI) research has been the development of technologies that augment human intelligence. This mission originates with V. Bush, Licklider, and Engelbart, who inspired many researchers such as Alan Kay at PARC in the development of the personal computer and the graphical user interface.
A natural extension of this idea in the Social Web and Web2.0 world is the development of technologies that augment social intelligence. In this spirit, the meaning of “Augmented Social Cognition” builds on Engelbart’s vision.

Beyond HCI researchers, scientists from diverse fields such as Computer-Supported Cooperative Work (CSCW), WWW research, Hypertext, Digital Libraries are feeling the impact of such systems and are publishing research papers that characterize, model, prototype, and evaluate various systems. Studies from behavioral microeconomics, organizational economics, sociology, ethnography, social network analysis, information flow analysis, political science, and conflict resolution are potentially relevant to Social Web researchers. Researchers are seeing a surge of new research on Web2.0 technologies distributed in a wide variety of disciplines and associated conferences. In this past year, I have attended conferences in these different fields to gain a sense of the horizontal effect that the Social Web is having on academic research.



• At the light-end of collaboration spectrum, we have researchers trying to understand the micro-economics of voting systems, of individual and social information foraging behaviors, processes that govern information cascade, and wisdom-of-the-crowd effects. HCI researchers have productively studied information foraging and behavioral models in the past, and are trying to apply them in the new social context on the Web. Economists are trying to understand peer production systems, new business models, and consumption and production markets based on intrinsic motivations.

Our own research on using information theory to study global tagging trends is an example here.

• At the middle of the collaboration spectrum, researchers are building algorithms that mine new socially constructed knowledge structures and social networks. Here physicists and social scientists are using network theories and algorithms to model, mine, and understand these processes. Algorithms for identifying expertise and information brokers are being devised and tested by information scientists.

Here we have been building a system called MrTaggy that uses an algorithm called TagSearch to offer a kind of social search system based on social tagging data. I'll blog with a screencast demo soon.

• At the heavy-end of the collaboration spectrum, the understanding of coordination and conflict costs are especially important for collaborative co-creation systems such as Wikipedia. Researchers had studied characteristics that enable groups of people to solve problems together or collaborate on scientific endeavors. Discoveries such as the identification of “invisible colleges” by Sandstrom have shown that implicit coordination can be studied and characterized.

Our research into coordination effects in Wikipedia is an example of research here.

The horizontal effect of the Social Web is changing academic research in various and important ways. The Social Web is providing a rich playground in which to understand how we can augment web users’ capacity and speed to acquire, produce, communicate, and use knowledge; and to advance collective and individual intelligence in socially mediated information environments. Augmented Social Cognition research, as explained here, emerged from a background of activities aimed at understanding and developing technologies that enhance the intelligence of users, individually and in social collectives, through socially mediated information production and use.

In part this is a natural evolution from HCI research around improving information seeking and sense making on the Web, but in part this is also a natural expansion in the scientific efforts to understand how to augment the human intellect.

The Social Web isn’t just a fad, but a fundamental transformation of the Web into a true collaborative and social platform. The research opportunity is to fully understand how to enhance the ability of a group of people to remember, think, and reason.

Monday, March 10, 2008

How to reduce the cost of doing user studies with Crowdsourcing


One problem we have been facing as HCI researchers is how to get user data such as their opinions or relevance judgements quickly and cheaply. I think we may have a good way of doing this with Amazon Mechanical Turk with crowdsourcing that we're about to report in the CHI2008 conference.


User studies are important for many aspects of the design process and involve techniques ranging from informal surveys to rigorous laboratory studies. However, the costs involved in engaging users often requires practitioners to trade off between sample size, time requirements, and monetary costs. In particular, collecting input from only a small set of participants is problematic in many design situations. In usability testing, many issues and errors (even large ones) are not easily caught with a small number of participants, as we have learned from people like Jared Spool at UIE.


Recently, we investigate the utility of a micro-task market for collecting user measurements. Micro-task markets, such as Amazon’s Mechanical Turk, offer a potential paradigm for engaging a large number of users for low time and monetary costs. Although micro-task markets have great potential for rapidly collecting user measurements at low costs, we found that special care is needed in formulating tasks in order to harness the capabilities of the approach. The special care turns out to mirror the game theoretic issues somewhat reminiscent of Luis von Ahn's work on ESP Games.

We conducted two experiments to test the utility of Mechanical Turk as a user study platform. Here is a quick summary:

In both experiments, we used tasks that collected quantitative user ratings as well as qualitative feedback regarding the quality of Wikipedia articles. We had Mechanical Turk users rate a set of 14 Wikipedia articles, and then compared their ratings to an expert group of Wikipedia administrators from a previous experiment. We had users rate articles on a 7-point Likert-scale according to a set of factors including how well written, factually accurate, neutral, well structured, and overall high quality the article was.

In one experiment, users were required to fill out a free-form text box describing what improvements they thought the article needed. 58 users provided 210 ratings for 14 articles (i.e., 15 ratings per article). User response was extremely fast, with 93 of the ratings received in the first 24 hours after the task was posted, and the remaining 117 received in the next 24 hours. However, in this first experiment, only about 41.4% of the responses appeared to be honest effort in rating the Wikipedia articles. An examination of the time taken to complete each rating also suggested gaming, with 64 ratings completed in less than 1 minute (less time than likely needed for reading the article, let along rating it). 123 (58.6%) ratings were flagged as potentially invalid based either on their comments or duration. However, many of the invalid responses were due to a small minority of users. So this appears to demonstrate the susceptibility of Mechanical Turk to malicious user behavior.

So in a second experiment, we tried a different design. The new design was intended to make creating believable invalid responses as effortful as completing the task in good faith. The task was also designed such that completing the known and verifiable portions would likely give the user sufficient familiarity with the content to accurately complete the subjective portion (the quality rating). For example, these questions required users to input how many references, images, and sections the article had. In addition, users were required to provide 4-6 keywords that would give someone a good summary of the contents of the article, which we can verify quickly later.

Instead of about 60% bad responses, there were dramatically fewer responses that appeared invalid. Only 7 responses had meaningless, incorrect, or copy-and-paste summaries, versus 102 in Experiment 1. We also had a positive correlation with the quality rating given to us by Wikipedia administrators, and was statistically significant (r=0.66, p=0.01).

These results suggest that micro-task markets may be useful as a crowdsourcing tool for other types of user study tasks that combine objective and subjective information gathering, but there are design considerations. While hundreds of users can be recruited for highly interactive tasks for marginal costs within a timeframe of days or even minutes, however, special care must be taken in the design of the task, especially for user measurements that are subjective or qualitative.

Reference is:
Kittur, A., Chi, E., Suh, B. Crowdsourcing User Studies With Mechanical Turk. In Proceedings of the ACM Conference on Human-factors in Computing Systems (CHI2008). ACM Press, 2008. Florence, Italy.

Paper is here.