New Issue: Orbital Catastrophe Ahead? Read Now

Parsing the Twitterverse: New Algorithms Analyze Tweets

Smarter language processors are helping experts analyze millions of short-text messages from across the Internet

Join Our Community of Science Lovers!

Researchers have been trolling Twitter for insights into the human condition since shortly after the site launched in 2006. In aggregate, the service provides a vast database of what people are doing, thinking and feeling. But the research tools at scientists’ disposal are highly imperfect. Keyword searches, for example, return many hits but offer a poor sense of overall trends.

When computer scientist James H. Martin of the University of Colorado at Boulder searched for tweets about the 2010 earthquake in Haiti, he found 14 million. “You can’t hire grad students to read them all,” he says. Researchers need a more automated approach.

One promising method is to develop programs that label words in tweets with parts of speech—such as subject, verb and object—and then use those tags to determine what each tweet is about. This method, called natural-language processing, is not a new idea, but applying it to short social text is new and growing. “That is just a huge area right now,” Martin says.


On supporting science journalism

If you're enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.


Scientists at the Xerox-owned Palo Alto Research Center recently developed one such program. It relies on text processors, called parsers, which are typically tested on news articles. Parsers can distinguish between words and punctuation, label parts of speech and analyze a sentence’s grammatical structure. But “they don’t do as well on Twitter,” says Kyle Dent, one of the Palo Alto researchers. He and his co-author wrote hundreds of rules to account for hash tags, repeated letters (as in “pleaaaaaase”) and other linguistic features perhaps not common in the Wall Street Journal. They will present their work on August 8 at an Association for the Advancement of Artificial Intelligence conference in San Francisco.

Dent and his colleagues also tried to use their program to distinguish between rhetorical questions and those that require a response. Businesses could use such a program to find what people are asking about their products. In a recent trial, their program classified 68 percent of 2,304 tweets correctly. “For a brand-new field, that sounds like a decent first attempt,” says Jeffrey ­Ellen of the Space and Naval Warfare Systems Command, which provides intelligence technology to the U.S. Navy.

Although Twitter-trawling technology is not yet ready to deploy, as a field, “it’s getting there pretty quickly,” Martin says. Once it matures, researchers should have access to an unprecedented trove of data about human behavior. For the first time in history, “watercooler talk” is recorded and publicly available, Ellen says. “A hundred years ago we just didn’t know what everybody was thinking.”

Scientific American Magazine Vol 305 Issue 2This article was published with the title “Parsing the Twitterverse: New Algorithms Analyze Tweets” in Scientific American Magazine Vol. 305 No. 2 ()
doi:10.1038/scientificamerican082011-2O9yHmrxt9lY2uyQSN0RNK

Subscribe to Support Independent Journalism

Great science journalism requires human expertise, time, effort and creativity. And it costs money. That’s why I and the journalists here at Scientific American hope you’ll join our community.

When you subscribe, you are supporting staff and freelance journalists who are passionate about telling science stories that are true, important and compelling. Our editors and reporters are often experts in their fields, which means they understand the nuances of big discoveries and can untangle the breakthroughs from the hype. With a subscription, you are also supporting rigorous fact-checking to ensure the words we publish are precise and accurate. And you’re supporting original illustrations, graphics and photos that bring you closer to an advanced laboratory, an ice sheet in Antarctica or a space mission in orbit. You’re helping us craft other types of high-quality journalism as well: Our newsletters are carefully written, edited and curated by staffers you have or will come to know and love. Our Science Quickly podcast is based on original reporting, collaboration with editors and scientists and exacting production.

Subscriptions keep this engine running so we can continue to deliver thoughtful, rigorous and independent science journalism to you. In an era of viral misinformation, this work is crucial. If you value what we do, I hope you’ll consider joining us as a subscriber

Thank you,

Jeanna Bryner, Editor in Chief, Scientific American

Subscribe