IN014-0008
Analyzing Newspaper Articles: Opportunities, Strategies, and Risks
Analyzing Newspaper Articles: Opportunities, Strategies, and Risks
Wednesday, 9 December 2020
Poster
Abstract:
Despite the difficulty in processing unstructured data sources, the potential that they hold makes the task perennially appealing. One of the most ubiquitous sources of unstructured data is textual data, which come in many forms, including technical or business reports, social media, or newspaper articles. While these sources contain important details, the information is diffuse and low value. Using a corpus of disaster-related newspaper articles, we demonstrate different techniques for introducing structure to these unstructured data sources. We present findings and insights generated from key term identification, topic modeling, bigram frequency counts, and sentiment analysis. Our results indicate that analyzing bigrams including a term believed to be relevant to the task at hand (e.g. "water" or "fire") can yield insights into how these terms are discussed in the textual data at hand. Such insights can in turn inform subsequent analysis and interpretation of results. This text analysis case study also highlights drawbacks associated with using some popular methods for analyzing unstructured textual data (e.g. explainability of topic modeling) and provides a path forward on how to combat these issues.