A comparison of three methods to determine the subject matter in textual data

George A Barnett; Christopher Calabrese; Jeanette B Ruiz

doi:10.3389/frma.2023.1104691

A comparison of three methods to determine the subject matter in textual data

Front Res Metr Anal. 2023 Jun 2:8:1104691. doi: 10.3389/frma.2023.1104691. eCollection 2023.

Authors

George A Barnett¹, Christopher Calabrese², Jeanette B Ruiz¹

Affiliations

¹ Department of Communication, University of California, Davis, Davis, CA, United States.
² Department of Communication, Clemson University, Clemson, SC, United States.

Abstract

This study compares three different methods commonly employed for the determination and interpretation of the subject matter of large corpuses of textual data. The methods reviewed are: (1) topic modeling, (2) community or group detection, and (3) cluster analysis of semantic networks. Two different datasets related to health topics were gathered from Twitter posts to compare the methods. The first dataset includes 16,138 original tweets concerning HIV pre-exposure prophylaxis (PrEP) from April 3, 2019 to April 3, 2020. The second dataset is comprised of 12,613 tweets about childhood vaccination from July 1, 2018 to October 15, 2018. Our findings suggest that the separate "topics" suggested by semantic networks (community detection) and/or cluster analysis (Ward's method) are more clearly identified than the topic modeling results. Topic modeling produced more subjects, but these tended to overlap. This study offers a better understanding of how results may vary based on method to determine subject matter chosen.

Keywords: cluster analysis; community detection; social media; text analysis; topic modeling.

Grants and funding

External funding was not received for the research. Our institution provides $1000 toward open-access publication fees.