A comparison of three methods to determine the subject matter in textual data-Reference-Cited by-同舟云学术

A comparison of three methods to determine the subject matter in textual data

Published:2023-06-02 Issue: Volume:8 Page:
ISSN:2504-0537
Container-title:Frontiers in Research Metrics and Analytics
language:
Short-container-title:Front. Res. Metr. Anal.

Author:

Barnett George A.,Calabrese Christopher,Ruiz Jeanette B.

Abstract

This study compares three different methods commonly employed for the determination and interpretation of the subject matter of large corpuses of textual data. The methods reviewed are: (1) topic modeling, (2) community or group detection, and (3) cluster analysis of semantic networks. Two different datasets related to health topics were gathered from Twitter posts to compare the methods. The first dataset includes 16,138 original tweets concerning HIV pre-exposure prophylaxis (PrEP) from April 3, 2019 to April 3, 2020. The second dataset is comprised of 12,613 tweets about childhood vaccination from July 1, 2018 to October 15, 2018. Our findings suggest that the separate “topics” suggested by semantic networks (community detection) and/or cluster analysis (Ward's method) are more clearly identified than the topic modeling results. Topic modeling produced more subjects, but these tended to overlap. This study offers a better understanding of how results may vary based on method to determine subject matter chosen.

Funder

University of California, Davis

Publisher

Frontiers Media SA

Subject

General Medicine

Reference52 articles.

1. Mining Text Data

2. 5. Issues in intercultural communication: a semantic network analysis;Barnett;Interc. Commun.,2017

3. An examination of the relationship between international telecommunication networks, terrorism and global news coverage;Barnett;Social Netw. Anal. Mining,2013

4. Gephi: an open source software for exploring and manipulating networks;Bastian;Proc. Conf. Web Soc. Media,2009

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Cascaded Semantic Fractionation for identifying a domain in social media;Frontiers in Research Metrics and Analytics;2024-03-01

2. Who Sets the Agenda for Climate Change in China? A Longitudinal Analysis of Primary Actors that Drive Online Discussions on Social Media;Environmental Communication;2024-02-07