Using Diachronic Static Word Embeddings to Detect and Characterize Term Toxification on Reddit
Files
Publication or External Link
External Link to Data Files
Date
Authors
Advisor
Citation
DRUM DOI
Abstract
Toxic discourse on the internet has a powerful influence on its users, and recognizing increasingly negative or derogatory usage of specific terms–such as ‘Karen’ or ‘woke’–can help signal potential “toxification” of an online space. Our research demonstrates that a novel approach building on prior work on diachronic static word embedding analysis can assist in tracking and understanding term toxification on Reddit, a popular online social media platform, by observing how the usages of words with multiple meanings (including a pejorative/negative one) change over time. This study can be expanded by using contextual word embeddings, expanding comparisons to other social media platforms, and examining word toxification in different languages.