Black user responses to Black style performances in generative AI outputs dataset

Loading...
Thumbnail Image

Related Publication Link

External Link to Data Files

Date

Advisor

Related Publication Citation

Abstract

This dataset contains video transcripts, metadata, images and user comments from a sample of TikTok videos. Dataset consists of .png and .jpeg image files and .csv files.

This dataset accompanies a submitted article about Black user responses to Black style performances in generative AI outputs. Novel audio-visual capacities have made explicit the performances of identity and ethnicity that are endemic to online existence, whether implicit — as with assumptions of whiteness being linguistically unmarked, or in deliberate performances of Blackness — through raciolinguistic associations of speech styles, and the production of phenotypical audio-visual markers of race by Generative AI technologies. GenAI’s new audiovisual offerings have expanded access to and commercialisation of Black cultural identity. In this paper we introduce a spectrum of Black responses to AI incursions into identity performance; from adoption — as with adopting kin, but also as tool uptake; to skepticism — engaging whilst critiquing and adapting technology to Black liberatory purpose; to refusal — the rejection of AI's exploitation and lack of verisimilitude in performances of Blackness sans Black embodiment.

Notes

Researchers collected a purposive sample of high-engagement videos and users within the field site, and overarching discourse on the topic. Identified commonly used terms and hashtags from this sample for systematic data collection. Researchers then created 'clean' research accounts and used TikTok's search function and filters to identify an initial sample of 236 videos (106 and 139 videos respectively) that were saved into two public 'collections' using TikTok's native organising function.

Popular terms derived from initial exploration were supplemented by TikTok's algorithmic recommendations of related hashtags/search terms. Terms used were: ‘ChatGPT’, ‘blackgpt’, ‘Black AI’, ‘Jamaica chatgpt’, ‘Jamaicangpt’, ‘ChatGPT patois’, ‘ChatGPT pidgin’, ‘ChatGPT creole’, ‘ChatGPT accent’ and ‘ChatGPT Ebonics’. Results were sorted by relevance and date posted, and researchers collated results as well as germane recommended videos suggested by TikTok's algorithm that did not explicitly contain predetermined search terms within captions or hashtags, but were thematically related.

Researchers assessed all content for relevancy, only including content in collection folders that met inclusion criteria as follows:

  1. Content generated by Black users (based on researcher's assessment)
  2. Content is predominantly in English (inclusive of non-standard/vernacular English)
  3. Content concerns GenAI AND Black identity, concerns meta-ontological discourse about (post-)human identity rather than simply prescriptive advice on using GenAI
  4. Content has significant engagement, >10,000 'likes' and/or >100 comments
  5. Content was posted between 25th September 2024 and 7th July 2025 (starting date of data collection process).

After removing content that did not meet inclusion criteria from the initial 236 videos, Zeeschuimer browser extension was used for data capture of the respective public TikTok collections. Duplicate entries were subsequently removed by 'filtering for unique items' using the 4CAT (Capture and Analysis Toolkit) interface, leaving 184 unique TikTok videos. Stijn Peeters. (2026). Zeeschuimer (v1.13.6). Zenodo. https://doi.org/10.5281/zenodo.18670373

Dataset was refined from 184 videos to a sample of 17 by selecting videos with highest engagement. Zeeschuimer browser extension was used for data capture of 683 comments with significant engagement (>10 likes) from remaining sample.

  1. Methods for processing the data: 4CAT (Capture and Analysis Toolkit) interface was used to analyse and process collected data. Processes included generation of .CSV files that included metadata as indicated in above file descriptions; anonymisation of datasets; download of TikTok videos and extracted audio files; transcription of video audio; download of thumbnails of selected videos; and creation of a image wall from extracted thumbnails. Peeters, S., & Hagen, S. (2022). The 4CAT Capture and Analysis Toolkit: A Modular Tool for Transparent and Traceable Social Media Research. Computational Communication Research, 4(2), 571–589. https://doi.org/10.5117/CCR2022.2.007.HAGE

  2. Environmental/experimental conditions: Data was collected by researchers located in USA/UK - noted for potential impact on location-based data gathered. Researchers both created 'clean' research accounts to reduce the influence of personal account preferences and search histories on algorithmic content-ordering processes that determine the content returned by search queries.

  3. People involved with sample collection, processing, analysis and/or submission: Rianna Walcott and Nessa Keddo

Rights