Network Contagion Research Institute Intelligence Report
7
profile’s search preferences (e.g., no accounts were followed, no prior searches were
performed, no engagements except views and saves were performed).
A standard collection methodology was followed for all keywords across each platform. In
conducting the TikTok data collection, the user began by typing the term into the Search field
and clicking on the first video that appeared. Subsequently, the user scrolled through each
subsequent video, saving each one. Every video was played for at least 15 seconds or until the
video concluded. Upon completing the recording session, the user navigated to the Saved page
on the User Profile to locate all the saved videos from the session. The upload date and the link
for each video were then copied into a spreadsheet.
For Instagram data collection, the user started by entering the term in the Search field and
selecting the first post that appeared. The user then scrolled through each subsequent post,
saving each one. Videos were played for at least 15 seconds or until they finished, and posts
containing multiple images were fully viewed by swiping through each image. After recording,
the user accessed the Saved page on the User Profile to retrieve all saved posts from the
session. The upload date and the link for each post were copied into a spreadsheet.
During YouTube data collection, the user entered the term into the Search field and hovered
over each video in the list, allowing it to play for 15 seconds (excluding YouTube Shorts and
videos that are in playlists). After completing the recording, the user scrolled back to the top of
the list and clicked on each video to copy the upload date and URL into a spreadsheet.
Coding Methodology
Following data collection, the first phase of analysis categorized content as either pro-China,
anti-China, neutral, or irrelevant. The search terms analyzed are inherently political for the CCP,
facilitating a clearer and more intuitive coding process than would be possible with broader,
more neutral terms such as “China” or “democracy”.
Blind coding12 was performed by two human analysts. In cases where there was a disagreement
in the coding determination, a third subject matter expert (SME) was tasked with arbitrating and
assigning a final coding category. Given the subtlety of messaging alongside the intermixture of
textual, visual, and audio semiotics that were assessed in making a coding determination, we
believe that human analysis tends to be more accurate and granular than AI or machine-driven
coding.
Despite reliance on human analysts for content coding, our average margin of intercoder
disagreement across all platforms and keywords was remarkably small. The average intercoder
disagreement across all keywords and platforms was measured as 12.81%, with a range of
12
In this context, blind coding connotes that the analysts categorizing video results were isolated from one another
to prevent conformity bias between their designations. Furthermore, this research was hypothesis agnostic in order
to minimize coder bias.