Posts by Collection

portfolio

publications

Increased Fox News Viewership Is Not Associated with Heightened Anti-Black Prejudice

Under review, 2023

Today’s media environment provides Americans with unparalleled choice in how or whether to watch political TV news. Prior studies have focused on the impacts of the growing range of ideological slants on vote choice. But the fragmentation of the media landscape may also increase variation in the coverage of race-related topics. With a large audience and programs that even some employees thought conveyed racism, Fox News provides a valuable case study. We use a population-based panel 2008–2020 to measure the associations between changes in self-reported Fox News viewership and race-related attitudes and thus bound Fox News’ likely effects assuming positive selection. Difference-in-difference models demonstrate that increased Fox News watching is not strongly associated with increases in Whites’ anti-Black prejudice or opposition to government assistance targeting Black Americans. However, those whose Fox News watching increased grew increasingly anti-immigration. These results indicate the limits of Fox News’ impacts on racial prejudice.

With Daniel J. Hopkins and Yphtach Lelkes.
Download Paper

Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels

Proceedings of the Sixth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS), 2024

Computational social science (CSS) practitioners often rely on human-labeled data to fine-tune supervised text classifiers. We assess the potential for researchers to augment or replace human-generated training data with surrogate training labels from generative large language models (LLMs). We introduce a recommended workflow and test this LLM application by replicating 14 classification tasks and measuring performance. We employ a novel corpus of English-language text classification data sets from recent CSS articles in high-impact journals. Because these data sets are stored in password-protected archives, our analyses are less prone to issues of contamination. For each task, we compare supervised classifiers fine-tuned using GPT-4 labels against classifiers fine-tuned with human annotations and against labels from GPT-4 and Mistral-7B with few-shot in-context learning. Our findings indicate that supervised classification models fine-tuned on LLM-generated labels perform comparably to models fine-tuned with labels from human annotators. Fine-tuning models using LLM-generated labels can be a fast, efficient and cost-effective method of building supervised text classifiers.

Nicholas Pangakis and Samuel Wolken. 2024. Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels. In Proceedings of the Sixth Workshop on Natural Language Processing and Computational Social Science (NLP+CSS 2024), pages 113–131, Mexico City, Mexico. Association for Computational Linguistics.
Download Paper

The Rise of and Demand for Identity-Oriented Media Coverage

American Journal of Political Science, 2024

While some assert that social identities have become more salient in American media coverage, existing evidence is largely anecdotal. An increased emphasis on social identities has important political implications, including for polarization and representation. We first document the rising salience of different social identities using natural language processing tools to analyze all tweets from 19 media outlets (2008–2021) alongside 553,078 URLs shared on Facebook. We then examine one potential mechanism: Outlets may highlight meaningful social identities—race/ethnicity, gender, religion, or partisanship—to attract readers through various social and psychological pathways. We find that identity cues are associated with increases in some forms of engagement on social media. To probe causality, we analyze 3,828 randomized headline experiments conducted via Upworthy. Headlines mentioning racial/ethnic identities generated more engagement than headlines that did not, with suggestive evidence for other identities. Identity-oriented media coverage is growing and rooted partly in audience demand.

Hopkins, D. J., Lelkes, Y., & Wolken, S. (2024). The rise of and demand for identity‐oriented media coverage. American Journal of Political Science.
Download Paper

Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI

International AAAI Conference on Web and Social Media (ICWSM), 2025

Automated text annotation is a compelling use case for generative large language models (LLMs) in social media re-search. Recent work suggests that LLMs can achieve strongperformance on annotation tasks; however, these studies evaluate LLMs on a small number of tasks and likely suffer from contamination due to a reliance on public benchmark datasets. Here, we test a human-centered frameworkfor responsibly evaluating artificial intelligence tools usedin automated annotation. We use GPT-4 to replicate 27 annotation tasks across 11 password-protected datasets fromrecently published computational social science articles in high-impact journals. For each task, we compare GPT-4 an-notations against human-annotated ground-truth labels andagainst annotations from separate supervised classificationmodels fine-tuned on human-generated labels. Although thequality of LLM labels is generally high, we find significant variation in LLM performance across tasks, even withindatasets. Our findings underscore the importance of a human-centered workflow and careful evaluation standards: Automated annotations significantly diverge from human judgment in numerous scenarios, despite various optimizationstrategies such as prompt tuning. Grounding automated annotation in validation labels generated by humans is essentialfor responsible evaluation.

Nicholas Pangakis and Samuel Wolken (2024). Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI. Proceedings of the International AAAI Conference on Web and Social Media.

Unpacking Media Bias in the Growing Divide Between Cable and Network News

Scientific Reports, 2025

The potential for a large, diverse population to coexist peacefully is thought to depend on the existence of a public sphere in which citizens are exposed to similar facts about similar topics. A generation ago, broadcast television news was widely considered to serve this function; however, since the rise of cable news in the 1990s, critics and scholars have worried that the corresponding fragmentation and segregation of audiences has caused this baseline of common understanding to be lost. Recent work documents that millions of Americans are loyal consumers of cable TV news stations. However, the implications of partisan segregation in TV news consumption depend on bias in content—which topics TV news programs talk about and the language they use to talk about them. Here, we measure bias in the production of TV news at scale by analyzing nearly a decade of TV news (Dec. 2012–Oct. 2022) on the largest cable and broadcast stations. We quantify the share of attention each station devoted to more than 20 politically significant topics as well as the linguistic similarity of different stations’ news coverage of those topics. We find that while broadcast news continues to cover similar topics with similar language, cable news stations have become increasingly distinct, both from broadcast news and from each other, diverging in terms of both content and language. This trend is driven by hard news as much as partisan commentary programs. Our results show that changes in the supply, not just consumption, of TV news are contributing to Americans’ polarizing media diets.

Hosseinmardi, H., Wolken, S., Rothschild, D. M., & Watts, D. J. (2025). Unpacking media bias in the growing divide between cable and network news. Scientific Reports, 15, 17607.
Download Paper

The Limits of De-Politicizing–and Also of Annotation: A Case Study in Russian Media Outlets’ Social Media Posts, 2016–2024

Proceedings of the International AAAI Conference on Web and Social Media (ICWSM), 2026

This paper examines how Russia’s February 2022 invasion of Ukraine affected political coverage on social media. We analyze more than two million posts from Russian-language media outlets on VK, Facebook, and Telegram, manually annotating 6,661 posts and supplementing this analysis with large language models. Political posts spiked after the invasion, and social media users became more likely to engage with political posts relative to non-political posts. The findings suggest that autocratic regimes may abandon de-politicization in favor of more invasive approaches when major events drive heightened engagement with the news, and the paper evaluates both the utility and limitations of applying LLMs to Russian-language text analysis.

Gong, L., Hopkins, D. J., & Wolken, S. (2026). The Limits of De-Politicizing–and Also of Annotation: A Case Study in Russian Media Outlets' Social Media Posts, 2016–2024. Proceedings of the International AAAI Conference on Web and Social Media, 20(1), 889–909.
Download Paper

Systemic electioneering from the evangelical pulpit: Evidence from a computational analysis

Proceedings of the National Academy of Sciences (PNAS), 2026

Religious institutions’ engagement in prohibited electoral advocacy is a growing concern for democratic governance. In the United States, such mobilization has been especially visible within the evangelical movement. This study examines the phenomenon using a corpus of 88,546 sermons from predominantly evangelical churches, transcribed from 63,683 hours of Sunday services spanning the 2020, 2022, and 2024 election cycles and a nonelection control period. Analysis of this corpus reveals that direct political advocacy and endorsements are widespread: 14.7% of churches engaged in this speech during the three months surrounding elections.

Jacob, M. S., Lelkes, Y., Wolken, S., & Westwood, S. J. (2026). Systemic electioneering from the evangelical pulpit: Evidence from a computational analysis. Proceedings of the National Academy of Sciences, 123(21), e2603911123.
Download Paper

talks

The rise of and demand for identity‐oriented media coverage

Published:

I presented “The rise of and demand for identity‐oriented media coverage,” co-authored with Dan Hopkins and Yphtach Lelkes, to Alexander Theodoridis’s seminar on causal inference in the Department of Political Science.

teaching