New Directions in Analyzing Text as Data

Text as Data 2026

Monday, October 5, 2026 UC Berkeley Submissions open through August 3

The New Directions in Analyzing Text as Data (TADA) meeting is a leading forum for interdisciplinary research on the study of politics, society, and culture through computational analysis of documents. Recent advances in NLP have the potential to revolutionize how we study human society. But using these tools effectively, reliably, and equitably requires continuous dialog between experts across computational methods, social sciences, and the humanities.

TADA 2026 is part of the Text as Data Association’s conference series, “New Directions in Analyzing Text as Data.”

Hosted at University of California, Berkeley

Mark your calendar

Key Dates

August 3

Submission deadline

One-page PDF proposals are due.

August 31

Notification of acceptance

Decisions are sent to all authors.

October 5

Conference day

Join us at the Krutch Theater on UC Berkeley’s Clark Kerr Campus.

TADA 2026 Schedule

Program

Day 1 · Monday, October 5

Krutch Auditorium

2601 Warring St, Berkeley, CA 94720 Map

  1. 8:00 AMRegistration check-in & breakfast
  2. 9:00 AMWelcome and introduction
  3. 9:15 AM
    Panel 1: MethodologyView papersHide papers

    Discussant: Diag Davenport

    1. 1
      Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements
    2. 2
      Multi-Perspective LLM Annotations for Valid Analyses in Subjective Tasks
    3. 3
      The Confounder Trap: Treatment-Encoding Representations in Causal Inference with Text
  4. 10:30 AM
    Panel 2: NarrativeView papersHide papers

    Discussant: Jeffrey Lockhart

    1. 1
      StoryScope: Investigating idiosyncrasies in AI fiction
    2. 2
      Projecting Agency: The Conversion of Power and Performative Pressure into Language
    3. 3
      Quantifying Racial Disparities in Media Representations of Gun Violence at Scale
  5. 11:45 AM
    Posters 1View postersHide posters
    1. 1
      Inferring Candidate-Voter Issue Alignment from Agentically-crawled Campaign TextKevin Foley
    2. 2
      The Garden of Forking PromptsAdvait Deshmukh, Nora Benedict, Melanie Walsh, Maria Antoniak
    3. 3
      How Local Is Local Law? Measuring Textual Standardization Across U.S. Ordinance CodesDiag Davenport, Carl Illustrisimo
    4. 4
      Can Large Language Models Read Industrial Policy? Expert Benchmarks and a Validated Pipeline for Sector Targeting in ChinaSiying Cao, Wei Chen
    5. 5
      Representing Evaluation in Collective Deliberation: Studying Gender Disparities in Wikipedia Deletion DiscussionsKhandaker Tasnim Huq, Giovanni Luca Ciampaglia
    6. 6
      Assumption Regression: A Failure Mode of Large Language Models in Social Science ResearchSarah K Dreier, Emily Kalah Gade, Sofia Serrano, Lucy H. Lin
    7. 7
      Setting the Gold Standard for Automated Policy Coding – The CAPBENCH-English Benchmark Dataset and Leaderboard for the Comparative Agendas ProjectMiklós Sebők
    8. 8
      Validating LLMs in social science: Emerging norms and epistemic threats from a systematic analysisMeera Desai, Dallas Card, Abigail Z. Jacobs
    9. 9
      The Impact of AI on Social Media-Based Research: Measuring AI-Generated Content Across Platforms, Modalities, and TimeMaty Bohacek, Sergey Sanovich
    10. 10
      An Empirical Investigation of Systematization Forms in GenAI EvaluationKimberly Le Truong, Emily Sheng, Hannah Washington, Solon Barocas, Hanna Wallach
    11. 11
      Persuasion Index: A Theory-Guided Framework for Persuasion AnalysisLiancheng Gong, Zhiyang Wang, Yiwei Xu, Julia Mendelsohn
    12. 12
      Beyond Topic Models: Analysing the United Nations Documentary RecordFrancesco Ignazio Re
    13. 13
      From Telegram Reports to Interception Data: Reconstructing Russian Aerial Attacks on UkraineTaylor Cox
    14. 14
      Measurement and Impacts of Patient AntagonismDaisy Lu
    15. 15
      Sociolinguistic Mobility: A Lens for LLM Pretraining Data CurationJessica Wang
    16. 16
      MMCRCHR: An Evidence-Grounded Agentic AI Research Workbench for the Humanities+Jeffrey Tharsen, Tianfang Zhu, Clovis Gladstone
    17. 17
      Affect is All You Need: Affective Polarization in Anonymous Political DiscussionsEmily Zou, Jennifer Pan
    18. 18
      Mapping the Space Between: Ecological and Relational Proximity in Environmental Governance NetworksXiao Shi
    19. 19
      LLMs in Inductive Work: A FrameworkHope Schroeder, Su Lin Blodgett, Hanna Wallach, Solon Barocas
    20. 20
      How Do Stories Influence Online Conversations?Joel Mire, Sophie Wu, Karina H Halevy, Jocelyn J Shen, Maria Antoniak, Steven R Wilson, Andrew Piper, Maarten Sap
    21. 21
      Understanding AI Cases in the U.S. CourtsDawson Petersen, Eilat Herman, Katherine Conhaim, Kevin Tobia, Nathan Schneider
    22. 22
      Three Years of r/ChatGPT: Societal Impact Evaluations from Social Media DataJessica Dai, Sean Garcia, Emma Pierson, Benjamin Recht, Nika Haghtalab
    23. 23
      AI-Mediated Conversations for Geopolitical Persuasion: Evidence from TaiwanTracy Weener, Ho Chun Herbert Chang
    24. 24
      Taking Turns: Natural Language Inference on Interview DataTom Einhorn
    25. 25
      Historical Vestige in Public Discussion: Commensurable Knowledge Graphs and What Persisted After the Reversal of China's One Child PolicyYehong Deng
    26. 26
      Algorithmic Collusion by LLM Agents: Extracting Pricing Policies and Price RepresentationsPatrick Y. Wu
    27. 27
      How Do People Challenge Racial Stereotypes Online? Counter-Story Detection Across Reddit CommunitiesUma Gunturi, Jimin Mun, Maarten Sap, Maria Antoniak
    28. 28
      A Pragmatics-Grounded Taxonomy of Norm Violations in Human–AI ConversationsJuhyun Oh, Sunnie S. Y. Kim, Hanna Wallach, Jennifer Wortman Vaughan, Hannah Washington
  6. 12:45 PMLunch
  7. 2:00 PMKeynote
  8. 3:15 PM
    Panel 3: ExpertiseView papersHide papers

    Discussant: David Bamman

    1. 1
      From Text to Large-n Datasets: Applying lessons about human coder disagreement to machine coders
    2. 2
      Decoding, Denoising, and Deciphering the Babylonian Astronomical Diaries with LLMs and Science-Informed Bayesian Models
    3. 3
      APPEAL: Attributing Persuasive Power to Expert-Specified Actions in LLMs
  9. 4:30 PM
    Posters 2View postersHide posters
    1. 1
      Evaluating the Use of Synthetic Data in Humanities PretrainingCraig Messner
    2. 2
      Measuring Family Identity Signaling from Political Campaign BiographiesJoy Youn Soo Kim
    3. 3
      Semi-Automatic Discourse Networks: Method, Software, and a First ApplicationRyan Williams
    4. 4
      In Your Own Words: Computationally Identifying Interpretable Themes from Free-Text Survey DataJenny Shan Wang, Aliya Saperstein, Emma Pierson
    5. 5
      The Ideal Surveytaker Problem: Language Models and the Limits of Silicon SamplesAlexis K Palmer
    6. 6
      Transition of Language in the Catholic Church: Analysis of Linguistic Shifts in Papal Encyclicals Pre vs. Post-Vatican IISimon Dovan Nguyen
    7. 7
      Universal Feature DictionariesKenny Peng, Jon Kleinberg, Nikhil Garg
    8. 8
      Measuring Academic Written Register at Scale: Automated Text Features in 16,505 Student EssaysSining Tao, Bailey Buchanan
    9. 9
      Selective Public Uptake of the Make America Healthy Again (MAHA) Campaign across Seven Online PlatformsHaoning Xue, Yue Li, Benjamin Lyons, Andy J. King
    10. 10
      Do LLMs Distort Public Opinion? Comparing Respondent Representation Across Topic Modelling ArchitecturesMasha Krupenkin, Andrew Gordon, David Rothschild
    11. 11
      What do language models know about the past? Recovering perspectival knowledge using large-scale language modelsLaura Nelson, Tom Einhorn, Yash Mali
    12. 12
      Standardizing Culture before Mass Schooling: The Examination State and Regulated Verse in Tang ChinaZhaomin Li
    13. 13
      The Survey Sampling Roots of Prediction-Powered Inference: Implications for Text as Data ResearchReagan Mozer
    14. 14
      Freezing and Thawing in Text: Binomial Order as a Measure of Cultural ChangeKatherine Van Koevering
    15. 15
      Subsidizing the Messenger: How Connected Owners Shape Economic NewsBahar Zafer
    16. 16
      The Art of Saying Nothing: Measuring Discursive Self-Censorship in Institutional Art Discourse Across Political RegimesXinyuan Hou, Jiedong Zhang
    17. 17
      Assessing Coders (Human and AI) Without Gold LabelsKarl Rohe
    18. 18
      Suspense in Short Fiction: Testing Theory Pluralism against LLMs’ Intrinsic ConceptsSvenja Guhr
    19. 19
      Measuring the Reach and Extremism of Podcasts that Feature U.S. PoliticiansBenjamin Roger Litterer, Amber Boydstun, Dallas Card, David Jurgens
    20. 20
      Are Economists Open to AI? A Text-as-Data-as-Survey Approach via Language ModelsLei Ge, Yi Wang
    21. 21
      What Do Conversation-Level Language Measures Measure? Evidence From a Randomized Experiment With an LLM Health ChatbotQiyao Peng, Jiaying Liu
    22. 22
      Which Documents Deserve an Expensive Second Look? Measuring Corpus Quality under a Fixed Adjudication BudgetNicholas James
    23. 23
      TOPICBench: A Systematic Evaluation of LLMs for Labeling Latent TopicsAya Salim, Adam Visokay, Henry George Xu, Michael Lachanski
    24. 24
      What Are They Gating? Context Engineering Evidence and Political Censorship in Large Language ModelsTongzhou Liu, Yifan Tian
    25. 25
      Deeds into Data: A Living Citation Graph for Georeferencing County Land RecordsOliver Payne Dahl
    26. 26
      Measuring Personal Culture in Open-ended RespondingAlejandro Sarria
    27. 27
      The Substance of Union Contracts: Measuring the Generosity of U.S. Collective Bargaining Agreements at ScaleMatthew Bramwell Bone, Prashant Garg, Chenxi Li, Zachary Parolin
  10. 5:30 PMConcluding remarks
  11. 6:00 PMEnd of day 1; dinner on your own

Day 2 · (Optional) Lightning Talks

Tilden Room at MLK Jr. Building

2495 Bancroft Way, Berkeley, CA 94720 Map

  1. 9:00 AMBreakfast
  2. 10:00 AMLightning talks 1
  3. 11:15 AMLightning talks 2
  4. 12:15 PMEnd of day 2; lunch on your own

Submit your work

Call for Presentations

  • A single-page PDF — LaTeX encouraged.
  • Anonymized — omit author names and affiliations
  • Preliminary results as a chart or visualization Optional
  • Figures and tables must be within the one page limit. References do not count toward the 1 page limit.

TADA 2026 is a non-archival conference; there are no formal proceedings, and papers presented at the conference will not be distributed publicly by the conference. Presenters are expected to provide a paper to their discussant two weeks before the conference. We welcome any work, so long as it hasn’t been previously presented at a TADA conference. We also welcome individuals to volunteer to serve as discussants.

Organizers for this year are: Amber Boydstun, David Bamman, Dallas Card, David Mimno, Ella Montgomery, Denis Peskoff, Brendan O'Connor, Brandon Stewart, Adam Visokay, Leah Windsor, Naitian Zhou.

Sponsors for this year are Bloomberg, Microsoft, and the UC Berkeley School of Information.

AI policy: AI use beyond copy editing is strongly discouraged. The conference organizers include authors of AI-detection research.

Getting here

Venue & Travel

Krutch Theater

Clark Kerr Campus · UC Berkeley

2601 Warring St, Berkeley, CA 94720
Open in Google Maps

By BART

Downtown Berkeley and Ashby are each 1.5 miles from Clark Kerr Campus. Buses and bike shares are available, and the area is pedestrian-friendly.

By car

Berkeley is a 30-minute drive from downtown San Francisco, longer in rush-hour traffic. Clark Kerr Campus has some parking.

By air

Oakland (OAK) is the closest airport; San Francisco (SFO) is likely cheaper. Both connect to Berkeley via BART.

Questions?

We’re happy to help with submissions, logistics, or anything else.

textasdata2026@gmail.com