Two research experiments comparing the cost-effectiveness of using LLM chatbots to persuade voters estimated that LLM-based persuasion costs between $48 and $75 per voter, compared to traditional methods that are equally persuasive but cost $100. However, we now know that traditional methods scale more effectively. The paper was published in. Journal of Experimental Politics.
Large-scale language models (LLMs) are artificial intelligence systems trained on large collections of text to generate human-like answers, explanations, stories, and conversations. It is increasingly being considered as a tool for political communication and campaigning, as it can quickly create consistent content and tailor responses to individual users.
Unlike traditional advertising, LLM chatbots can conduct long conversations, responding to voters’ concerns and tailoring the discussion to that person’s beliefs, values, and emotional responses. These capabilities have raised concerns that governments, political groups, or foreign actors could use LLMs to manipulate public opinion, spread propaganda, and deepen social divisions on an unprecedented scale.
This technology could make it cheaper and easier to mass-produce personalized political messages, including misleading content that appears believable and spontaneous. However, being able to craft a persuasive message does not automatically mean that such a message will influence a large number of voters.
Political persuasion involves not only persuading people who encounter a message, but also capturing their attention and persuading them to engage with the message in the first place. However, the difficulty in getting large numbers of people to participate in political chatbots currently severely limits their real-world impact.
Study author Zhongren Chen, a researcher at Yale University, and his colleagues conducted two experiments testing different methods of political persuasion, including those involving chatbots. They also measured immediate and long-term attitude changes rather than measures of perceived persuasion. Finally, we considered issues related to ensuring voter engagement with LLM content. This allowed us to estimate the real-world threat that LLMs may pose to democratic processes.
Participants in the first experiment were 5,150 people who were recruited online to complete an online survey as part of a study comparing traditional human persuasion methods and LLM-driven methods on attitudes toward immigration policy.
They were randomly assigned to one of four experimental conditions. Participants in the first condition watched a video unrelated to immigration, which served as a placebo condition. In the second condition, human persuasion, participants watched a 3-minute video of a human advocate making a pro-immigration case. That person was a teacher who shared personal reasons for supporting the proposed immigration policy.
In the third condition, participants had an interactive conversation with an AI chatbot disguised as a human that presented pro-immigration arguments. In the fourth condition, participants engaged in a conversation with an AI chatbot that presented pro-immigration arguments while identifying itself as an AI. Both chatbot conditions used Claude 3.5 Sonnet. Immediately after the experimental condition and 5 weeks later, participants rated their agreement with the advertised policy.
The second study tested the same persuasion method on three different problems. The first was that illegal immigrants should be eligible for in-state college tuition. The second idea was that transgender people should be allowed to use the bathroom that corresponds to their gender identity. Third, the federal minimum wage “should not be raised” from the current $7.25 per hour to $15 per hour. The inclusion of this third policy was intended to test the effectiveness of persuasion in both liberal and conservative directions. Finally, the study authors conducted numerical simulations to estimate the cost-effectiveness of different persuasion methods.
Results from the first study showed that both human and chatbot-based persuasion methods were effective compared to a placebo condition. However, there was no difference in effectiveness between the chatbot and human persuasion conditions immediately after treatment and after 5 weeks. Additionally, there was no difference in persuasiveness between the two AI-based conditions. Overall, this study suggests that AI chatbots can be just as persuasive as watching a human video.
A second study found similar results. Although there were some differences across topics, the results showed either no consistent differences in persuasion or that human persuasion was slightly more effective.
Simulations conducted by the study authors showed that LLM-based persuasion costs between $48 and $75 per voter, compared to $100 per voter for traditional methods. However, they note that traditional methods currently scale more efficiently.
“LLM does not yet offer significant potential for large-scale persuasion, but this may change as capabilities improve and scalable exposure techniques become feasible,” the study authors concluded.
This study contributes to the scientific understanding of the political persuasion capacity of LLMs. However, the functionality of LLM depends on the model used, the resources allocated to it, and the persuasion methods used. Studies using different persuasion methods or testing other LLM models may not yield the same results. Furthermore, participants were informed that they were working on AI as part of their scientific research, which may have increased their perception of AI’s neutrality. Skepticism about AI in the real world can weaken these persuasive effects.
The paper, “A Framework for Assessing the Persuasion Risks of Large-Scale Language Models that Chatbots Bring to Democratic Societies,” was authored by Zhongren Chen, Joshua Kalla, Quan Le, Shinpei Nuclear-Sakai, Jasjeet Sekhon, and Ruixiao Wang.

