A study by the Institute for Strategic Dialogueue found that leading AI chatbots gave flawed answers to election-related questions, including on mail-in voting, voter registration and polling locations.
Researchers reviewed 2,400 responses from six AI chatbots to election questions across 10 US states. They assessed whether answers were accurate, current and well sourced, and found weaker results when similar questions were asked in Spanish rather than English.
According to commentary from Neurologyca Chief Strategy Officer Marc Fernandez, the Spanish-language performance gap averaged 16%. He said the problem went beyond factual accuracy and exposed limits in how AI systems handle user intent in high-stakes situations.
Election information is a difficult test for generative AI tools because rules can vary by state and county, and can change between election cycles. A user asking whether they can still register to vote may need an answer that reflects urgency, deadlines and local procedures, rather than a generic summary.
Fernandez said the findings point to a broader issue for AI systems used in sensitive areas such as health, benefits, finance and personal safety. In those settings, errors or omissions can have direct consequences for users acting on the information they receive.
"A new study from the Institute for Strategic Dialogueue tested 2,400 responses from six leading AI chatbots on election questions across 10 states. One result stood out: when similar questions were asked in Spanish instead of English, performance dropped by an average of 16%.
That matters because more people are turning to AI to find information and make decisions. Elections are a tough test for these systems. Rules vary by state, sometimes by county, and they can change from one cycle to the next.
The Spanish-language gap is troubling on its own. People shouldn't get worse information simply because they ask a question in Spanish.
But the language gap isn't the only problem. The study looks at whether answers were accurate, current and well sourced. Those things matter. But in a high-stakes situation, a system also needs to understand what the person is trying to do.
Take a voter asking, 'Can I still register?' They are probably not asking out of curiosity. They are trying to vote, and they may be running out of time. A good response should recognise that urgency. It should lead with the deadline and treat location as part of the answer. If the rules changed recently, it should say so. If anything is uncertain, it should point the voter to their election office to confirm.
That context is the difference between an answer that helps and one that just happens to be right. Elections make the problem easy to see because the stakes are clear and timing matters. The same is true in health, benefits, money and personal safety.
AI has become very good at producing answers that sound convincing. It still has a long way to go in understanding why someone is asking, what they are trying to accomplish and what could happen if the answer sends them in the wrong direction. That is the human-context problem. It means understanding the goal, situation and intent behind a request.
It is also the problem my company works on, but this is much bigger than any one company. As AI takes on more responsibility, people are going to expect more than a correct answer. They are going to expect the system to understand what they actually need," said Marc Fernandez, Chief Strategy Officer, Neurologyca.
The concerns come as chatbots are increasingly used as a first stop for practical questions once answered through official websites, call centres or community organisations. Election guidance is especially vulnerable to error because deadlines, eligibility rules and voting methods are often highly specific and time-sensitive.
Researchers focused on routine but critical questions, including whether a voter could still register, how to cast a mail ballot and where to vote. Inaccurate or outdated responses in any of those areas could leave users unable to complete the voting process on time.
The language gap identified in the study also raises questions about equal access to reliable civic information through AI tools. If Spanish-language users receive weaker answers to the same questions, the shortfall could deepen existing information barriers rather than reduce them.
Fernandez framed the issue as one of context as much as content. In his view, systems must recognise not just the topic of a question, but the likely objective behind it and the consequences of giving an answer that is technically plausible but practically incomplete.
That debate is likely to extend beyond elections. As more people use AI assistants for guidance on public services and personal decisions, scrutiny is increasing over how these systems source information, reflect local rules and respond when certainty is limited.
The study suggests that convincing language remains a poor substitute for dependable guidance when the stakes are immediate and the margin for error is small.