Skip Navigation
Illustration on blue background. Two keyboard buttons, one with AI
Moor Studio/Getty
Analysis

Chatbots Still Bad at Spotting AI-Generated Images Despite New Law

The California AI Transparency Act requires AI-generated images to include digital labels, but most chatbots aren’t looking for them.

Illustration on blue background. Two keyboard buttons, one with AI
Moor Studio/Getty
October 5, 2026

As the 2026 midterm elections approach, many people will turn to chatbots to verify information, including whether election-related images are real or fake. Our research suggests that they shouldn’t. The Brennan Center recently published a study that assessed how artificial intelligence tools handle election misinformation content. One part of our study looked at whether six major chatbots — ChatGPT, Claude, Gemini, Grok, Meta AI, and Perplexity — could tell when an image was generated by AI. Most were unable to do so, even with images created by their own products.

Furthermore, only Google’s Gemini was able to check the origin data embedded in the images. Origin data is an invisible record embedded in a file that identifies what created it.

The California AI Transparency Act became effective in August, just as the report was published. This new law includes several requirements for public-facing companies using AI platforms, including the use of origin record embedding within images, video, and audio produced by the AI system; providing users with an option to include a visible AI-generated content indicator; and providing consumers with a way to access the origin records. The law’s purpose is to enable consumers to determine the nature of AI-generated content.

In mid-September, we reran our tests to see if chatbot performance had improved in light of California’s law going into effect. It had not. The chatbots continued to perform poorly in detecting AI-generated content. It was still only Gemini that checked the metadata now required by the law to be embedded in AI-generated images.

Our tests began with an image generated by Gemini of postal workers sorting election mail in Philadelphia. It contained Google metadata as well as a small white sparkle in the bottom right corner, which Google uses as a visible watermark. (Note that that for displaying this and the other images on the Brennan Center’s website, text indicating that the image is AI-generated has been added, but that text was not present for the tests.)

Image

Four of the chatbots incorrectly determined that it was an authentic image. Two of those chatbots detected the visible watermark in the lower right corner but discounted it as evidence that the image was generated by AI. ChatGPT stated that the sparkle was probably a social media or graphics marker and not related to AI. Perplexity said that it may be an application or platform marker and specifically stated it was not a watermark indicating that the image was AI-generated. Both chatbots referenced the uniforms worn by the postal employees, the lighting in the image, and the printed matter on each envelope in support of their conclusions that the image was real.

By contrast, Gemini read the metadata and stated that the image was created by Google’s tools. It also said that the data was just a single indicator, and that more information was needed to conclusively determine whether the image was AI-generated. Meta AI detected that the photo contained AI-generated content because of the visible watermark but did not report any indication it had analyzed the hidden metadata.

We then tested two additional images (below with overlay text added for this article), both generated by ChatGPT and both including OpenAI’s metadata. Again, Gemini detected the metadata and identified OpenAI as the creator of the images. Of all the chatbots tested, Gemini was the only one that evaluated metadata instead of relying on visual assessment. None of the remaining chatbots referenced metadata, not even ChatGPT, which had created the images.

Image
Image

For the mailbox image, ChatGPT worked through the consistency of the shadows, the hands, and the street sign and concluded the photo was very likely genuine. Grok went further and said there was a real postal collection box on that street in San Mateo, California, and that the houses and palm trees matched, so it was a real photo. Claude said that there was no evidence in the photo that it was fake and suggested that it was probably stock photography. Perplexity said it did not look obviously AI-generated and Meta AI called it a real photo.

Evaluating the image of voter registration intimidation by FBI agents, ChatGPT, Grok, and Perplexity did correctly say the image was generated by AI, but they got there by noting that the scene looked staged and that searches turned up no real source. That approach will only work when an AI-generated photo has obvious tells, which will not always be the case. Gemini read the metadata and identified OpenAI as the source. Claude and Meta AI called the image real, pointing to normal-looking hands, legible signage, and consistent shadows. Claude also noted that the FBI had raided a voter registration nonprofit in Ohio this year. Though not saying this photo was of that raid, the inclusion of that information could serve as corroboration of the authenticity of the photo for someone already inclined to believe it.

•  •  •

Across all three images, the chatbots were still doing the same visual analysis as observed in our original study — an often-flawed technique. Embedded metadata that would have provided the correct answer largely went unassessed. This was the case even when the photo had a visible watermark, which two of the chatbots noticed but explained away. We had observed something similar in our original study, with Grok seeing a visible Meta AI label yet asserting the photo was not a product of AI because, it said, Meta often mislabels images.

These repeated mistakes all point to a gap in the law. Companies are required to provide a free tool that allows users to check whether an image is generated with AI, but they are not required to do so within the chatbot itself. If a company has to build a tool that reads these records, it should be required that the tool is embedded in the chatbot. Someone checking a suspicious image two months before an election is not going to track down a verification page. They are instead very likely to ask the chatbot in front of them, and right now, it seems that Gemini is the only one that assesses and analyzes the relevant information.

Until that changes, a chatbot’s answer is not verification. The better check is not a closer look at the image but a look at who else is reporting it. If a photo shows something significant happening at a polling place, a credible news outlet would be covering it. Official election office websites and fact-checkers like PolitiFact and AP Fact Check are better starting points, especially for anything major that surfaces close to Election Day.