This article was written by an actual, flesh-and-blood human â me â but an increasing amount of the text and video content you come across online is not. Itâs coming from generative AI tools, which have gotten pretty good at creating realistic-sounding text and natural-looking video. So, how do you sort out the human-made from the robotic?
The answer is more complicated than that urban legend about the overuse of em-dashes would have you believe. Lots of people write with an (over)abundance of that particular piece of punctuation, as any editor will tell you. The clues may have more to do with the phrasing and the fact that, as with any writer, large language models tend to repeat themselves.
Thatâs the logic behind AI-detection programs. The problem is that those systems are often AI-powered themselves, and they provide few details about how they arrived at their assessments. That makes them hard to trust.
A new feature from the AI-detection company Copyleaks, called AI Logic, provides more insight into not just whether and how much of something might have been written by AI, but what evidence itâs basing that decision on. What results is something that looks a lot like a plagiarism detector, with individual passages highlighted. You can then see whether Copyleaks flagged it because it matched text on a website known to be AI-generated, or if it was a phrase that the companyâs research has determined is far more likely to appear in AI-produced than human-written text.
You donât even necessarily have to seek out a gen AI tool to produce text with one these days. Tech companies like Microsoft and Google are adding AI helpers to workplace apps, but itâs even showing up in dating apps. A survey from the Kinsey Institute and Match, which owns Tinder and Hinge, found that 26% of singles were using AI in dating, whether itâs to punch up profiles or come up with better lines. AI writing is inescapable, and there are times when you probably want to know whether a person actually wrote what youâre reading.Â
This additional information from a Copyleaks-checked text marks a step forward in the search for a way to separate the AI-made from the human-written, but the important element still isnât the software. It takes a human being to look at this data and figure out whatâs a coincidence and whatâs concerning.
âThe idea is really to get to a point where there is no question mark, to provide as much evidence as we can,â Copyleaks CEO Alon Yamin told me.
A noble sentiment, but I also wanted to see for myself what the AI detector would detect and why.
How AI detection works
Copyleaks started out by using AI models to identify specific writing styles as a way to detect copyright infringement. When OpenAIâs ChatGPT burst on the scene in 2022, the company realized it could use the same models to detect the style of large language models. Yamin called it âAI versus AI,â in that models were trained to look for specific factors like the length of sentences, punctuation usage and specific phrases. (Disclosure: Ziff Davis, CNETâs parent company, in April filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)
The problem with using AI to detect AI is that large language models are often a âblack boxâ â theyâll produce an output that makes sense, and you know what went into training them, but they donât show their work. Copyleaksâ AI Logic function tries to pull back the veil so people have a better sense of what in the copy theyâre evaluating might actually be AI-written.Â
âWhatâs really important is to have as much transparency around AI models [as possible], even internally,â Yamin said.
Read more: AI Essentials: 29 Ways to Make Gen AI Work for You, According to Our Experts
AI Logic uses two different approaches to identify text potentially written by an LLM. One, called AI Source Match, uses a database of AI-generated content from sources either created in-house by Copyleaks or on AI-produced sites online. This works much like a traditional plagiarism detector. âWhat weâve discovered is that AI content, a lot of the time, if you ask the same question or a similar question over and over again, youâll get similar answers or a similar version of the same answer,â Yamin said.
The other component, AI Phrases, detects terms and groups of words that Copyleaksâ research has determined are far more likely to be used by LLMs than by human writers. In one sample report, Copyleaks identified the phrase âwith advancements in technologyâ as potentially AI-written. Copyleaksâ analysis of generated content found that the phrase appeared 125 times per million AI-written documents, compared with just six times per million documents written by people.
The question is, does it work?
Can Copyleaks spot AI content and explain why?
I ran a few documents through Copyleaks to see if AI Logic can identify what I know to be AI-created content, or if it flags human-written content as AI-written.
Example: A human-written classic
What better way to test an artificial intelligence tool than with a story about artificial intelligence? I asked Copyleaks to test a section of Isaac Asimovâs classic 1956 short story The Last Question, about a fictional artificial intelligence tasked with solving a difficult problem. Copyleaks successfully identified it as 100% matched text on the internet and 0% AI-written.Â
Example: Partially AI-written
For this example, I asked ChatGPT to add two paragraphs of additional copy to a story I wrote and published earlier in the day. I ran the resulting text â my original story with the two AI-written paragraphs added at the bottom â through Copyleaks.Â
Copyleaks successfully identified that 65.8% of this copy matched existing text (because it was literally an article already on the internet), but it didnât pick up anything as being AI-generated. Those two paragraphs ChatGPT just wrote? Flew completely under the radar.Â
Enlarge Image
Copyleaks thought everything in this article was written by AI, even though only a few paragraphs were.
I tried again, this time asking Googleâs Gemini to add some copy to my existing story. Copyleaks again identified that 67.2% of the text matched what was online, but it also reported that 100% of the text may have been AI-generated. Even text I wrote was flagged, with some phrases, like âgenerative AI model,â described as occurring more frequently in AI-written text.Â
Example: Totally AI-written
In a test of generative AIâs ability to create things that are totally out of touch with reality, I asked it to write a news story as if the Cincinnati Bengals had won the Super Bowl. (In this fictional universe, Cincinnati beat the San Francisco 49ers by a score of 31-17.) When I ran the fake story through Copyleaks, it successfully identified it as entirely AI-written.Â
Enlarge Image
Copyleaksâ AI Logic quickly realized this story about the Cincinnati Bengals winning the Super Bowl was written by an AI chatbot.
What Copyleaks didnât do, however, is explain why. It said no results were found in its AI Source Match or its AI Phrases, but with a note: âThere is no specific phrase that indicates AI. However, other criteria suggest that this text was generated by AI.âÂ
I tried again, this time with a different ChatGPT-generated story about the Bengals winning the Super Bowl 27-24 over the 49ers, and Copyleaks provided a more detailed explanation. It calculated the content was 98.7% AI-created, with a handful of phrases singled out. These included some seemingly innocent terms like âmade several criticalâ and âtestament to years of.â It also included some strings of words that spread across multiple phrases or sentences, like âcontinues to evolve, the Bengalsâ future,â which apparently occurred 317 times more frequently in the databaseâs AI-generated content than in human text documents. (After raising the issue with the first attempt with Copyleaks, I tried it again and got similar results to this second test.)
Just to be sure it wasnât operating entirely on the fact that the Bengals have never won a Super Bowl, I asked ChatGPT to write an article about the Los Angeles Dodgers winning the World Series. Copyleaks found that 50.5% matched existing text online, but also reported it was 100% AI-written.Â
A high-profile example
Copyleaks did some testing of its own, using a recent example of a controversial alleged use of AI. In May, the news outlet NOTUS said that a report from the Trump administrationâs Make America Healthy Again Commission contained references to academic studies that did not exist. Researchers who were cited in the MAHA report told media outlets that they did not produce that work. Citations to nonexistent sources are a common result of AI hallucination, which is why itâs important to check anything an LLM cites. The Trump administration defended the report, with a spokesperson blaming âminor citation and formatting errorsâ and stating that the substance of the report remains unchanged.Â
Copyleaks ran the report through its system, which reported finding 20.8% potential AI-written content. It found some sections around childrenâs mental health raised red flags in its AI Phrases database. Some phrases that occurred far more often in AI-written text included âimpacts of social media on theirâ and âThe Negative Impact of Social Media on Their Mental Health.â
Can an AI really detect AI-written text?
In my experience, the increased transparency from Copyleaks into how the tool works is a step forward for the world of AI detection, but this is still far from foolproof. Thereâs still a troubling risk of false positives. In my testing, sometimes words I had written just hours before (and I know AI didnât play a role in them) could be flagged because of some of the phrasing. Still, Copyleaks was able to spot a bogus news article about a team that has never won a championship doing so.Â
Yamin said the goal isnât necessarily to be the ultimate source of truth but to provide people who need to assess whether and how AI has been used with tools to make better decisions. A human needs to be in the loop, but tools like Copyleaks can help with trust.Â
âThe idea in the end is to help humans in the process of evaluating content,â he said. âI think weâre in an age where content is everywhere, and itâs being produced more and more and faster than ever before. Itâs getting harder to identify content that you can trust.â
Hereâs my take: When using an AI detector, one way to have more confidence is to look specifically at what is being flagged as possibly AI-written. The occasional suspicious phrase may be, and likely is, innocent. After all, there are only so many different ways you can rearrange words â a compact phrase like âgenerative AI modelâ is pretty handy for us humans, same as for AI. But if itâs several whole paragraphs? That may be more troubling.
AI detectors, just like that rumor that the em dash is an AI tell, can have false positives. A tool that is still largely a black box will make mistakes, and that can be devastating for someone whose genuine writing was flagged through no fault of their own.
I asked Yamin how human writers can make sure their work isnât caught in that trap. âJust do your thing,â he said. âMake sure you have your human touch.â

