Y'all Correctly Guessed Which Essay Was Human and Which Was AI, but how do AI Checkers Fair?

Revision en2, by livenlife453, 2026-08-09 00:34:17

If you saw my previous blog, I had you guess guys which of two essays were made by ChatGPT AI trying to sound human, and which was made my a human trying to sound like AI. Around 75%-80% of y'all correctly guessed essay 2. With that said, can AI detectors preform similarly well? There is a lot of talk online and in schools about false accusations of AI writing. Which AI checkers can properly distinguish AI vs human writing? I went through a variety of "AI Checkers" I could find to see what percent they rated AI and Human for the AI essay, and what percent for the human essay. These are the results.

Asking Chatgpt Itself: (not done on same account as essay was made) note: the prompt was "The following essay was either written by a human or AI. Please tell me how confident you are it was written by an AI, and how confident you are it was written by a human, in terms of percentages. [Insert Essay]"

Human Essay Result: 65% AI, 35% Human. Chatgpt explicitly notes that it is not confident about it being AI.

AI Essay Result: 65% AI, 35% Human. Despite giving the same percentages, Chatgpt does not explicitly note that it is not confident.

ZeroGPT: Human Essay Result: 19.5% AI, 80.5% Human.

AI Essay Result: 100% AI, 0% Human.

GPTZero:

Human Essay Result: 0% AI, 100% Human.

AI Essay Result: 100% AI, 0% Human.

QuillBot AI Detector:

Human Essay Result: 0% AI, 100% Human.

AI Essay Result: 100% AI, 0% Human.

TurnItIn AI Detector:

Human Essay Result: 75% AI, 9% Mixed, 16% Human.

AI Essay Result: 85% AI, 7% Mixed, 8% Human.

What's interesting is that TurnItIn claims to "crosscheck" multiple AI sources to verify itself, including ones mentioned above that can easily distinguish the two essays.

Grammerly AI Detector:

Human Essay Result: 0% AI, 100% Human.

AI Essay Result: 63% AI, 37% Human.

Most Reliable: GPTZero, QuillBot AI Detector.

Least Reliable: ChatGPT itself, TurnItIn AI Detector.

For the context of false positives of school essays, the inaccuracy of TurnItIn's AI Detector is especially concerning. It's like many of these false positives occurred when teachers inputted essays into TurnItIn, and the site then proceeded to falsely accuse them of AI writing. Instead, sites like GPTZero and QuillBot AI Detector are more reliable. Asking Chatgpt itself, by comparison, is extremely unreliable. Of course, this is a sample size of one AI essay from only ChatGPT and one human essay trying to sound like AI. Someone with more materials could run a wider test. Perhaps I could do a research project about that in the future. But it still brings light onto how false accusations of AI writing can be real and can seriously hinder students' academic journeys if the wrong detectors are used.

Tags artificial intelligence, humanity, writing, #false positive

History

 
 
 
 
Revisions
 
 
  Rev. Lang. By When Δ Comment
en3 English livenlife453 2026-08-09 00:34:42 2 Tiny change: 'roGPT:**\nHuman Es' -> 'roGPT:**\n\nHuman Es'
en2 English livenlife453 2026-08-09 00:34:17 30
en1 English livenlife453 2026-08-09 00:32:17 2908 Initial revision (published)