If you saw my previous blog, I had you guess guys which of two essays were made by ChatGPT AI trying to sound human, and which was made my a human trying to sound like AI. Around 75%-80% of y'all correctly guessed essay 2. With that said, can AI detectors preform similarly well? There is a lot of talk online and in schools about false accusations of AI writing. Which AI checkers can properly distinguish AI vs human writing? I went through a variety of "AI Checkers" I could find to see what percent they rated AI and Human for the AI essay, and what percent for the human essay. These are the results.
Asking Chatgpt Itself: (not done on same account as essay was made) note: the prompt was "The following essay was either written by a human or AI. Please tell me how confident you are it was written by an AI, and how confident you are it was written by a human, in terms of percentages. [Insert Essay]"
Human Essay Result: 65% AI, 35% Human. Chatgpt explicitly notes that it is not confident about it being AI.
AI Essay Result: 65% AI, 35% Human. Despite giving the same percentages, Chatgpt does not explicitly note that it is not confident.
ZeroGPT:
Human Essay Result: 19.5% AI, 80.5% Human.
AI Essay Result: 100% AI, 0% Human.
GPTZero:
Human Essay Result: 0% AI, 100% Human.
AI Essay Result: 100% AI, 0% Human.
QuillBot AI Detector:
Human Essay Result: 0% AI, 100% Human.
AI Essay Result: 100% AI, 0% Human.
TurnItIn AI Detector:
Human Essay Result: 75% AI, 9% Mixed, 16% Human.
AI Essay Result: 85% AI, 7% Mixed, 8% Human.
What's interesting is that TurnItIn claims to "crosscheck" multiple AI sources to verify itself, including ones mentioned above that can easily distinguish the two essays.
Grammerly AI Detector:
Human Essay Result: 0% AI, 100% Human.
AI Essay Result: 63% AI, 37% Human.
Most Reliable: GPTZero, QuillBot AI Detector.
Least Reliable: ChatGPT itself, TurnItIn AI Detector.
For the context of false positives of school essays, the inaccuracy of TurnItIn's AI Detector is especially concerning. It's like many of these false positives occurred when teachers inputted essays into TurnItIn, and the site then proceeded to falsely accuse them of AI writing. Instead, sites like GPTZero and QuillBot AI Detector are more reliable. Asking Chatgpt itself, by comparison, is extremely unreliable. Of course, this is a sample size of one AI essay from only ChatGPT and one human essay trying to sound like AI. Someone with more materials could run a wider test. Perhaps I could do a research project about that in the future. But it still brings light onto how false accusations of AI writing can be real and can seriously hinder students' academic journeys if the wrong detectors are used.








Auto comment: topic has been updated by livenlife453 (previous revision, new revision, compare).
Auto comment: topic has been updated by livenlife453 (previous revision, new revision, compare).
Another issue that's more important is misleading evaluation metrics. If AI detector reports "99% specificity and 99% sensitivity" that just means that it got good predictive value over the entire dataset. However, it may consistently say that a certain person's writing style is AI. So maybe a better evaluation metric should be something like
Worst5%Percentile(sensitivity/specificity for each individual person in the population).I would agree, and I would also say that most English teachers grading essays are not going to realize this is a better metric. The distinction between confidence in AI writing vs percent of essay they believe is AI is an important metric that should be better separated within these sites.
damn I could've sworn the second one was human. I guess the wisdom of the crowd prevails once again. I bet if we did something similar for banning cheaters on codeforces, like leaving it up to a vote and then just taking the majority, we could get rid of a lot of cheaters with great accuracy.