If you saw my previous blog, I had you guess guys which of two essays were made by ChatGPT AI trying to sound human, and which was made my a human trying to sound like AI. Around 75%-80% of y'all correctly guessed essay 2. With that said, can AI detectors preform similarly well? There is a lot of talk online and in schools about false accusations of AI writing. Which AI checkers can properly distinguish AI vs human writing? I went through a variety of "AI Checkers" I could find to see what percent they rated AI and Human for the AI essay, and what percent for the human essay. These are the results.
Asking Chatgpt Itself: (not done on same account as essay was made) note: the prompt was "The following essay was either written by a human or AI. Please tell me how confident you are it was written by an AI, and how confident you are it was written by a human, in terms of percentages. [Insert Essay]" Human Essay Result: 65% AI, 35% Human. Chatgpt explicitly notes that it is not confident about it being AI. AI Essay Result: 65% AI, 35% Human. Despite giving the same percentages, Chatgpt does not explicitly note that it is not confident.
ZeroGPT: Human Essay Result: 19.5% AI, 80.5% Human AI Essay Result: 100% AI, 0% Human.
GPTZero: Human Essay Result: 0% AI, 100% Human AI Essay Result: 100% AI, 0% Human.
QuillBot AI Detector: Human Essay Result: 0% AI, 100% Human AI Essay Result: 100% AI, 0% Human
TurnItIn AI Detector: Human Essay Result: 75% AI, 9% Mixed, 16% Human. AI Essay Result: 85% AI, 7% Mixed, 8% Human. What's interesting is that TurnItIn claims to "crosscheck" multiple AI sources to verify itself, including ones mentioned above that can easily distinguish the two essays.
Grammerly AI Detector: Human Essay Result: 0% AI, 100% Human. AI Essay Result: 63% AI, 37% Human.
Most Reliable: GPTZero, QuillBot AI Detector. Least Reliable: ChatGPT itself, TurnItIn AI Detector.
For the context of false positives of school essays, the inaccuracy of TurnItIn's AI Detector is especially concerning. It's like many of these false positives occurred when teachers inputted essays into TurnItIn, and the site then proceeded to falsely accuse them of AI writing. Instead, sites like GPTZero and QuillBot AI Detector are more reliable. Asking Chatgpt itself, by comparison, is extremely unreliable. Of course, this is a sample size of one AI essay from only ChatGPT and one human essay trying to sound like AI. Someone with more materials could run a wider test. Perhaps I could do a research project about that in the future. But it still brings light onto how false accusations of AI writing can be real and can seriously hinder students' academic journeys if the wrong detectors are used.



