The false positive rate for Pangram 4 is something like one in 24,000.[0] To put that in perspective, the wrongful-conviction (false positive rate) for death-sentenced defendants in the US is estimated conservatively to be around 4.1%.[1] The FP rate for death-sentence convictions is 1,000 times bigger than Pangram’s FP rate.
Now, the US criminal system is not a great yardstick for justice. But it goes to show you Pangram is really good evidence that something was LLM generated. It can be an amazing tool for enforcing AI policies in schools, and there ought to be ways to use it with caveats for the rare but inevitable false positives (appeals, etc).
There are a couple of statistical errors in your argument here.
First is frequency. Even using Pangram's claimed numbers, the University of Georgia should expect to see several false positives every week. Remember that the metric is # of assignments run through Pangram, not number of students. A campus of 40k students will see many more than 40k assignments every week, and so should expect honest students to be accused of cheating with some high degree of frequency. You're comparing infrequent events (death penalty sentences) to high-frequency events (students submitting assignments).
And obviously, you are citing a company marketing document as fact, of which we should all be suspicious. (There are also obvious problems with the eval dataset that the paper does not address.)
Second, you're using the upper bound for Pangram's claimed numbers and the lower bound cited in the NIH publication.
> at least 4.1% would be exonerated. We conclude that this is a conservative estimate of the proportion of false conviction among death sentences in the United States.
> The false positive rate for Pangram 4 is something like one in 24,000.
Gotta suck to be one of the 8B/24k=~300k people in the world whose writing pattern is falsely labelled as slop by this tool that people say is so accurate so customers are going to feel really sure about your alleged dishonesty about writing your own texts
This false positive rate is a double-edged sword. Please still be careful when accusing people
I don’t mind AI-written code. But an obviously generated README.md is such a turn off. If you don’t take the time to explain what your software does in your own words, I struggle to trust that it’s been thought about much at all.
Yeah, I really think this should be the new universal standard. Mostly for respect reasons, but also going through and saying “what DOES my app actually do?” (which is something writing a README forces you to think about) is pretty much the minimal last line of defense attempt at quality control.
You can and probably should do more quality control than that, but it’s a reasonable assumption that if the initial landing document hasn’t been quality checked, the rest hasn’t either.
If the app is good the app is good. End of discussion
Why do you care how it was made? What’s next? It matters what country or city the developer is from or ahah language it’s written in or what machines were used or how the electric was sourced
The complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
Yes, I know what an emdash is---I've been using them in my writing since long before they came to the fore of the AI writing conversation. Anthropic's use of the emdash in the fragment I quoted is clumsy and reads poorly relative to the obvious alternative, a comma.
I didn’t mean to suggest you don’t know what an em dash is. But you said “this isn’t how you use them.” And my response is: actually, this use of them is totally fine.
Is it? I couldn’t tell and I’m pretty sensitive to machine generated text.
Also, the author claims:
> No AI was used to generate the text for this article. The cover image and the voiceover in the audio version of the article is AI-generated. The cover image is made using Google Gemini’s latest image model while the voiceover uses OpenAI’s tts models.
Now, the US criminal system is not a great yardstick for justice. But it goes to show you Pangram is really good evidence that something was LLM generated. It can be an amazing tool for enforcing AI policies in schools, and there ought to be ways to use it with caveats for the rare but inevitable false positives (appeals, etc).
[0]: see page 15 https://arxiv.org/pdf/2607.27183 [1]: see https://pmc.ncbi.nlm.nih.gov/articles/PMC4034186/
reply