Hacker Newsnew | past | comments | ask | show | jobs | submit | droidjj's commentslogin

The false positive rate for Pangram 4 is something like one in 24,000.[0] To put that in perspective, the wrongful-conviction (false positive rate) for death-sentenced defendants in the US is estimated conservatively to be around 4.1%.[1] The FP rate for death-sentence convictions is 1,000 times bigger than Pangram’s FP rate.

Now, the US criminal system is not a great yardstick for justice. But it goes to show you Pangram is really good evidence that something was LLM generated. It can be an amazing tool for enforcing AI policies in schools, and there ought to be ways to use it with caveats for the rare but inevitable false positives (appeals, etc).

[0]: see page 15 https://arxiv.org/pdf/2607.27183 [1]: see https://pmc.ncbi.nlm.nih.gov/articles/PMC4034186/


There are a couple of statistical errors in your argument here.

First is frequency. Even using Pangram's claimed numbers, the University of Georgia should expect to see several false positives every week. Remember that the metric is # of assignments run through Pangram, not number of students. A campus of 40k students will see many more than 40k assignments every week, and so should expect honest students to be accused of cheating with some high degree of frequency. You're comparing infrequent events (death penalty sentences) to high-frequency events (students submitting assignments).

And obviously, you are citing a company marketing document as fact, of which we should all be suspicious. (There are also obvious problems with the eval dataset that the paper does not address.)

Second, you're using the upper bound for Pangram's claimed numbers and the lower bound cited in the NIH publication.

> at least 4.1% would be exonerated. We conclude that this is a conservative estimate of the proportion of false conviction among death sentences in the United States.

The one commonality is that


> The false positive rate for Pangram 4 is something like one in 24,000.

First of all: says them.. Do you think they might have an incentive to boost their numbers?

Second, this is on existing text, which might be in their training data, no?


> The false positive rate for Pangram 4 is something like one in 24,000.

Gotta suck to be one of the 8B/24k=~300k people in the world whose writing pattern is falsely labelled as slop by this tool that people say is so accurate so customers are going to feel really sure about your alleged dishonesty about writing your own texts

This false positive rate is a double-edged sword. Please still be careful when accusing people


I don’t mind AI-written code. But an obviously generated README.md is such a turn off. If you don’t take the time to explain what your software does in your own words, I struggle to trust that it’s been thought about much at all.

I didnt even need to open the link, a ghibli-generated profile pic already says so much about someone's priorities.

So what

Yeah, I really think this should be the new universal standard. Mostly for respect reasons, but also going through and saying “what DOES my app actually do?” (which is something writing a README forces you to think about) is pretty much the minimal last line of defense attempt at quality control.

You can and probably should do more quality control than that, but it’s a reasonable assumption that if the initial landing document hasn’t been quality checked, the rest hasn’t either.


If the app is good the app is good. End of discussion

Why do you care how it was made? What’s next? It matters what country or city the developer is from or ahah language it’s written in or what machines were used or how the electric was sourced

Who cares!


A README is one of the most important pieces of a project, if one can't even advertise their own project with their own words, it signals low effort.

After all, the people visiting one's repo on GH likely have access to the same AI tools, too.


Linus Tech Tips dropped a video on this an hour ago, so the prices are already skyrocketing sadly.

Link: https://youtu.be/lmLh29LE2W4?is=QdRwtQS3xyyPt44E


Yuck... They increased a bit ($150+?) after the GPU unlock came out, but they're nuts now.

The price is probably coming from when Luna was first released. OpenAI slashed the price by 80% at the end of July.

Yes! Good catch, thanks - I'll fix that.

I had the prices wrong on Sol and Terra as well - they've all had price drops:

https://openai.com/index/advancing-the-price-performance-fro...

Sol discount is until November 21, 2026 according to https://developers.openai.com/api/docs/changelog

  Luna — costs in cents
  +--------+--------+---------+
  | Effort | Before | After   |
  +--------+--------+---------+
  | max    |   7.83 |    1.57 |
  | xhigh  |   4.24 |    0.85 |
  | high   |   2.46 |    0.49 |
  | medium |   1.26 |    0.25 |
  | low    |   0.76 |    0.15 |
  | none   |   0.71 |    0.14 |
  +--------+--------+---------+
  Per million tokens:
  Before: $1 input / $6 output
  After:  $0.20 input / $1.20 output

  Sol — costs in cents
  +--------+--------+---------+
  | Effort | Before | After   |
  +--------+--------+---------+
  | max    |  48.55 |   32.37 |
  | xhigh  |  24.11 |   16.08 |
  | high   |  10.38 |    6.92 |
  | medium |  10.55 |    7.03 |
  | low    |   8.33 |    5.55 |
  | none   |   5.90 |    3.93 |
  +--------+--------+---------+
  Per million tokens:
  Before: $5 input / $30 output
  After:  $4 input / $20 output

  Terra — costs in cents
  +--------+--------+---------+
  | Effort | Before | After   |
  +--------+--------+---------+
  | max    |  32.09 |   25.67 |
  | xhigh  |  14.67 |   11.74 |
  | high   |   3.74 |    2.99 |
  | medium |   3.46 |    2.77 |
  | low    |   3.47 |    2.78 |
  | none   |   2.60 |    2.08 |
  +--------+--------+---------+
  Per million tokens:
  Before: $2.50 input / $15 output
  After:  $2 input / $12 output

The complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.

When AGI comes it will come as a pelican and gobble up all these troublesome little fishies who gripe and whine and moan about pelicans.

I'm always tired of seeing at the top of every new model release post on here. I say Simon should just keep it to Twitter.

I have never seen a Pelican on X which used to be called twitter in about 1872. Keep up!

Em dashes are commonly used to add emphasis, even where you would ordinarily use a comma. Their flexibility is why many people love them! See https://www.merriam-webster.com/grammar/em-dash-en-dash-how-...

Yes, I know what an emdash is---I've been using them in my writing since long before they came to the fore of the AI writing conversation. Anthropic's use of the emdash in the fragment I quoted is clumsy and reads poorly relative to the obvious alternative, a comma.

I didn’t mean to suggest you don’t know what an em dash is. But you said “this isn’t how you use them.” And my response is: actually, this use of them is totally fine.

Well, I disagree! https://news.ycombinator.com/item?id=49527581

It's fine in the sense that when a bad writer writes something I can usually understand what they're trying to say.


It's also certainly not 2x-4x cheaper than getting GLM 5.3 inference elsewhere.

Compare: https://www.coralbricks.ai/docs#models vs https://openrouter.ai/z-ai/glm-5.3#providers


Is "Salem" a tongue-in-cheek nod to being a small, quirky competitor to Boston Dynamics (Salem, MA : Boston, MA :: Salem Robotics : Boston Dynamics)?

If so, as someone who lives close to Salem, I like it :)

Good luck to you guys.


Haha it wasn't, but we might steal that for our story going forward.


is this based in Salem MA? If so, nice


Cofounder here. we're based in Austin, TX, but yes, the name alludes to Salem MA. We're bringing inanimate objects to life through robotics, haha.


First thing I thought of too, I love it


Salem Statics


Boston Dynamics vs Salem Statics, haha. Seriously though, we are very complimentary to BD, since Spot is one of the platforms we deploy on.


Is it? I couldn’t tell and I’m pretty sensitive to machine generated text.

Also, the author claims:

> No AI was used to generate the text for this article. The cover image and the voiceover in the audio version of the article is AI-generated. The cover image is made using Google Gemini’s latest image model while the voiceover uses OpenAI’s tts models.


> Also, the author claims:

You can slap an arbitrary disclaimer on anything; it doesn’t mean it’s genuine.


Very true, I just want to trust people.


I would like, personally, if the image was a stock photo. It being AI feels in poor taste.


Pangram reports "100% AI" with "high" confidence...


I think this is a GitHub issue (of course it is) as I'm not able to search for anything atm.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: