Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Claude 3.7 Sonnet Thinking scores 33.5 (4th place after o1, o3-mini, and DeepSeek R1) on my Extended NYT Connections benchmark. Claude 3.7 Sonnet scores 18.9. I'll run my other benchmarks in the upcoming days.

https://github.com/lechmazur/nyt-connections/



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: