I did everything by hand for so long and miss it, now I just read thousands of lines day in day out. But there are interesting gains such as lowered barrier to quickly run adhoc perf tests, do some formal verification, you can just add more guardrails for the thing that gets shipped. The thing I hate though is when managers sell their generated code and you are left shoehorning it to production, they take away the interesting part of the job.
Misleading article, they are selling to corporations already who in turn have plans to create a few fabs in Europe. So it is maybe more about European companies who are not buying, but that rests on difficult capital availability and fragmented nature of the continent. Europe is not one country after all, someone wants a fab in Germany, but the why is it not in France or Sweeden.
> we found no immediate increase in key engineering metrics such as monthly pull requests and lines of code
> To establish a before-Copilot baseline, we used data from April, May, and June 2024. After-Copilot data was represented by the period of September, October, and November 2024
> GitHub Copilot usage varied significantly among engineers, the tool demonstrably fostered positive changes in perceived engineer value, reduced time spent on various engineering activities, and boosted motivation and perceived skills
> subsequent monitoring of PRs and LOC for participants from December 2024 to May 2025 showed no statistical improvements
> the implications of more advanced capabilities, such as retrieval-augmented generation (RAG) over enterprise codebases or deeper engineering workflow integrations, warrant separate investigation
I think it is just out of date, habits have changed as well. I have not seen much gain personally at that period except in the last 12 months. Also, models not named, token counts not shown. Not to mention it was the older autocomplete + chat that were in use, these days it is much more advanced with RAG, cli use, MCPs, etc.
Despite the article looks like it talks about Claude, in reality it describes a new type of a job - a synthesis of data science, research, comp science, plus industry specific knowledge. Another extremely important thing is to have access to all related research in some programmatic way, this is for exploration, I do not think many have such access. Finally, you need to be prepared to read all those generated results and judge them effectively to pick the strands worth pursuing further. I bet you could do it with any model and your own harness, even authors admit they use their own to manage multiple sessions which hints that claude is not enough.
Just for the context UBI kind of already exists in some countries but in a different form. If you are in Ireland and do not work you get the benefits provided you have somewhere to live and can pick them up in the post office (~200eur/week). The problem is that it solves only part of the issue because people have to live somewhere, and sure the gov will even contribute towards it (they do already) but somehow not enough is being built. The benefit is basically not enough to get the mortgage, nor is there housing available.
Not vr/ar but 3d visualisations and walkthroughs were a thing like 20 years ago. I did try to leverage it with the clients but what happened was that once people could see everything they had more opinions about how to change it. It helped me to get clients though. Sometimes it would go on for quite a ”few” iterations. I think it is more scalable not to do it, i.e. not to be on the cutting edge, unless the customer pays premium for it. I am pretty sure vr/ar has the same challenges where people are like “oh I thought the ceiling was higher, can we increase it by 2 inches and move the stair case a bit?” then you do it and something is wrong again.
I'd see that as evidence that the visualization is doing its job. It's much cheaper to discover "I wish the ceiling were 2" higher" before anyone starts building than after. The challenge is reducing the cost of each iteration. That's where newer tooling starts to matter.
It looks like the message here is “make sure to use $100k worth of Claude when doing any analysis or evaluation” and the given examples show that prior effort could have been improved or made faster. But to me 100k is an opportunity cost, and there is a possibility that these results are not reproducible, so spending it on some researcher or a grad student would buy you more in a long term. If it was 1k then sure it is worth throwing at a large problem space to find things, like using fuzzing.
Went through the comments here and there and one thing to note is that there was a question about who do you think should have won instead. This is a good question because it is possible that all submissions were like this or there were ones that looked just worse. It would be quite useful to know who came close as well in this case. If you knew which submissions were good you could have a process to revoke the prize and give it to someone else in case of fraud or negligence or similar.
Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.
Assuming this description is accurate, if they were all like this then none of them should have won. They should have all been disqualified and the organizers should have looked at themselves in a mirror for a long time.
I think the underlying problem here is that no single human brain has enough glycogen in reserve to thoughtfully process all the AI slop. It simply cannot be done by mortals.
I've noticed this over and over again with "professionals actually prefer LLM responses" studies. Typically the human generated responses seem better to me on a quick sample, but if I had to review 50 of them I'd probably start taking lazy shortcuts; using superficial language aptitude or factual comprehensiveness instead of critically reading.
It does seem like the human judges here might have given credit for e.g. a 20pg arXiv paper without actually reading it. I can blame them professionally but emotionally I have nothing but sympathy. I truly hate LLMs.
reply