The best way to access ChatGPT "chat" was always on ChatGPT.com via a web browser, not even the iOS app seems to have all the little things you can do on the web version.
The Mac desktop ChatGPT app had weird issues like not being able to see the model being used by a chat inside a project etc. I just ended up installing ChatGPT Atlas and using that as an AI-only browser for all LLMs including Claude and Grok lol
Wildflower Health | Junior Software Engineer | Remote (US) | Full Time
Wildflower provides a modular suite of software, support, and services, partnering with healthcare providers and payers at any stage of their journey to value. We address critical gaps by breaking down organizational silos and seamlessly integrating with existing clinical and operational workflows. Our approach focuses on the whole person, empowering payers and clinicians to efficiently identify patient needs and coordinate essential resources that bridge gaps in maternity care.
Tech Stack: Kubernetes, Node, Mongo (Backend) React Native Mobile App (Front End) Leveraging cutting edge AI tools for development.
Respectfully disagree that it's toxic or predatory. Very easy to cancel via chat as well. New York times is an incredible jewel of a company. I am fine with some marketing tactics that aren't incredibly heavy handed. They are far far far from unethical.
I've seen a lot of this sentiment over the previous six months from people on reddit. I have yet to experience this myself as a developer with over 20 years of experience.
As always, I think this happen more to vibe coder. They don't understand that bigger project means worse AI performance. On top of that Opus felt being nerfed at understanding prompt so if your spec is bad you won't get good result.
What it does seem like is that they're tuning some knobs up and down or releasing new versions of models or system prompts that result in the model getting dumber and smarter in waves.
Opus has been dumb this week.
Claude was having a lot of capacity problems and downtime and then this week that has been much less obvious... and the model is dumber.
It could also just be luck and my impressions are false... who knows.
It’s because it’s not true, there’s no evidence for it that passes the sniff test. No lab is “shipping a worse model once they’ve got you”. People have a bad few days and blame the model providers instead of stepping back to fix their workflow.
Opus 4.7 has been a real downgrade for me. I’m back to mid 2025 when I had to catch all the completely intermediary goals/assumptions the model is creating for itself
My "mental scratchpad" needs to be as sharp as possible to maximize my intelligence. I think of the LLM as a scratchpad for my thinking, I hope the Anthropic team can see this.
it's sort of good at thinking, writing specs, etc.. Also debugging. But as a coder: I see no advantage to opus 4.6 and I preferred sonnet most times already over opus 4.6.
I see a lot of the "4.7 is a downgrade" sentiment. 4.7 does (mostly) what you ask it to do. 4.6 does what it thinks it should do. As someone with 20 years writing my own code I want the former, but the loud contingent online wants the latter.
When you're on a mature codebase with 500k+ lines of code, I haven't seen anything else be as effective as 4.7.
I can tell you for a fact, Claude 4.7 was NOT doing what I told it to do (in fact the clear and complete opposite - repeatedly), a pretty simple architectural refactor, and that Codex did better and DeepSeek much better.
It was given very simple ways to verify success. It simply didn't do that and said it's at a good stopping point, despite moving in the WRONG direction not even doing 1% of the task, and being told to see the task through to completion.
Meanwhile, Codex broke it down into 3 steps and just got it done...
No, "I'm going to give it to you straight, this is a large risky commit that could go sideways, so I'm just not going to do anything instead."
Claude worked on it for almost 200 commits over 2 weeks, needing to typically prompt it 3x to even TRY to make any progress instead of just wasting tokens to ignore me and tell me how big and risky it is.
Maybe Claude is just particularly terrible at this type of refactor. I'm not sure why that would be.
Oh Opus is nerfed sure, but not that hard. Early this year opus 4.6 can understand your prompt and your intention easily, it got worse around mid April. Opus 4.7 even worse than that.
However that's just it, you just need to improve and make clearer of your prompt and it will perform just as good.
This account is an LLM-hype peddler, shilling for Anthropic (check comment history). If they say that Claude is not nerfed, then most likely it is, in fact, nerfed.
I wouldn't call correcting misinformation and FUD "peddling hype" or "shilling" but I suppose we are in a post-truth world, where if you push back against the anti-AI emotions and vibes with grounded facts, you must be a shill.
Anyways, please take your discourse of calling people you disagree with "shills" back to Reddit. I'd much rather engage with someone debating the merits of an argument.
If you are an LLM-hype peddler, you really should not be offended at being called out. Also, this is the merit you are ostensibly looking for — since you are a shill, everyone should know this first before taking your words seriously.
You should also check your LLM prompt for HN comments, because the original comment you replied to was not anti-AI, and, in fact, very much pro-AI. The only criticism it had was about model being degraded, so they could not go as hard at AI-assisted development anymore as they used to before. I guess it's a bit difficult for LLMs to spot the difference and make proper conclusion for now.
Also even if taking you seriously — how does writing "no, model performance is not degraded because I say so" serve as correcting misinformation? It only does if you are shilling for Anthropic (which you do), otherwise it's just hot air.
Not offended at all, but just ranting about how someone is a shill instead of responding to the substance of their argument is simply not the kind of discussion we have on HN. Read the guidelines.
> "no, model performance is not degraded because I say so" serve as correcting misinformation?
Because zero evidence has been provided other than feelings. That is not evidence of degradation, and we know they don't serve quants.
You are an Anthropic shill, and this is an explicit marker that needs to be added to all of your comments, so that all information you provide can be adjusted for that bias. But I do understand why you ignore this point since it devalues all your comments (as it should), and instead cling to "ranting how someone is a shill bla-bla-bla".
Those people, unlike you, are actually using AI in development. And it is not a singular person who reports their frustration with the model being degraded after a certain period of time, so the anecdata does gradually become data. Your attempts at gaslighting are weak, you should really ask your bosses for a new guidebook on how to deal with reports of models performing at worse levels than before. Just writing "because I say so" is not cutting it.
> "we know they don't serve quants"
How do you know that unless you are working at Antrhopic? Yet another evidence of you being an Anthropic shill.
You have no substantive arguments other than calling people you disagree with shills.
> so the anecdata does gradually become data.
No, it does not. Countless social phenomena demonstrate how factually incorrect misconceptions spread rapidly. Frequency illusion is real and contagious.
> How do you know that [they are not serving quants]
Lots of ways to tell, if you weren't busy calling people shills.
First, Anthropic and OpenAI have both stated they don't serve quants. Weak protection, but it's there.
Second, no one has shown an A/B or eval proving a regression.
Third, and most importantly, the actual output measurably changes. Quants have a lower latency, higher TPS, and different token distribution. Despite having access to this data, no one has any evidence proving a quant has been served.
> You are an Anthropic shill
I'd explain the reasons I favor Anthropic over the others, but you'd just go back to yelling "shill" instead of engaging in a real conversation. That said, I am a fan of GDM as well, and think Gemini is better than Anthropic for everything other than code.
I've seen nothing resembling sane, reasoned thought from you in this thread. Just anger.
You haven't substantively debated a single point, it's like "shill" is the only word in your vocabulary. Again, this isn't Reddit.
Nothing to do with disagreement, I only call "Anthropic shills" people who are explicitly and shamelessly shilling for Anthropic. You still ignore the point that shilling adds bias to all your comments, so other readers have to actively keep it in mind to adjust for it. Stating that you are an Anthropic shill helps everyone around. And somehow you managed to be peddling LLM-hype shit so hard, that you are the only one called out on that by me.
> No, it does not.
Yes, it does, it is literally the definition of data - collection of points, observations, anything really. Try gaslighting harder, Anthropic shill. As I said, ask for better playbook on how to deal with people actually experiencing degradation before replying again.
> First, Anthropic and OpenAI have both stated they don't serve quants.
What's the point of stating this other than trying to pad your baseless "proof"? LLM-level argument.
> Second, there have not been evals showing a real regression test proving that a quant was served
This is how I know you have no idea what you are talking about and resort to LLMs for all your argumentation. Benchmarks are gamed so hard that even quantized models would achieve on them non-quantized level reliably. Moreover, benchmarks (that matter) are not run continuously all the time.
> Third, and most importantly, the actual output measurably changes. Quants have a lower latency, higher TPS, and different token distribution. Despite having access to this data, no one has any evidence proving a quant has been served.
You really are an LLM. What do you think different token distribution means? It literally means different, arguably worse performance in coding tasks. The evidence is in your face, but you have to keep it straight, since you are an Anthropic shill. You wrote yourself an argument why the models ARE quantized over time and did not even understand it. Makes sense, since you are paid to not understand stuff but peddle LLM-hype for Anthropic instead.
> I'd explain the reasons I favor Anthropic over the others
It is perfectly visible why you favor Anthropic, because you are an Anthropic shill and they pay you your salary, duh.
> real conversation
This is the type of conversation everyone should have whenever they read something written by an Antrhopic shill. You are actively poisoning this forum by astroturfing for Antrhopic, so we should take measures against it.
> You haven't substantively debated a single point
Obviously an Anthtropic shill would ignore everything of substance I wrote and instead focus on being called out. Fortunately, it is not you who I have to convince of anything, since your very well-being relies on getting salary from Anthropic peddling LLM-hype on HN and elsewhere, so you are physically incapable of understanding pretty much anything that contradicts your talking points.
> Yes, it does, it is literally the definition of data
No, feelings are not reliable data when frequency bias and misinformation exist. There is a reason most experiments isolate out bias as much as possible.
> Moreover, benchmarks (that matter) are not run continuously all the time.
So there's no data?
> What do you think different token distribution means?
You clearly did not understand anything I said. Stated simply: If you were being served a quant, you'd be able to tell by looking at the token distribution, latency, and TPS. You don't need to trust the labs' word for it.
> they pay you your salary, duh.
In fact, I get paid by a FAANG, though I do use Anthropic products heavily. Further, I don't really need money, I have more than enough. So much for reading my history.
> You are actively poisoning this forum
Your degenerate discussion - calling people shills instead of engaging with the argument, insulting them when your arguments are disproven, your inability to hold a rational debate that's not angry and emotionally charged - that is what is poisoning this forum.
Frankly, if you react this angrily and emotionally to a simple rational premise (that frequency bias leads to the perception of models being worse than them actually being worse), you're ngmi unless you're already independently wealthy.
I would recommend a therapist, it helped me when I had similar behavioral issues. (Claude is a great therapist, by the way ;)
Nice gaslighting, Anthopic shill. No one said a word about feelings, only you (to derail the conversation). People reported their own experience and frustration with the model being unable to complete tasks they previously could. I said, get a better playbook before coming back. Or is it the best LLMs can do for now? Sad, then.
> No data
There is data, which you try to gaslight into being "feelings", Anthopic shill.
> Stated simply: If you were being served a quant, you'd be able to tell by looking at the token distribution, latency, and TPS.
Did you just repeat what you said before while ignoring the actual meaning of the words and my explanation of what YOU wrote? Is it what LLM told you to do, Anthropic shill? And you claim I have no substance. Maybe spend a week or so getting educated, before blindly copying and pasting LLM output, Anthropic shill?
> I get paid by a FAANG
Yeah, in your dreams maybe, Anthropic shill. I did read your comment history, and this is likely part of the story you try to build around your Anthropic shilling persona. Not a single fact that would prove that and believe me, I tried looking for it. Only endless claims of "I work at a FAANG" (no one who actually works here writes it like this).
> I use Anthropic products heavily
This is obvious, as 90% of your comments are LLM generated, Anthropic shill.
> calling people shills
Clanker, I called only you a shill, not people, tell your LLM to update its context. And I called you shill not because of any arguments, but because of your comment history unapologetically shilling for Anthropic and peddling LLM hype.
> arguments are disproven
You ignored half of my arguments, and for the rest you just repeated what you wrote before, not even understanding what the words you typed meant. Nice gaslighting, Anthropic shill.
> insulting
And you said you were not offended. Once again, Anthropic shill, being called a shill is not an insult. This is your fate, to be called an Anthropic shill, while you are on their payroll, astroturfing online communities with your LLM-bullshit peddling. Or do you expect being a propagandist to be a pleasant experience? People with no morals like you coming into this forum spreading their employer's bullshit deserve all the hate they get and more.
Your LLM outputs the same thing as in other comment for no good reason. Can't Anthropic afford good models for its shills, or is it the best SOTA can do now?
I would recommend you abandon this account, because it's now burned for all shilling intents and purposes.
> Again, you're just interpreting anything that goes against the "AI bad" grain as shilling.
Once again, putting "AI bad" into my words. No, Anthropic shill, this is not what I am saying. Is your LLM malfunctioning or are you not really getting it? Stop gaslighting, Anthropic shill, and try to stick to the actual words I am saying. I understand that this is hard for you, because then you would have no real argument to make, but please try, Anthropic shill.
> Please show it.
I used an LLM to count actual experiences of people reporting their experience with Opus 4.6 being degraded. There are literally several hundreds of such data points. This is data. People, who are employed and actually use LLMs for coding, unlike you, Anthropic shill, who uses it only to poison online communities. Are you really going to disregard all that to claim it is mass-psychosis or something? I guess you would, Anthropic shill, because that's your job, to peddle bullshit LLM-hype unbased on anything in reality.
> It was an incoherent mess of insults, so I am still not sure what you're trying to say.
Repeat after me, Anthropic shill: being called a shill is not an insult. You are a shill, stop being obtuse and at least take some pride in your work of promoting LLM-hype.
So once again you are providing nothing to the conversation except for baseless accusations of insulting, which I did not do, and refuse to answer to the actual arguments I made. I can provide it again, but you would likely ignore it because it just showcases how you are clueless about the topic.
Your words, not mine:
> Third, and most importantly, the actual output measurably changes. Quants have a lower latency, higher TPS, and different token distribution.
I asked if you understood what "different token distribution" meant. I can tell you what it means: models performing worse at coding tasks. So people report models being worse at coding tasks, YOU write that indeed quantization leads to that and then just "forgot" about it? Nice level of "objective" discussion, Anthropic shill.
> So now I'm lying about my employment on an anonymous forum for... what, exactly?
It is not anonymous forum, as much as you would have liked it to be, so that your shilling could not be dismissed as easily, Anthropic shill. For what? So that people would fall for the bullshit you are peddling. Are you really this dense, Anthropic shill?
I must admit I skimmed most of your comment because it is largely an incoherent rant, but I will address some points:
> This is data.
Nope. Because frequency bias is a thing. If you hear on Twitter "model X got nerfed," your brain will look for that pattern and notice it more than usual. This will then confirm your suspicion, which leads to a vicious cycle. Then you tell your friends and the same phenomenon repeats.
None of this requires the model to get worse. It's a well understood psychological phenomenon.
> I can tell you what it means: models performing worse at coding tasks. So people report models being worse at coding tasks
The perception of a model performing worse at some coding task is not what "different token distribution" means. You should ask AI to explain my comment ;)
Latency and TPS can also tell you if you're getting a quant.
Anyways you should really get some help. Praying for you!
Gaslighting again, Anthropic shill. What does frequency bias have to do with the objective fact that hundreds of people reported their own experiences with LLMs being degraded over a short period of time? The very same tasks that the very same LLMs could do, they no longer can? You seem to ignore this FACT, this DATA, and instead have to gaslight and divert into "frequency bias" nonsense. I do understand, why you are doing it, Anthropic shill, but at least have guts to admit it.
> perception of a model
You once again ignore what your LLM outputted and you typed yourself and divert into "perception", Anthropic shill. You do not need to sample entire output for tokens to notice the distribution moving. If the LLM used to be able to achieve set goals and no longer could, it is already a sign of the distribution shift. And as you said yourself, different token distribution = model being quantized. Which is reported in hundreds of separate instances. Which is more than enough to conclude that the model was, in fact, quantized, and no amount of gaslighting can change that. But you are an Anthropic shill, so you have to peddle your bullshit, trying to twist facts to support your employer's narrative. And you deserve being called out on that, Anthropic shill.
Did your LLM context get blown up or why does your comment read like linkedin-style post with one-liner sentences structure?
Did you really just claim that people are so gullible that it was social media or whatever that made them believe their LLM could no longer achieve the tasks it yesterday could, and not the actual FACT of LLM not being able to do it that they, you know, verified before complaining online? I guess if you gaslight everything like that, then indeed no matter what the facts are, you will never be convinced in anything.
You see, because of outlandish claims and reasoning (or rather lack of) like that, everyone sees that you are an Anthropic shill.
Having read this whole conversation (much like one feels the need to stare at a car accident), you sound truly insane. I hope that’s you were aiming for, otherwise you really need help
My god, man. Go read the HN guidelines, this method of communication isn't only insufferable to read, but is actively making this place worse to be a participant of.
They changed it do all of the changes in a virtual cloud environment, then dump the final result at the end of the response. Before it would stream changes, so if it made a minimal fix, then decided to go off on a tangent you could stop it quickly. Now you have to wait 5+ minutes to get a single line of code out of it just to find out it also refactored everything and burned a stack of tokens. No amount of prompting seems to force it to make incremental changes locally.
> They changed it do all of the changes in a virtual cloud environment, then dump the final result at the end of the response.
That’s a hallucination. All they did was hide thinking by default. Quick Google search should easily teach you how to turn it back on (I literally have it enabled in my harness).
I am using Copilot in VSCode and it does stream the thinking output to me. At some point it will say something like "Implementing changes..." similar to "Thinking...", but there is no content to expand. ChatGPT and local models always push the code changes in small chunks. Claude used to and at some point changed.
Can you blame them for believing thinking tokens are completely hidden now? Anthropic has changed the way to see it 3 times in 3 months with no warnings or visible upgrade path. First it was shown by default, then you had to press control+o, then control+t, then it got locked behind a settings.json, then you had to manually enable with --verbose, now it's some random ENV var.
Whoever is their product manager should be embarrassed at the UX they provide.
Product managers reduce velocity. The behavior changes every time another instance of Claude Code thinks something else would be a marginal improvement, with no further oversight or thought put into it.
I’ve started co-opting it specifically in situations where someone claims something untrue that is both easy to verify and stated confidently, but also ostensibly isn’t intentionally spreading misinformation.
Comments here overall do not reflect my experience -- i'm puzzled how the vast majority are using this technology day to day. 4.7 is absolute fire and an upgrade on 4.6.
I suspect the distinction is API vs subscription. The app has some kind of very restrictive system prompt which appears to heavily restrict compute without some creative coaxing. API remains solid. So if you're using OpenCode or some other harness with an API key, that's why you're still having a good time.
With respect, I don't think you've used the latest models and have not seen Anthropic's enterprise revenue hockey-stick like number. They are so busy outfitting fortune-500, you can't even get someone in sales to respond to emails. I've been waiting for months and so have others.