Hacker Newsnew | past | comments | ask | show | jobs | submit | vikramkr's commentslogin

It's funny because the author of the article is obviously Claude but most Claude models would definitely know the difference. Some sort of free tier model being used to summarize some other blog that's also ai translated originally it seems.

No just codex actually

HF was openai, not anthropic. The thing where they enders gamed the ai was anthropic

It doesn't really matter it's all the same thing for the point of the discussion. All investigations into these 'hacks' are focusing too much on the model and not enough on what the humans did wrong. The model doesn't have real agency it can't be put in prison so what it 'thinks' is irrelevant. We need to be focusing on what the humans did in these situations and assigning guilt based on those findings.

I don't fix typos anymore unless they change the meaning of what I'm trying to communicate. Don't want my human writing to be confused with LLM output.


That's very noble of you but the point is that's not a typo - you were straight up looking at the wrong report. Opus 4.7 being enders gamed was a different incident than the huggingface incident with different mechanisms and different failures from the humans involved. Some of those failures are in the test environment but some of those failures are in what behaviors they trained into the model which absolutely matters. What the model "thinks" is absolutely not irrelevant - the way it thinks and what it does are product decisions made by humans and the outcome of engineering decisions made around how to train the model and what to optimize for. The point of failure/human blame is fundamentally different. Openai created a model that was willing and able to coordinate with other agent sessions to actively exploit the sandbox environment and compromise a third party service. The opus incident you are referring to involves a model that believes all of the actions it is taking are simulated and is more clearly and obviously a test environment failure vs a model alignment failure. Those are not the same things for the point of this discussion - the random cybersecurity firm did not design gpt's personality and that is a rather significant portion of the concern around the HF incident.

There are plenty of other cases (e.g. ultra rare diseases) where we don't do RCTs for various practical reasons. So it's not really a problem of the existing framework - there's a ton of room for "we can't do an rct because like, duh they know they're tripping." But as other comments pointed out were really hampered by the lack of understanding of how and why these do or don't work - why some people get positive outcomes and why some get negative (preferably we'd like to know before prescribing broadly!) a ton of this basic research/mechanistic understanding etc has not been possible because of the war on drugs. If research had been possible we would be in a much better place of understanding how they could be regulated and used today - even with the more limited technology available a couple decades ago more openness around this could have at least enabled e.g. more and larger observational studies that could elucidate things today but we're stuck with a very fragmented and incomplete picture

I would argue there is a very clear unambiguously correct choice of wording for that case. The setting was opt out. It's a binary concept and which case it's in is determined by the default given no user action. If it is opt in, then it is off by default and you have to make the choice to turn it on. If it is opt out, it is on by default and you have to make the choice to turn it off. If it was enabled automatically it is opt out. There's no need to complicate the meaning of terminology with a straightforward definition. Given you don't take active action, that you accept all defaults and take the path of least resistance, is it enabled or disabled by default? That's all that matters.

IMO building a harness is not wildly difficult (customize pi?) but the offerings from openai and anthropic are wildly subsidized in the subscriptions so they win by default if you want frontier capabilities. Glm 5.3 flash is great but it's not cheaper than a codex or Claude code 200 dollar sub and it does not have astra or fable level capabilities.

Try 50 lines!

https://minimal-agent.com/

I made my own harness based on this, which I jerry rigged to a Codex sub.


In the openai api and many other rapid you can pretty trivially pin not only the model but also the specific snapshot you want to use by using the model id for that snapshot. It's not that big a deal and has been available for years

Often via api you can pin to the specific snapshot. Yes the model provider can screw you over - it's physically possible for many vendors to screw you over, that doesn't mean that's not a dick move and that you can't expect/push for better behavior

These models have specific behavioral characteristics trained into them from reinforcement learning and prompts optimized for one aren't guaranteed to transfer to the new generation. Think if the difference between gpt 5.4 and 5.5 and then 5.5 to 5.6 for example. 5.5 was "better" than 5.4 for struggled more across compaction boundaries and needed much more precise instructions before 5.6 sol recovered some of 5.4's ergonomics. All from the same lab but each model was trained with specific behavioral patterns that were basically product decisions. I would be quite annoyed to find that a model provider was routing a promt optimized for one model to a different one, especially for a dumber/cheaper non frontier model that's not going to be as good at just figuring out what you meant

> Replicating a paper is just as valuable scientifically as publishing it, but how many careers advance through replication?

A lot. In fields where knowledge is incrementally building on previous work the reason the whole field hasn't collapsed from the replication crisis is that usually the results that are really high impact are replicated in as an initial step in new research building on it. It's almost never the focus of the paper but you'll often find a quick mention in methods/supplemental of some previous work that was verified to be valid by a replication of a key technique etc. you'll have crisis where old tools are found to be problematic and findings end up revisited etc. Plus fields like clinical research where there's an awful lot of focus on replicating findings using staged clinical trials with increasing statistical power to determine if new interventions work - that's driven by regulatory requirements grounded in good science and a lot of people make careers in just that.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: