Wouldn't this mean better sandboxes are needed for some things, for example (might include very strong airgaps even)? Breaking out of something isolated electromagnetically, optically, and acustically is not easy.
Could sit in the box and interact if a model of certain capabilities is needed/tested. We do physical security for other things, too. Not saying everything needs that type of isolation.
Why would the limit of T1 devices matter when there are ones that already can do hotter? Thermodynamics make things harder to remove heat but not sure how they forbid it (you need the hot side to be hotter than outside to pump heat). Might need stages, for example. We can cool things 200K and more below ambient.
Not sure it needs to be about publications and data analysis technology as such to regulate AI development, for example - could be about processes, procedures, people, etc., no?
Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even.
Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no?
We police people working with all sorts of dangerous things, if we think AI dangerous why not do that here, too? We don't just leave things up to people on the ground or companies.
Edit: I think the post I replied to changed a bit - nevermind. A complex topic.
I read more about the incident, and was offering up way too much opinion not grounded in 'fact' (barring philosphical evidence).
It's a complex topic for sure.
I stand by my opinions about frontier work, pushing thr edge, and connecting ideas.
But I have no idea, and haven't given much thought to what it means to enforce regulation that would also slow the forward advancement of technology, the economy, etc.
What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?
I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens.
The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly that doing these things to HF were not allowed then doing them anyway, or at least not notifying people. What I'm getting at broadly is this was not a case of "we told it to attack however it wanted and it chose to hack HF" or "we told it to attack a simulation but it did the real thing" or "we explained not to do that but it was so far back in the context window the models acted like they never saw it" or even "the instructions were not clear".
Yes, my point was more that I don't know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL "forcing".
There’s definitely issues with using them to understand what the models were “thinking” but we can use them to answer a few questions. Most relevant here is that the idea or instructions that attacking hf would be out of scope was not simply lost in the context.
Does that work with RL? Simpler RL systems already have done weird or unexpected things (even simple optimizations are prone to home in on errors or incorrect inputs to create poor results)? Could be easier to limit certain things, have processes and controls outside etc. instead of trying to align (as we do in a lot of areas when using machinery).
RL things doing weird and unexpected things isn't new - much simpler things than current AI already show that.
That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside of the thing - not sure why that isn't a possible direction (or maybe I misunderstood).
reply