Hacker Newsnew | past | comments | ask | show | jobs | submit | RandomLensman's commentslogin

Not sure that is enough for forecasting as the function to be estimated could change over time in random ways.

Wouldn't this mean better sandboxes are needed for some things, for example (might include very strong airgaps even)? Breaking out of something isolated electromagnetically, optically, and acustically is not easy.

That works as long as no one ever interacts with the models, which would make the models themselves useless.

Could sit in the box and interact if a model of certain capabilities is needed/tested. We do physical security for other things, too. Not saying everything needs that type of isolation.

Why would 41C be the limit for the hot side? How does it work now in hotter places?

In short: Things degrade.

More: https://angliancompressors.com/blog/how-ambient-temperature-...

There's a T3 class for up to 52 degrees Celsius. While it's higher than 41 degrees, it's not significantly higher:

https://zerohvacr.com/blogs/t1-t2-t3-explained-climate-class...


Is 52C a physical limit?

Yes. The compressor can't pump more heat outside since it's already hot from working and trying to pump out the extracted heat as well.

For T1 class of devices, that 41 degrees is also a physical limit, not an artificial cut, because of their design.

You know, thermodynamics.


Why would the limit of T1 devices matter when there are ones that already can do hotter? Thermodynamics make things harder to remove heat but not sure how they forbid it (you need the hot side to be hotter than outside to pump heat). Might need stages, for example. We can cool things 200K and more below ambient.

Early data processing machines were used in the Holocaust, for example. Or cipher devices played important roles in WWII.

Mongols did it on horseback. It's mostly a matter of will/capability and no scruples.

Why do we need modern weapons for defense at all then? Horses, bows & arrows plus will not enough?

Because attacks have evolved. Defense follows the attackers.

Good luck hitting a Mach 15 missile with a trebuchet.


Genocides are more common in the 20th century, because they are easier. Due to technology.

Not sure it needs to be about publications and data analysis technology as such to regulate AI development, for example - could be about processes, procedures, people, etc., no?

Not sure we need to experience all possible issues to mandate certain things. We don't do that in other areas either, no?

Nevermind, I have no idea honestly.

Not just on prediction but in parts also based on just not wanting certain risks. We can and do deem some things inherently risky, up to the point of banning them even.

Why wasn't it airgapped, for example? How was the action not allowed? Or do you mean in some weak sense, not in a hard not possible? RL systems doing weird and expected things wouldn't exactly be new, no?

We police people working with all sorts of dangerous things, if we think AI dangerous why not do that here, too? We don't just leave things up to people on the ground or companies.

Edit: I think the post I replied to changed a bit - nevermind. A complex topic.


I read more about the incident, and was offering up way too much opinion not grounded in 'fact' (barring philosphical evidence).

It's a complex topic for sure.

I stand by my opinions about frontier work, pushing thr edge, and connecting ideas.

But I have no idea, and haven't given much thought to what it means to enforce regulation that would also slow the forward advancement of technology, the economy, etc.


Big if.

What was the inner state there? How would something not being allowed expressed internally? Maybe such language is one way to elicit certain behavior but not a statement of what was permissible?

I'm referring to their transcripts of the reasoning and output tokens - this doesn't go into the detail of evaluating hidden states as there's also iirc evidence of better models having one internal state but putting something misleading down in the "reasoning" tokens.

The either output or reasoning tokens, or perhaps in the messages they were sending each other on the boards they created, have them saying explicitly that doing these things to HF were not allowed then doing them anyway, or at least not notifying people. What I'm getting at broadly is this was not a case of "we told it to attack however it wanted and it chose to hack HF" or "we told it to attack a simulation but it did the real thing" or "we explained not to do that but it was so far back in the context window the models acted like they never saw it" or even "the instructions were not clear".


Yes, my point was more that I don't know whether parsing those outputs as a human is a useful thing to do or not (even though it is in human language of sorts). What machines mean or want elecit might be different from a human interpretation, especially in relation to any RL "forcing".

There’s definitely issues with using them to understand what the models were “thinking” but we can use them to answer a few questions. Most relevant here is that the idea or instructions that attacking hf would be out of scope was not simply lost in the context.

Does that work with RL? Simpler RL systems already have done weird or unexpected things (even simple optimizations are prone to home in on errors or incorrect inputs to create poor results)? Could be easier to limit certain things, have processes and controls outside etc. instead of trying to align (as we do in a lot of areas when using machinery).

RL things doing weird and unexpected things isn't new - much simpler things than current AI already show that.

That said, we have a lot of experience working with (potentially) unaligned machines and things of various degrees of risk (from heavy machinery, to pathogens, to humans) and the approaches include various measures and procedures to control, contain, limit, etc. that are outside of the thing - not sure why that isn't a possible direction (or maybe I misunderstood).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: