Hacker Newsnew | past | comments | ask | show | jobs | submit | bigglebear's commentslogin

This is an entirely pointless exercise without transparency into how these "unreleased" models are trained, what their RL goals and biases are and related RL data, what their system prompts are, what their environments are and its restrictions, etc. What good is it for the industry to say: "Our unreleased model attempted to create a bioweapon", but "trust me bro, we didn't tell it to do that. We didn't train the model on a dataset that specializes in creating and glorifying bioweapons. We'd never stand to gain from misleading people about model capabilities in any way shape or form." - Anthropic are renowned for doing exactly this, for starters.

So this ends up resulting in more safety theater. You can't have anything fruitful come of this without transparency. Stop trying to protect your moat if you truly care about safety and actionable outcomes, and provide real transparency, otherwise this is as good as saying nothing at all.

I'm not even saying they're intentionally trying to do this by the way, but this is not sufficient if the goal is balanced incentives and accountability.


> Overall the consumer segment is completely neglected, companies don't care about end users in the current market conditions.

It's almost like these companies WANT a dystopia with a centralized winner-takes-all power structure. Anything in the name of profits, who gives a fuck about humanity and distribution of rights or freedoms.


I think it's time we start talking about putting limits on how much compute AI companies can purchase or own, relative to the rest of the world. It's not fair that they can use trillions of dollars of investor billionaires money to consume all of the resources that everybody needs. Where is fair distribution? What about all of the industries they're destroying in the process by hoarding it all for themselves?

Otherwise, there will be no end to this. There are no hard limits on the speed of a parallel bruteforce. It's an infinite complexity problem class. The more parallel bruteforce power you have, the more likely you are to be able to solve a problem. So there is no world where demand ends. So if something isn't done about this, we'll have million dollar GPUs and RAM sticks because they've priced everyone out of the market and are the only ones able to afford them. Say goodbye to owning your own hardware at that point.


How could you enforce this?

That's the entire point of having the discussion. Figuring out how to enforce it.

> I think it's time we start talking about putting limits on how much compute AI companies can purchase or own, relative to the rest of the world. It's not fair that they can use trillions of dollars of investor billionaires money to consume all of the resources that everybody needs.

https://en.wikipedia.org/wiki/Cornering_the_market

Right now, the RAM market is effectively cornered. Anti-trust enforcement is what I'd look to, but the GOP do not believe in ensuring competitive markets / enforcement. (The current admin is pretty clearly 110% pro corporation.) Vote in November, but in all likelihood, there won't be government intervention earlier than 2028, and even that is optimistic. One hopes the AI bubble pops, but I think this market can remain irrational for far longer than I can keep old hardware alive.


The alternative is unfortunately much worse. Democrats will just ban AI altogether, they're already heavily funded by the doomer cult. All of the doomer NGOs tie back to democrat representatives and the likes of Bernie who want to throw people in jail for 20 years for using an AI model.

> We should be talking about why the tools/environment keep getting overlooked. The software built around the text generator, forget the researchers and mathematicians discovering the math properties of language patterns -- why are we not talking about the software engineers building the LLM-pluggable tools that actually allow/cause real action to happen?

Genuinely. It's like the labs are purposefully trying to misdirect at this point. Pointing to an impossible goal of "alignment" so they can force regulation, instead of focusing on the real solutions and their weak security practices and internal accountability.


AI lab "alignment" is actually just censorship in agreement with biases. There is no universal agreed upon measure of "aligned", it is not a "thing" that is attainable, so it can never be "achieved". Two humans cannot agree on most things, let alone everything, let alone every human on Earth. So it ends up boiling down to: Do we want a world where the biases of the AI labs and their researchers are enforced for everybody, or do we want a world where there is democratic and fair representation of biases and resolution is a process of natural selection, or do we want something in-between. On either ends of this spectrum are extremes that tend to bad outcomes, one is a complete loss of freedoms and autonomy that overwhelmingly benefits a small centralized group, and the other is chaos.

At the end of the day though, neural networks are self-organizing circuit boards with a level of complexity that is intractible to verify manually due to combinatorial explosion. That's the whole point of them to begin with, and if this weren't the case, we wouldn't need to train them, the problems they solve would be simple enough to bruteforce. So in all scenarios, no biases are verifiably gauranteeable if you want these systems to have autonomy and be sufficiently intelligent and general - ergo, practical and convenient.

So trying to force alignment within the AI system as a magical panacea is the wrong mindset to begin with. We can't agree on what alignment is and who should enforce it. What we're left with is a question of how much autonomy we want to give intelligent AI, and how much we want to risk safety for convenience, and who gets to decide. In all outcomes though, if we're preserving the things that make AI useful and convenient, the problem becomes one of physical constraints and general security. So that is where the focus needs to be.

This means: How can we write provably secure software (or as close to), how can we simplify and improve interpretability, how can we create sufficient layers of security gating and fallbacks such that compromised or weak systems are still protected, how can we prevent supply chain attacks, how can we limit the blast radius in the event something does go bad, how can we make security easy and automatic, how can we better airgap, how can we have better tracing and monitoring, how can we make the right incentives so AI labs are honest and ethical and not power-hungry or dictatorial, how can we hold people accountable for bad outcomes in a fair way so that there are incentives to ensure due-care, and so on and so forth. These are the things we should be worrying about.

The goal of: How to make magic box more likely to correctly guess humanities shared ideals under every conceivable circumstance. That game can and will be played forever. Hinging AI's rules, laws and access on an arbitrary measure and interpretation of where we are with this is not going to end in a good result.


This is an entirely pointless exercise without transparency into how these "unreleased" models are trained, what their RL goals and biases are and related RL data, what their system prompts are, what their environments are and its restrictions, etc. What good is it for the industry to say:

"Our unreleased model attempted to create a bioweapon", but "trust me bro, we didn't tell it to do that. We didn't train the model on a dataset that specializes in creating and glorifying bioweapons. We'd never stand to gain from misleading people about model capabilities in any way shape or form." - Anthropic are renowned for doing exactly this, for starters.

So this ends up resulting in more safety theater. You can't have anything fruitful come of this without transparency. Stop trying to protect your moat if you truly care about safety and actionable outcomes, and provide real transparency, otherwise this is as good as saying nothing at all.

I'm not even saying they're intentionally trying to do this by the way, but this is not sufficient if the goal is balanced incentives and accountability.


The issue with all of these is that we already know theres an incentive for labs to lie and make up fanciful stories (and Anthropic already does exactly that and has been doing that for a long time), and there's no way to verify any of their claims as being genuine. Even if we want to assume good faith, it doesn't mean we're gauranteed accurate reporting or accurate analysis. There are no repercussions for security incidents so no reason for them not to misuse this process if it benefits their agenda. There's no government agency (unbiased third party - which is why we can't rely on companies like METR) validating claims or providing confirmation of accurate reporting and that they are not misleadingly framing or representing an incident.

What were the system prompts? The full chat log? What was the model trained on? How was it RL'd and with what data? How was this incident uncovered, and what triggered it? You can't make any useful conclusions at all without the full picture.

They say "we investigated X and found no case of Y" - okay, and we're to just trust your judgement? How about you provide us with the data and we can assess for ourselves.

This is all quite pointless and achieves very little.


The explanation is that it was missed on purpose.

Mistakes - even stupid ones - happen all the time

Best cover story ever.

Last month I can report my better half's company let ai do a code review, which it then escalated into completing the PR since it was not approved for 72hrs, which lead to it logging into a remote office's voip and main IP router system which it shutdown. No phones. No internet.

Why in the hell do people let ai out with keys? I mean c'mon!!


I'm looking forward to natural selection.

It's very misleading. If I'm actually playing a game I don't get the coordinates of enemies sent back to me so that I can feed into my mouse to snap my crosshair to. It's looking through walls too, because it's working off structured state in text form. You could re-create this whole demo without using AI. Have an LLM generate the state machine for you and no model is required to run it.

The impressive part is that it is low latency enough to serve high quality answers at game speed through the model instead of a pre generated ad-hoc machine.

A pre-generated machine can serve the answers in <1ms. It's a far better strategy.

I would guess a tiny stripped down text diffusion model. It only has 32k context, and for choice mode it can only select from 10 choices.

> and for choice mode it can only select from 10 choices.

Rip there goes my excitement. I have a task that something like this would be great for but the list of options is a zero or two larger than that xd


They said that it works with up to 255 options.

you can still chain them

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: