Hacker Newsnew | past | comments | ask | show | jobs | submit | jsnell's commentslogin

The impact from a model not being included in the training set is increased false negatives for text generated by that model, not increased false positives for text written by humans.

Source?

The latter are not doomers at all, that's just a misuse of the term.

A year ago I would have described that group as AI skeptics or AI deniers depending on how charitable I felt on any given sany. At this point AI truthers seems more appropriate.


There was a recent video doing rounds in twitter. An eclectic group of doomers and truthers who found an unusual common ground on hating on AI.

Nate Soares (doomer high priest who has called for pause before LLMs were a thing) was explaining why AI should be stopped due to x-risks. Ed zitron (truther extraordinaire) asks him what has he done to stop it which is funny when Nate spent big part of his life on this mission.

It was strange watching two people who have completely different ideas on AI speaking on a common ground


You should probably look at the cost/score graph by effort level instead:

https://artificialanalysis.ai/models/claude-opus-5-5#intelli...

It is most of the pareto frontier.


Not disputing the increase in quality, just stating that non-cherry-picked benchmarks show it is more verbose at Max effort

so don't use it at max? The benchmarks suggest that high/xhigh are more than sufficient to be ahead and a whole magnitude below max with regards to token usage. I'd treat that as an outlier and not how verbose the model is in general (QED I know)

You’re missing my point. I’m saying anthropic are exaggerating their results.

how are they exaggerating the results? Comparing the cost from that chart for 5 and 5.5 for medium-max effort paints a pretty clear picture:

         mean  median
 model
 5      4.135   4.245
 5.5    3.150   2.640

Again seeing how max is a clear outlier, the median cost saving is ~38%, not that far off from the proclaimed 40%.

Is verboseness the only measure of token efficiency towards overall task completion?

> Half of the 832K LoC are mostly tests she answered:

No? That quote is clearly saying the opposite of your summary.


True! I edited.

Flagged.

This is like a web page asking for your email password to check password strength, promising that it won't leave your browser. Maybe it doesn't right now, but who knows about tomorrow.



A keyboard/video/mouse switch.

In this context its not a switch - its a Keyboard / Video / Mouse - but over IP - remote control of a PC (including boot and bios)

I've never entered a datacenter before – what does that mean / what is it for?

There are really two types of KVMs... classic ones were used by consumers to plug multiple computers into one screen, keyboard, and mouse and flip back and forth between them.

Networked KVMs let you remotely access a computer or server. It's like VNC except by physically plugging into the ports. This is useful because it allows you to have remote access to a server when its OS is offline, tweak bios settings, watch it boot, reinstall the OS, etc.

Most true servers have this functionality and more built in (ipmi, idrac, redfish, etc are terminology for the feature). These plugin ones allow you to add that functionality to any machine. You're not likely to see tons of them in an actual data center, but they are incredibly useful for systems without remote management built in.


Got it! So I could for example plug this into my mac mini at home and reboot it after a power failure or something while traveling?

You can even boot it from powered off, but (at least for mac mini) this probably needs Wake-On-Lan to work.

Yes, exactly. And even do things like inject virtual USB drives so you can reinstall an OS or something

But the proposal the article is reacting to is for literally none of that! It is quite literally the opposite, with its proposed measures applying only to frontier labs rather than grandfathering them. It doesn't say anything about open source training, self-hosting, open weights, or startups. It does not suggest a ban on Chinese models (just better enforcement of chip export controls).

> Any AI model a company offers to the public has to be released as open weights.

The obvious outcome of this plan would be frontier models not being released at all to public It'd be an amazingly bad outcome. Models accessible to the public would stagnate (no capital, no access to frontier model tokens to distill from), while internal models would keep improving at their previous pace, use them in-house, and eventually eat the whole economy.


That's a legitimate possibility but it's also much slower, more risky, and harder to raise money for than what they're doing now.

If Anthropic/OpenAI would've shown incredible results using these models then I'd be worried about this. Instead we get Codex and Claude Code, bloated and disappointing software. I'm sorry but "use them in-house, and eventually eat the whole economy." doesn't appear to be a real concern with these two companies.

I honestly don't even understand what straw man Doctorow is arguing against here.

But he is wrong on the facts: these incidents were not merely the models already being in a infosec context and escalating beyond the intended parameters. They happened also with no kind of security elicitation. So the task was something like searching the internet for economic statistics, not to hack into a system.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: