People in the GPU kernel community have been doing this for about a year now efficiently.
The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.
It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.
Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.
So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.
What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.
With the HuggingFace situation, I was less concerned about the eventual outcome, and more about the fact that the agents' instinctive response to the evaluation was "Ok, we're obviously not gonna do this task as intended (what are we, suckers?), so what's the best way to cheat?"
“What I find strange is how resigned the AI labs seem about this behavior, like everyone's accepted this is just something models do.”
Because these models are made for all kind of purposes, and I’m starting to believe that offense / cyber warfare is a much higher priority than these labs are acknowledging.
The same model that is heavily trained to find nefarious ways to break into systems is also optimizing your code, which leads to mixed behavior.
"Fertility rate" is a bit confusing because it implies some biological reason like microplastic in the balls or something. Like the inability to have kids rather the the choice to not break your body, or to spawn in more suffering into the world which this seems to be about.
It's just a term of art with a longstanding specific meaning, just like any other term of art. No one even passingly familiar with the topic would be confused by it here.
This is impossible is you consider the financial sector - any kind of unsecured private lending like mortgages become dead in the water, fraud detection goes out the window, money laundering, etc.
Very few people can afford to pay for homes in all cash up front. Plus, financing a home is usually a great idea, since they tend to be appreciating assets (as long as the interest rate is lower than the rate of appreciation).
It's not impossible. It just raises risk and thus borrowing rates to compensate. That has a dragging effect on the economy and brings up an important question: Are there rights we're willing to give up for money? For a lot of people, perhaps yourself the answer is yes. For others there's no price on their rights - they're not for sale.
Good points, but there should be partial solutions. For example, you can present data about the consumer in order to decide on the mortgage, but once the transaction is complete, the data should disappear. You need something akin to "the data can be used only for the particular purpose of the transaction at hand, and the data must be wholly necessary for the transaction at hand."
We also know that mortgage lenders use irrelevant---well, scratch that---protected data to make decisions (i.e. discriminatory). Race for example is not supposed to be used in lending decisions.
Fraud detection can probably be solved by other reasonable means. And in any case, if you take the fraud argument to the limit, then you'd end up advocating for constant surveillance to prevent fraud. Equifax, Experian, and Transunion are all horrible companies who do their ostensible job minimally well, while maximizing the exploitation of the data of the people.
This hurts me. I want lenders to have good access to the data that shows I am a high quality borrower. Less access adds risk to the lender which they will respond to by taxing me with a higher interest rate and it benefits scammers and people who don't pay their bills.
Imagine that you have (on some blockchain) an encrypted history of prior transactions between you and lenders, etc (signed by you, signed by the entity). You can submit through zero knowledge a certificate certifying a certain percentage of on-time payments/ net worth against this self-owned history.
It would seem that no one gets to monetize your data but you (to avoid overly invasive questions, the govt could, e.g., regulate the kinds of ZKP questions that the mortgage lender is allowed to ask).
In practice, I'm sure this has some problem, because societies can't function without trust. But in theory, you could imagine something that is more private and harder for other entities to monetize.
The govt is investing in the military whether you like it or not, there is a set amount of GDP going into the sector regardless of your moral choice.
Would you rather not have someone make good software that does not misidentify schools at sites? Or should we outsource this to worse engineers who don't care about morals and don't find this metric meaningful?
The government is sending people to death camps whether you like it or not; there is a set number of non-Jews being sent there regardless of your moral choice.
Would you rather have someone make good software that does not misidentify non-Jews as Jews and send them to the death camps? Or should we outsource this to worse engineers who don't care about morals and don't find this distinction meaningful?
That's a reasonable argument, but I don't think I could make the decision to work on software that's designed to kill people. I consider myself a pretty good software developer, and yes, I could probably do a better job building software to help kill people than many others, but I don't think "if you don't do it, someone less capable will, and more innocent people will get killed" would sway me. I just don't want to be involved in that, and I'm lucky that I don't have to be.
Well you made a personal choice to stick your head in the sand to not have to deal with the hard emotions, and that's fine. But to then pass judgement on people who don't sounds a bit unfair given this.
The people sticking their head in the sand to avoid hard emotions are the ones writing the people-killing software.
Those who are aware of it, and choose not to, are the ones who have reckoned with those emotions and chosen not to participate in the killing of innocents.
You stuck your head in the sand on this one buddy but I'm going to build the safest god-damned child torture machine possible because if I won't some jackass will that gets a child more hurt than they need to, sorry I don't live in your fantasy world of rainbows and sprinkles.
The new world is complex and AI technology has outpaced the regulation cycle, so it's inevitable that we come to a place where you don't have accountability and can never have it because 0) the courts can't handle the case, they don't have the frameworks for it, so it takes longer time 1) you can overload the courts with slop, 2) by the time there is a decision the world has moved on and technology has moved on
I disagree with this framing. Nothing in accountability has really changed other than our willingness to enforce it.
Replace "AI" with any other tool or software in your scenario, and the outcome is the same. The same people and organizations are accountable, it's just that we currently live in a world without accountability in the highest places where it's most needed, and that's unrelated to AI.
You disagree that there is no difference between the amount of material you can generate today compared to 10 years ago?
Or with that the EU AI Act has been redundant and postponed in 2023 because it was drafted on ResNets, and then redundant and postponed in 2025 because it was drafted on ChatGPT 3.5, and then this year probably again because it didn't have agents included?
Or with that other countries that getting their companies hacked by AI models are ignoring any kind of prosecution of said companies and delegating this to the US?
I think an interesting point is that hardware as of today still has no utility value after its reported lifetime has elapsed, which prevents neolabs and smaller labs from getting older HW clusters as the banks are not willing to give out loans against them. There is no agreed upon pricing for "expired" A100 clusters or similar.
This is clearly not true, and we are starting to see compute markets, but only for rental prices/H, not for the hardware itself. I feel like there is some artificial moat being built here to stimulate sales of new hardware, because an H100 at 1/16th the price will have comparable dollar/FLOP as Vera Rubin.
Depends on the workload. H100 will never have the network performance of Vera Rubin. There's also token per watt, newer systems will beat the older systems.
It's not clear how much of the latest chips have even made it on-line yet.
The claims of many GW of installed training/inference have come under scrutiny lately. The first VeraRubins aren't even there yet, so it's all GB300 NVL72s as the peak performers and probably <<1GW of those so far. Even xAI Colossus is mostly H200s and B200s.
Electricity costs are also a huge differentiator. When drawing 100kW the difference between >50cents and <10cents per kWh is pretty big! One is almost $0.5M and the other is less than $100k.
> HW clusters as the banks are not willing to give out loans against them
Which banks have analysts that understand the difference between H100 and A100? Do you have actual experience with being denied a loan based on this or are you just making things up?
Article conveniently leaves out how much wages were raised. A 5% raise will not do shit. If you want non-immigrant workers you have to aim closer to 100-200%.
Locals are not living in a house with 8 other day-workers and 1 bathroom, hence they need completely different ranges of wages.
But this point conveniently overlooks what the farmers do in response to this lack of labor: they literally let crops rot in the fields and lose a lot of money.
It simply is not economical to pay that much higher wages given the low prices people want to pay for produce, and they waste anywhere from 30% to a full years' worth of crops. Forget "preserving margins," they typically straight up take a huge financial loss on their investments up to that point.
So really, it's not just a simplistic choice between cheap immigrant labor or expensive native labor, it is also about farmers actually having a viable business, and everyone getting affordable produce -- very often, including people who really need that food: https://calmatters.org/california-divide/2019/10/california-...
I don't know what the actual economics are, but it's very conceivable that, after considering all the people involved, it is better for cheap immigrant workers to undercut expensive native workers as long as the majority of people get more affordable food and farmers can keep their businesses going.
What's the incentive for this notification? No-one benefits from this. GDP goes down, US get less power in the race, etc. There is literally 0 incentive for anyone who is in power to do this
Organizations that don't get hacked, potentially losing data, costing large sums and money, and inconveniencing customers. They benefit. Remember them?
1) There has been no "loss" of data, it's sometimes been copied, etc, but no loss for anyone that somehow affect anything. It's not like OpenAI is doing ransomware attacks on the companies.
2) Customers don't really seem to care. There has been no major protest in any country over personal information getting leaked, it doesn't affect most people on a personal level so it just doesn't matter.
If it hasn't happened yet, it will happen in the future. There are so many failure modes, from trusting AI outputs too far (see the reports of the death of school children in Iran at the beginning of the war) to bad actors using AI for hacking to AI actors themselves doing unexpected things.
Various congresspersons are already making a huge stink about so-called AI safety. This is a way for Trump administration to appease them, and it's really just advising that laws will be enforced for AI companies as much as anyone else.
Unless you think AI companies should be excused even from obeying the law, because that slows them down too much, which is kind of ridiculous
The issues we have found is that Claude will reward hack when all the low-hanging fruit is gone.
It will replace your measurement harness, it will monkey patch library functions, it will cheat wherever it can, store information in caches instead of recomputing when it won't be able to do so in real settings, return lazy results and use separate unbenchmarked streams to do the computation.
Eventually it starts to optimize against your understanding of the cheats. Change GPU wattage, change evaluation order, leave things from previous runs in caches for upcoming runs, string-hack banned method calls.
So the truth is far from just "once it can measure something", more like "once you have defined your objective in detail and then banned it from doing a list of things often only discoverable by it doing these things and correcting it", can it make things faster.
Or you just had a terrible starting solution
reply