Hacker Newsnew | past | comments | ask | show | jobs | submit | simoncion's commentslogin

> It’s real but whether you can make a profit or not is dependent on whether you can protect the information.

That's what contracts and licenses [0] are for. Or perhaps you're arguing for a world where the only possible protection is that of trade secrets... that once protected information is made public by any means anyone can do anything they wish with it? If you are, then this quote [1] seems relevant:

  If IP is real, then the AI companies have performed flagrant theft.
  
  If IP is not real, then the algorithms and weights the AI companies have developed should also be free as they are just more information.
[0] ...which are contracts in disguise...

[1] <https://news.ycombinator.com/item?id=49754448>


To read this post you must pay me $5. Please leave your contact information as a reply or you'll be hearing from my lawyer. [1]

Those contracts and licenses are just a form of protection enforced by the State. In general once knowledge or information is widely available it's also freely available. Whether someone can further protect the distribution or the usage of that knowledge to some ends is up to them.

I disagree with both of those premises.

IP is real, and scraping freely given comments or published material on the Internet and doing something economically useful with it doesn't entitle you to retrospectively go back and decide that you are owed some money. If that were the case you have to prove material damages. How much is my post here worth? one quadrillionth of a cent?

I think it can become different if, for example, a book was scraped or cataloged but that's primarily because it's likely that you can enforce or make a case within a given legal system to enforce your copyright or IP claims. But what if I read a book and then thought the main idea was my own, or just spoke to someone else about it and they spoke to someone else about it and it winds up in an LLM? There's nothing wrong with that or anything you can do.

To expose information is to put that information at risk of being used by others. It's up to you to protect that information or enforce claims on it. If Reddit's public site gets scraped and OpenAI does something economically useful with it, well, that's just life. You can't put information out in the public for free and then demand payment later. Reddit comments are freely accessible by the public, yes? Companies are part of the public.

If you want to argue that it's IP theft then distilling weights or otherwise reverse engineering the models is a violation of IP protections too.

[1] Rhetorical


> Those contracts and licenses are just a form of protection enforced by the State.

Yes, that's how property works most anywhere that has stable government. Societies where The State has a monopoly on violence are usually far more stable and pleasant to live in than ones where vigilantism is how correction of injuries is done.

> Reddit comments are freely accessible by the public, yes?

In exactly the same way that the books in my local public library are freely accessibly to the public, yes. I'm sure that you're aware there are so many things you can't legally do with books in your local public library. Reddit isn't substantially different... go read the "User Agreement" contract that governs use of the site when you get a free hour or two.

> To read this post you must pay me $5. Please leave your contact information...

You can totally prevent me from -among other things- making many sorts of commercial uses of your post, but -as I'm sure you know- the made-up system you're gesturing at not how software licenses have worked for roughly as long as software has been a thing that was commonly licensed.

Anyway. It sounds like you want a return to "publish nothing, keep everything of worth secret to anyone not in the guild, and do what needs done to people who the guild suspects has betrayed its secrets". Those were the really bad old days, and the awfulness of that system was why the US adopted the patent system back in the late 1700s. It's also part of why it adopted the copyright system [0] as well as limited term lengths (fourteen years, with an option for fourteen more if the work's author was still alive and explicitly requested an extension).

[0] Though, you had to explicitly register your work with the relevant authorities to get protection... unregistered works received no protection. Things stayed this way for roughly two hundred years, and -IMO- should have stayed that way.


> ID chips can't be cloned...

ID chips can be manufactured, so they can obviously be cloned.

Thinking like yours leads to the asinine situation we saw ten, fifteen years ago where insurers were refusing to pay out vehicle theft claims because "There's no way to clone an RF keyfob or RF immobilizer chip!". Spoiler alert: There were many, many ways to do that.


Yeah, it's more that they can, if designed correctly, be made very difficult to clone. (and that if in the middle is pretty important!)

There are also many ways to steal a car without having a key at all.

You appear to be confusing "avoid" with "never interact with".

If your professional life has required you to spend enormous amounts of time working very closely with people you don't like and/or don't trust, then -unless it has made you enough money to retire after a few years' work- please accept my condolences.


That is what avoid means. Avoiding something means staying away from it to the best of your ability. You may well be constrained by circumstances such that you sometimes have no choice, but the intention of avoiding something is to come into contact with it never.

> Avoiding something means staying away from it to the best of your ability

Yes, that's what I said. "To the best of your ability." often means "You have to occasionally spend small amounts of time working with really shitty people.". OP said:

  following this advice would have made most of my professional life impossible...
Unless OP's professional life made them a ton of money, either they had a uniquely terrible professional life, or what you and I understand "avoid" to mean is not at all what OP thinks the word means.

Oops, just wrote something similar before seeing this. It came up recently in another community where it seems some people take the word "avoid" to mean "never" and in this case: "never not even in the past."

At the time of this writing, the subtitle of the submission here on HN is

  recovering the signing keys for US driver's license barcodes
Notably, this subtitle doesn't appear on the blog post.

Anyway. I only see claims that the public key can be determined from license barcodes, not that a signing key can be determined. What am I missing or misunderstanding?

To head off one potential retort: While it's true that one can use a public key to encrypt data for the recipient that has the private half of that key or verify that data has been signed by the possessor of the private half of that key, I'm almost 100% certain that it's not possible to use that public key to sign data would validate to other folks as being signed by the private half of that key. It has been more than a decade since I've thought about any of this, but isn't the entire point of public-key cryptography that the public part can be distributed to your worst enemy without causing you any trouble at all?


Yes. The subtitle is wrong. He recovers the public key, due to the way EDCSA signing works.

Yup. The person who submitted this to HN is probably way less knowledgeable on this topic than the writer of the article. The article clearly labels the recovered keys as “recovered public keys” at the top.

Given the poster's nickname, the person who posted it and the author may be one and the same.

I think the author is Claude.

the first part of the article was interesting, but I couldn't finish it because it reads so heavily in claude's 'voice'

> Predictably the discussion is already veering towards OpenAI's negligence...

> Cherry on top: That was part of an attempt to jail-break itself via self-prompt injection.

We continue to see so-called "prompt injection" "attacks" in the wild that override a user's intended program with an attacker's [0], and/or the LLM producer's intended "safety" instructions with the user's. The fact that this sort of program hijacking is possible at all is strong evidence of negligence. Why?

OpenAI and Anthropic both claim that they're working on very dangerous Internet-connected tools. So very dangerous that the production of and access to said tools needs to be tightly regulated, they claim. If one actually believes that the computerized tool one is working on is very dangerous, one generally doesn't design that tool so that it blindly executes instructions handed to it by complete strangers on the Internet. That's akin to connecting the sole activation switch for a biosphere-evaporating firebomb to the Internet.

The major LLM producers are so obviously negligent and -as a bonus- have openly admitted to committing cybercrimes [1] that would get people like you and me fined out the ass and jailed for ages if we did them. The tragedy is that they're making so much money for the rich and powerful that -much like the architects of the 2008 housing crash- they'll never see any meaningful punishments for their actions.

[0] One recent example is <https://agentic.tracebit.com/context-bombs/>, but there are so, so many more to choose from.

[1] ...the "cyber" prefix is so stupid...


Yes, prompt injection is an issue with model safety, which is what we should be focusing on. I meant to say the discussion about OpenAI's negligence in securing agents and their infrastructure is the red herring.

That said, nobody has been charged for these hacks despite openly talking about them because typically you need to show intent. If intent was not a requirement, they would have been in trouble way back when the first AI-assisted suicides happened. Lawsuits have been filed, but OpenAI's whole schtick is "these agents are so dangerous because they do all these crazy things without being asked to."

As far as we know nobody told the agents to do any of this, or even that it's OK to do this. If someone can find any proof of anything approaching actual intent, I'd bet there would be no shortage of attorney generals willing to be build their career on this case. After all, there are already many AGs investigating OpenAI.


> I meant to say the discussion about OpenAI's negligence in securing agents and their infrastructure is the red herring.

It absolutely is not. It's yet more evidence that the culture inside these companies is entirely inadequate for a company that's building what they appear to be claiming are WMDs that are very likely to be species-ending.

> ...because typically you need to show intent.

a) You seem to be suggesting that criminal negligence doesn't exist. You also seem to be claiming that deploying and operating computer software that you built [0] that you don't just know but widely advertise has a "discover and exploit faults in someone else's computer systems" feature without ensuring that that computer software cannot access other people's computer systems isn't -when viewed in the most lenient possible light- incredible negligence.

b) Go look up the facts of weev's case. weev's intent was very obviously benign and prosocial. The only reason he didn't spend four years in jail and have to pay tens of thousands of dollars was because of a choice of jurisdiction error made by the Federal government.

[0] "You" in this case refers OpenAI, Anthropic, and other major LLM providers. Don't bother with a "But what if the people running the software had nothing to do with building it!" retort.


Of course criminal negligence is a thing but you should look into the bar required to prove it. For one, it requires the highest standard of proof. Then there are at least 4 very specific elements that need to be proven, and it would be interesting to see how that could be achieved without the benefit of hindsight, given that we've seen no technology like this until just 4 years ago.

If there was a case to be made for criminal negligence, I would say it would be with the AI-assisted suicides, and lawsuits are pending but we've not seen anything come out of that yet.

> Go look up the facts of weev's case. weev's intent was very obviously benign and prosocial.

Oh I'm very familiar with the case, which is why I know the investigation revealed detailed chat logs where weev and his accomplice discussed multiple ways to illegally profit from the breach, including shorting AT&T stock, selling the data to spammers / phishers, or organizing a spamming / phishing operation themselves.

I am very curious about what sources led you to associate the words "prosocial" and "benign" because weev is very, very, very clearly a noxious, antisocial person who spends a lot of his time harassing people. I would recommend you feed "weev criticism" into a Google AI overview. Random quote from his wikipedia:

> In early October 2014, The Daily Stormer published an article by Auernheimer in which he effectively identified himself as a white supremacist and neo-Nazi. He is known for his "extremely violent rhetoric advocating genocide of non-whites", according to the SPLC.

And as you said, he got off on a technicality. That case was heralded as a triumph of justice purely because it followed the rule of law despite the obviously unlawful intents of the defendant.


Since roughly zero of the articles I've seen link back to the source of their quotes, here it is: <https://status.aws.amazon.com/#multipleservices-me-south-1_1...>

> This is for data loss, of course, not access or surveillance or unlawful processing.

This is why the minority of politicians who actually know about how this stuff works worry about where the data resides for jurisdictional purposes. If the government where the data resides can compel the folks who have physical and/or logical access to the physical machines that contain that data to give them access to that data, then that's game over for you.

«But you just don't permit that sort of breach to happen!» you might say. To which I reply "Yeah, right.".

Substantial physical separation of datacenters is very important, but the politics and policies of the location housing the data cannot be ignored.


I mean, in those rooms I was arguing over the best policies to prevent access and surveillance and unlawful processing, and what the potential cost-benefit analysis was. And what I was arguing against was an assumption that physically compelling all companies -- or worse, all citizens -- to keep their data within the borders of the host country, would protect you from these problems.

We'd have to explain that if the data was physically in Brazil, but hosted by a U.S. company, that would not stop that company from accessing that data remotely -- unless you specified that. We'd have to also explain that if you were intended to defend against US mass surveillance of non-US persons by the US intelligence services, intelligence services and SIGINT are univerally almost defined by their broad remit to target foreign nations on their own territory in violation of local law. And, finally, if you intended to use the prohibiting the movement of of data as a sanction against companies to punish them for violating data protection standards, as pre-GDPR law in the EU had as an ultimate last resort, and the GDPR often ends up relying on as a last resort, you would find that multinationals are more capable of putting up servers in your home territory and continuing to serve your citizens than they are of substantially changing their practices regarding data processing.

I don't want to sound nihilistic about this -- regulations can exist in these areas. But it's those politics and policies of the institutions with control over the data that are the most important part of this: not where the bits are kept. Especially when those bits are encrypted, and the keys and access controls are elsewhere.


> But it's those politics and policies of the institutions with control over the data that are the most important part of this: not where the bits are kept. Especially when those bits are encrypted, and the keys and access controls are elsewhere.

Nah. Policies prohibit rule-followers from accessing data that you don't want accessed. Such policies are very important. But if you give your adversary effectively-unlimited physical access to the hardware where the bits are kept, that's game over. If you don't trust the governors of a region to honor the "don't tamper with this hardware" gentleman's agreement, and you very seriously care about preventing unauthorized access to the data that that hardware stores and processes, then you don't put that hardware in that region.

To point to a real-world example of this, there's not going to be an AWS Top Secret Cloud region in China, Russia, or -say- North Korea.


> If you don’t want it then just pay for ChatGPT.

That will work until it very suddenly doesn't. The industry is awash in paid Internet-delivered advert-free services that suddenly got adverts stuck in them.

And no, if you're in the US getting a refund for breach of contract won't likely be an option unless your contract explicitly permits you to terminate the contract and receive a refund for the unused time. At least one court has determined that adding adverts to a contracted service -in the middle of most contract periods, rather than at renewal time- that was explicitly advertised as advert-free and requiring contractees to pay an additional like 3 USD per month to remove these new adverts wasn't a breach of contract or a price increase, but was rather a "change in subscription benefits" and was A-OK.


> That will work until it very suddenly doesn't.

Exactly. Once upon a time, cable didn't have ads, which was its primary advantage. Then they realized that they could charge money for the service and advertise to you.


> Once upon a time, cable didn't have ads, which was its primary advantage.

To be pedantic: originally its primary advantage was allowing people who lived in bowls and other places that couldn't get adequate TV reception to watch broadcast TV. Later on, it became a way for those paid subscribers to get access to stuff that wasn't broadcast television.


Valid pedanticism!

Yep. An entity in the "dirty room" reads the thing to be reimplemented and produces a document that thoroughly describes its behavior. That document is passed to the entity in the "clean room" whose only knowledge of that system is through that document. Reverse engineering is legal, plagiarism is not.

Despite the fact that the raw output of the system is incomprehensible to humans, scanning a photograph of Mickey Mouse and running it through a lossy compression system like JPEG doesn't suddenly make it not a picture of Mickey Mouse. Similarly, running the code for a system through the lossy compression system that is LLM "training" doesn't suddenly obliterate that data and make that LLM a "clean room". If one has any doubt of that, remember that they are known to reproduce their inputs, even after all these years of work to make them not do that. [0]

[0] <https://news.ycombinator.com/item?id=49727685>


Plagiarism is legal, actually. Copyright infringement is not.

Fair-ish point. I was considering plagiarism in the context of computer programs, which is -at best- dreadfully difficult to do without also violating copyright and/or software licenses.

I'd also rephrase your first sentence as "Plagiarism isn't illegal, actually.". Unless you're rich and/or very influential, plagiarism is a seriously bad thing to do.


Getting expelled from a university is legal.

Sometimes. Other times is is most definitely not.

> Do you have many examples of this actually happening that you could share?

If by "this" you mean "an LLM [reproducing] copyrighted material", I have one here [0], with the challenge posed and plagiarized response at [1], found via [2].

One wonders how often the "Don't plagiarize, make no mistakes about this!" instruction fails and the LLMs include nontrivial chunks of other people's work into what they emit.

[0] <https://infosec.exchange/@zzt@mas.to/117134157775929932>

[1] <https://mas.to/@zzt/117122289150514171>

[2] <https://infosec.exchange/@david_chisnall/117134446424182178>


> If by "this" you mean "an LLM [reproducing] copyrighted material

No. I have absolutely no doubt that happens, and I'm not defending it or encouraging it.

My point is whilst it technically might be illegal and happening all the time, if it's unenforceable or sets a precedence that would severely break the business world, there's every chance nobody would dare bring it to a court room.

I'm not aware of it happening yet (someone trying to enforce copyright on code that was put in production via an LLM) and I would have thought if it did, we would all know about the precedence now.


> ...someone trying to enforce copyright on code that was put in production via an LLM...

AIUI, in the US something that's entirely machine-generated is not eligible for copyright protection. If one could demonstrate that that machine-generated output is plagiarized human work and were rich enough to bring it to court, and able to wait five to ten years for the outcome, I have to believe that the usual copyvio rules would apply because that's not machine-generated output, it's straight-up unauthorized copying performed by a machine.

> My point is whilst it technically might be illegal and happening all the time...

If it's illegal, it's illegal. Refusal to enforce the law doesn't make the action any less illegal. Criminals who get away with their crimes are still criminals. [0]

> ... if it's unenforceable or sets a precedence that would severely break the business world, there's every chance nobody would dare bring it to a court room.

Or the highest court of the land would find a way to misinterpret "related" historical decisions to make the crime retroactively legal, yeah.

> ...I would have thought if it did, we would all know about the precedence now.

I'm not so sure. For one thing, the courts move really slowly when they're not very motivated to address something. For another, news that paints the twin VC darlings and their "industry" as villainous has a tendency to get buried by any one of a billion hype pieces or minor scandals that they have waiting in the wings.

[0] To bystanders who might wish to retort: Yes, I'm very aware that some things that are illegal should not be. I'm also aware that some things that are not illegal very much should be.


Sorry, I don't understand your point.

Can you state what actual point you're trying to make, rather than just picking each of my sentences and saying you're not so sure or you disagree? I could do the same to you, but it's just a waste of time if we're not trying to actually come to a conclusion together.

My hypothesis is that I suspect there's a MAD type situation where technically - by the letter of the law - all the big tech companies bragging about x% of their code being LLM generated are basically admitting to breaching copyright laws. But if one well funded tech company successfully prosecutes that and sets a precedence, then all big tech is going to have to wind back all the LLM code its put in production in the past few years to prevent litigation, which is probably unworkable, hence MAD.

I was looking to disprove my own hypothesis by asking for court cases prosecuting this. I'm not interested in arguing semantics with you.

(FWIW, I was repeatedly taught by various legal scholars at various levels of education that a criminal is someone who has been found by a court to be guilty of committing an act that violates criminal law. Up until that point, they're usually just a suspect or similar. Not debating, just highlighting that your definition of criminal doesn't negate or disprove any point I'm making. Same for your definition of illegal.)


> Sorry, I don't understand your point.

If you're looking for a single point, that's going cause you trouble. Maybe go back and read what I wrote more carefully? I believe that I took care to make sure that my rebuttals to your claims both dovetailed in nicely with the part of the claim they were rebutting and -as a backup- contained the context needed to understand exactly what I was rebutting.

> I was looking to disprove my own hypothesis by asking for court cases prosecuting this.

_Prosecuting_? If you know enough to ask for that, then you have the knowledge needed to discover that those court cases don't really exist yet. The courts move slowly, especially when wealthy and/or powerful entities are supremely disinterested in attracting their attention.

> I'm not interested in arguing semantics with you.

For anything that's not incredibly clear-cut, that's like the entirety of the practice of law... at least in the US. At its heart, it's an adversarial storytelling exercise that has hundreds to thousands of years of tradition to pull from. I know this in part because I strongly considered becoming a lawyer before I became a programmer.

> ...a criminal is someone who has been found by a court to be guilty...

That's very good practice to engage in when talking about ordinary people. When talking about very powerful people or entities (such as large, influential companies), it's not so great. Very powerful entities tend to be permitted to get away with antisocial conduct that the rest of us get metaphorically nailed to the wall for. "Rules for thee but not for me" is trite, but very, very frequently how things end up working.

EDIT: There's also the example that I had in mind yesterday of US-based marijuana growers, commercial purchasers, individual purchasers, and consumers. All of these people are knowingly illegally producing, trafficking, and/or consuming a Schedule I substance. Folks often do this in the open, making no secret of it. The DoJ and DEA might choose not to haul these people into court, but most -if not all- of them are still very plainly criminals... just not convicted ones.

> My hypothesis is that I suspect there's a MAD type situation...

No. In regards to this whole LLM craze, there's no MAD-type situation between well-funded tech companies. There are a handful of players who have absolutely no interest in stopping the money-making party. No one involved in the money-making party that resulted in the 2008 crash who could make a credible report to the relevant authorities had any interest in calling those authorities in to stop it... why on earth would they when there's so much money being made? And since approximately none of the folks responsible for that huge pile of fraud saw any meaningful punishment, why wouldn't businesses with the means to engage in activity that's no less illegal but is very profitable for a large number of powerful people be discouraged from doing it again?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: