Hacker Newsnew | past | comments | ask | show | jobs | submit | mlmonkey's commentslogin

> What empty, bare land?

Have you been to Death Valley, my friend? All of Southeastern California is empty, barren land. Death Valley, Mojave Desert, etc.


> Have you been to Death Valley, my friend?

Well, no, but I’ve been close and I can imagine some features of the area that might have impacts on cost of construction and ongoing maintenance that are related to why it was empty land even before it was a National Park (and I can think of additional problems with California trying to build solar generation in a national park—different ones under the current federal administration than would be historically normal, but either way it would problematic.)


Guess how fast the wind blows across the surfaces. There's a reason why these places currently host massive wind turbine farms.

They should give credit where credit is due. Solar panels on top of irrigation canals was pioneered in the Indian state of Gujarat back in 2012[1]

[1] https://www.autonocion.com/us/arizona-solar-panels-irrigatio...


You need to credit someone for "hey lets use solar panels to shade things that would benefit from shade"?

They don't mention Gujarat specifically, but they do mention the prior projects in India and Arizona (which your link discusses) in the article.

I noticed Terry Tao is not in this "advisory group" ...

Given that he apparently feels exploited after an appearance in some of their earlier marketing material, I imagine he's not overly keen on working more for their PR department, or this extension thereof.

> During this event, OpenAI requested an interview concerning my vision of the future of AI and mathematics. I accepted, and spoke with them for perhaps an hour. I had done similar interviews in various venues, and I assumed that, as with these other cases, they would eventually post the entire interview online, which talked about both the possibilities and risks of AI much as I have done in these other interviews. As it turned out, they only used a few snippets of that interview for that infamous advertisement instead.

https://terrytao.wordpress.com/2026/09/07/finite-time-blowup...


Why is the domain `pdp1173.com` when the machine is a PDP 11/83 ?

There's a couple eyebrow raising things about this post.. like the fact that there's nothing but vague details in text and one photo with an AI prompt like alt-text. I think they're just trying to get clicks for that banner ad.

Sadly, the site does seem AI-flavoured to a rather off-putting extent.

And, as you say, the details are so vague - it's exactly what you might expect if an LLM wrote them based on reading a spec sheet rather than any practical experience.

In particular, connecting that DEQNA to a modern network must have been a huge challenge all in itself and I'd have expected there to be a series of blog posts about just that. Maybe it's legit (and I'd love more detail if it is), but right now the presentation screams "thinly disguised ad".


I think he has actually posted videos of the restoration before (they've popped up in my youtube feed in the past), but I've never watched them, so dunno. The LLM site design and advertisement front-and-center are certainly a... choice....

But, more importantly, why do you reckon DEQNA would be hard to connect to a modern network? NetBSD still supports it when running on a q-bus VAX: https://github.com/NetBSD/src/blob/trunk/sys/dev/qbus/if_qe....

It has some pretty broken behavior regarding multicast/broadcast traffic IIRC, but nothing putting it behind a vlan or separate LAN altogether wouldn't solve. I've got a little lsi-11 system under my bed I'm hoping to, one day, get the Fuzzball [0] distribution running on. The 22-bit address extension wirewrap the previous owner did needs to be redone entirely, so it's been a "later" project for years now, but I am hoping my little DEQNA will be able to ping the Internet still.

[0] https://en.wikipedia.org/wiki/Fuzzball_router


I was assuming that something of that vintage would only have been able to handle a 10BASE5 transceiver, but a wee bit of googling suggests that an UTP transceiver works well enough.

Needs a patched ROM and recompiled kernel, though! See http://www.cosam.org/computers/dec/pdp11-23/20080402.html (and the two subsequent installments)

(That cosam.org site is more what I'd be expecting for a project like this, but maybe I'm just too old and grumpy for my own good!)


Dave posts frequently about his work, some on YouTube and some on Facebook. You have to be in the right groups on FB.

You and several other people here are making unsubstantiated claims about how you're super sure this is real in a way that makes it seem less real. You guys seem to know this Dave character and are doing him a solid, right?

Dave Plummer's a well known YouTuber. He pops up in a couple of DEC groups that I'm in on Facebook. However I'm not a fan. He seems to take a lot and contribute only self-promotional videos.

The photo on this page looks very AI heavy. I'm not an expert but I really don't think those light panels are real (if they're real they're not contemporary to the machines they're in).

(Edit: I should qualify that I don't mean the panel with the octet switches, that looks roughly right - I mean the one under the 8" disks and the one on the vaguely Vax-pedestal looking thing)


Okay gotcha, sorry I was jumping to conclusions.

Given the AI-heavy page, I don't think your skepticism was unreasonable.

While probably it's legit because I've seen Dave discussing the process on the YT channel (which I no longer watch) he unfortunately does have form https://news.ycombinator.com/item?id=39813625

Having TMOG on a .org when it sells commercial licences doesn't sit well with me given the history.


FWIW, .com was taken so I registered .org, and it's only used for the free version. The PRO version lives at tmog.store, which should be pretty plain.

> The PRO version lives at tmog.store

https://tmog.org/pro.html


Oh, yeah - this popped up on HN yesterday too and also met with scepticism: https://news.ycombinator.com/item?id=49772362

Oh well, so this PDP restoration probably isn't real. Disappointing.


"Probably isn't real". Interesting theory.

Well be real, you could do a lot more to make this convincing. Couple the lack of details with some other posts you've made that seem to be an attempt at advertising your software (the one with a big banner on this page), sprinkle in the AI generated website and the fact that the domain name says a different model than the page does.. this is going to raise some eyebrows.

You put all this work into a project like this, all you would need to do to prove it's real is post a selfie with the hardware.. but "interesting theory" is all you're going to say?


What's not real about it?

I'm a little confused by your comment.. especially considering the fact that you edited it and changed the link from and old comment thread from 2024 to whatever this dead post is from 2012. What does having form mean? Who is Dave?

Sorry, lost the last digit, fixed now. This restored PDP-11/83 project is Dave Plummer's

Because it started as a 15Mhz 11/73 but I later upgraded it to an 18MHz 11/83 CPU. And now it's a Mentec M1, but it's still an "83 class" machine. That's my thinking, anyway.

Anecdotally, I have found the same. I spend a lot of time with these frontier models, brainstorming, etc. and the drop in performance from, say, week 1 to week 8 is often massive. Whereas in the beginning, it seemed like a capable research assistant, by the end of week 8 or so it starts acting like a puppy dog eager to make its 'master' happy for a few treats.

Reminds me of how slot machine users swear the odds have changed on a machine.

also when someone says you just have to prompt it a certain way it reminds me of people who think they can get better results out of a slot machine by pressing buttons in a certain order

The providers of these models also design the UX similarly to slot machines (run it x amount of times for better results, multiplying your spend) this isnt a coincidence and they're playing into the gambler mentality, and probably hire UX designers that specialize in this.


I am skeptical of this as well but slot machines are programmable and the house can change the odds.

Wtf are you on my dude. Anthropic UI is designed like a slot machine? Hiring slot machine specialists? Sometimes I can’t believe im even on HN anymore with comments like this.

I think some of it comes from that they do not publicly let you see the random seed. So each time you ask the answer is different (like a slot machine) and if they let users use the random seed it would let people much more accurately assess if an underlying model changed somehow (same seed and same input will always have the same output).

Of course the closed Anthropic would never share this, it would definitely take away the 'magic' feeling of the AI


With models there are a bunch of other dials that can be tuned even if the model itself remains exactly the same.

Are those dials set the same across all hardware configurations and clusters? Does model behavior average out the same across different hardware?

There are just too many different buttons that can be set to really trust a provider either not to directly commit fraud, or indirectly commit fraud with system complexity affecting the output.


To my understanding, with batched inference and other "optimizations" you wouldn't get the exact same token predictions even with temp=0.0.

Well, running an LLM X amount of times does give you better results provided you are willing to select the best one out of the X yourself.

But I agree with your general point. One of the reasons subscription plans are cheaper because they modulate usage in this way based on demand. They can also recover compute more coarsely via usage resets (which give positive PR).


At that point I might as well do it myself

Well yeah. If for some task you find it easier to just do it yourself then you should. But you can improve the situation even if you can't entirely automate it by automating parts of the verification thereby making it easier to human-do larger verifications when X>1. But in many cases even that is not possible.

The progress however is such that the number of tasks that you can do with >p% automated and X=1 keeps increasing. So many times just waiting works. Of course, here also it changes from field to field. There are some tasks at which AI hasn't even gotten started, others where it has already peaked, others where it's increasing slowly, and others where it's increasing fast.


IMHO (not a paper writer, but read a lot during my grad school years), the Genie is out of the bottle. The only way forward, as I see it, is using LLMs for reviews also. Basically, filter all submitted papers with an LLM and ask it to summarize it, find the biggest weaknesses and main strong points, etc. that a human can then use to review the paper. Basically, LLM-as-a-reviewer .

Personally, I would love to see a conference where people are explicitly encouraged to use LLMs for doing the work and writing the papers, and LLMs are used to review them too.


There should at least be a code of professional conduct where authors state the extent to which LLMs were used. (This would also help not wasting time by asking some “authors” about “their” paper.)

Journals themselves should make policies about the extent to which they allow the use of LLMs. In some areas it might be considered more benign than in others.


There is. NLP conferences, like the ACL family (and the ARR) require you to disclose in the paper if you used LLMs for writing and coding, which ones and how. Whether every author is honest is a different matter.

This already happens. It's clear when a reviewer used an LLM, and it's very annoying for the authors that have to respond to what are usually low quality, superficial reviews.

Boy do I have news for you! :-D

"low quality, superficial reviews" have always been around. Reviewing is most often an unpaid, thankless job and many times reviewers barely put in the effort.


There is still a difference. Previously, if the review was low quality and superficial, it was also quite short and easy to answer. Now I receive fully LLM-written reviews of 2-3 pages (or more) of superficial comments asking for tons of additional text and experiments that are almost always outside the scope of the work. Answering such a review requires a lot of work with absolutely no gain.

I hear you. I just happened to submit a paper to Neurips this year, and the "reviewer" sent back so many questions, etc. that the character limit of the response could never accomodate my response. So I just addressed a few of the main points and called it a day.

I am sure an LLM review could be made much shorter with some prompting.


The character limit thing is real. Another problem is that even if you manage to address all of their points, they're not required to engage with that. They can just ask the LLM to write a basic response back and say they won't change the score.

Similar experience to sibling: a disinterested human isn't going to ask for 30 different detailed things to add to, or change in, the paper that could technically be improved but aren't worth the additional page real estate. An LLM reviewer definitely does this.

To be fair, low quality superficial reviews were also not uncommon before LLMs…

maybe but as a paper writer, the quality of reviews and reviewers have gone down because of LLM-as-a-reviewer too, because LLM reviews seem to regurgitate the limitations section of the paper, and are highly influenced by the way things are phrased in the paper rather than the actual substance.

but I'm hopeful that some middle ground will be found in the future


I dunno man, I use LLM's to review things before I send them in and I often feel like I've been handcuffed to a chair, had a bright light shone in my face and had to account for all the mistakes I've made in my life.

They can be brutal.


Completely agree. I’ve had a Nature paper and a Nature Neuroscience paper published this year and the reviews from Opus were much more thorough and complete than those of the human reviewers. I still benefitted from the human reviewers, but now I consider an LLM review essential.

I think that’s a lot of risk of anchoring reviewer bias. I’d be more comfortable with a triaged review where the editor’s office uses models to score whether a human editor should evaluate a paper to potentially send out for review, then the editor makes their own assessment, and the reviewers continue to do their job unassisted

The entire point of an academic paper is to add to the sum of human knowledge. How can an LLM trained on a subset of human knowledge possibly even begin to accurate evaluate such a paper?

I trust an LLM to review that the language used in the paper is grammatically correct, but not to evaluate new information for accuracy.


This is very bad logic

1) Humans also are trained on a subset of human knowledge. 2)A lot of papers are just about experimenting something, and then applying simple stats. Eg empirical studies, around 1/3rd of published papers. Like, we tried this drug or did this experiment, from a sample size X here are the results. An expert is needed to maybe comment on the conclusion/hypothesis of the underlying suspected mechanism, but LLMs are still very useful on catching bad statistics or p hacking (so so common)


Schmidthuber has an answer: https://arxiv.org/abs/0812.4360

For a more practical approach you need to use proxies: https://zby.github.io/commonplace/articles/what-an-automated...


Yes. I want reproducability, open-sourcing, accessibility, correctness, and most of all: usefulness. I don't care how it was written or reviewed, as long as some assurances regarding above things can be made, and I don't see why LLMs would get in the way of that.

What can't be gotten rid of fast enough is the notion that having written something is meaningful on its own. Making something that looks right was a level above total novice: now it's the floor.


Here is my proposal (or actually our proposal - me and my agents): https://zby.github.io/commonplace/articles/what-an-automated...

Extend this beyond review. The most value we would get is from quality checking existing published papers.

Nope, I’ve tried this, it’s awful.

For a start LLMs love LLM generated text, so you are boosting papers people never had any input in.

Secondly, LLMs in my experience are good at small issues, but fail totally at the whole paper being obviously poorly constructed, or clearly fake.


> Let's see if it (or crossbreeds) can be tamed.

Clearly you are not a cat person ... the question is not if it can be tamed, but how quickly it can tame the humans around it ... ;-)


All along I had thought that "AGI", "RSI", etc. were at the model level: but this paper seems to be talking about "agents", etc. I'm not sure having a swarm of agents explore a problem space in parallel via brute force is what "AGI" is about. I'd be happy to be proven wrong.

AGI and RSI are both meaningless terms, meaning whatever you choose them to mean.

RSI is the new sexy. Models are RSI-ing themselves towards the singularity, these folks' agents are RSI-ing themselves towards mastery of their training environments, and my pet cat is RSI-ing himself into the best cat that he can be.


> Ask yourself, how does Google -- a company that famously does everything -- benefit from SWEs outside of Google having access to powerful coding models?

Catch is, even Googlers internally do not have access to top-tier models. (or did not until recently, when apparently Claude was made accessible to the SWEs internally).


Googlers now have internal access to frontier models.

Funny enough, after trying it, I went back to G3.8


Yes, agreed, 3.8 with thinking is quite good.

The goal is to kill us all .... j/k. :-)

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: