Hacker Newsnew | past | comments | ask | show | jobs | submit | irthomasthomas's commentslogin

He was kicked out of Canada because they thought the Technocracy movement which he promoted was too communist.

And Claude often identifies as Qwen or Deepseek when prompted in Chinese.

Except for all the exceptions, which are many, like the Jack the Ripper police files which where denied to the public.

They're just hiding misconduct because the government does an enormous number of things that would be national scandals if they weren't allowed to bury them.

I doub't it. The main issue is not cost, though they do get expensive as context grows, but intelligence. A frontier model like fable becomes as dumb as haiku after 200k tokens. They have been stuck at ~1M context/200k useful context for 18 months, now, with little sign of advancement. A model with a 10M context window that retains it's intelligence up to 2M tokens would be a big breakthrough.

> A model with a 10M context window that retains it's intelligence up to 2M tokens would be a big breakthrough.

so the true next frontier might not be just raw intelligence but rather larger context window?


Or better context curation - less lossy compression saving back to context. Maybe even jettisoning part context into an external semantic store instead of conpression. Or placing less data into context to start with.

Or a combination of all those things.


What are your favourite quotes from it?

Let me be clear, he's no cormac mccarthy(my favorite prose is the passage from blood meridian describing the raiders). My bar for fan fiction prose is basically whether it annoys me, its a low bar. Most of what makes hpmor good is the cleverness with which yudkowsky subverts the original harry potter world to try to make a rationalist argument, whether that be by inverting a character or playing off the harry/voldemort relationship in thoughtful ways.

'"Voldemort," said the old wizard. "I understand him now at last. Because to believe that the world is truly like that, you must believe there is no justice in it, that it is woven of darkness at its core. I asked you why he became a monster, and you could give no reason. And if I could ask him, I suppose, his answer would be: Why not?

They stood there gazing into each other's eyes, the old wizard in his robes, and the young boy with the lightning-bolt scar on his forehead.

"Tell me, Harry," said the old wizard, "will you become a monster?"

"No," said the boy, an iron certainty in his voice.

"Why not?" said the old wizard.'

as you can see its nothing special, just good enough


I see, thanks. It reminds me a little of those spam emails which are sprinkled, intentionally, with obvious spelling errors. They aren't trying to trick the average person, they are trying to filter for a much smaller, more valuable audience.

fwiw, I don't think that was intentional here. He's just more interested in "ideas" than prose or people. A more extreme example of Andy Weir. Definitely a niche. I read it years ago while backpacking and it perhaps needed an editor but the overall struggle/war and reimagining of logical fallacies of the HP universe's magic were fun.

yes exactly. Fan fiction readers largely do not care about grammar or nice prose, just has to not be overly distracting which unfortunately most fan fic fails at

It is getting a lot harder for those people to justify using OpenAI to assist such endeavours. Afterall, OpenAI might just front-run you if they hear a rumor you solved some marquee problem that they can brag about in PR campaigns.

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.


It's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...

It took them two years to finally get him out with the help of the government in Shenzhen

https://www.nme.com/news/arm-china-finally-ousts-rogue-ceo-t...


And this is why building governance out of a pack of cards is a risky business...

If you cannot escalate a governance issue to a higher power, the controls are not worth the paper they are written on.


Wait China pulled off a coup inside ARM?

> For instance the task naming in the task file starts with an optimistic 1, 2, 3, 5, 5a but then eventually gets to 8a, 8a1, and then ends up with 8b2c2b3 and “8b2c2b2b checkpoint1”. The code that it produced got ever more wild. I don’t want to bore you with what it tried to build, but here are some example pieces of the interpreter changes:

  Hardcoded constants everywhere
  Multiple same-line macro invocations in C
  Random indexes in production code
  Hideous tokenizer code in C
https://lucumr.pocoo.org/2026/9/7/astra-why/

Astra scores the same on DeepSWE 1.1 (~75%) as Gemini Flash 3.8 and Deeepseek Flash 4.1 So general coding ability has plateaued, for now. Also consider the context windows. 1M token models where a breakthrough two years ago. Today they are still limited to 1M. In fact, if you don't want intelligence to drop off a cliff, you are really limited to 200k tokens.

Gemini Flash is a joke for coding. If you can get the same output as you can get with Sol/Astra I'm impressed. Not to mention that Antigravity is awful.

It is not a universal opinion at all that general coding ability has plateaued.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: