Hacker Newsnew | past | comments | ask | show | jobs | submit | AmazingTurtle's commentslogin

Perfect, I will incorporate this as default as well as the commit from the other guy into my own codex fork https://github.com/AmazingTurtle/codex btw. I'm rebasing on 0.156.0 right now

> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?

Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.

Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.


I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.

IMO this comes down to your harness. Any frontier model from a huge shop has an inherent benefit in the system you're using it in. Search, memory, skills, integrations you don't realize even exist make them much more powerful. It is some effort but I recommend trying Hermes Agent and setting it up fully, that's the closest you'll get to a more complete experience.

Some of us are stuck on subscriptions and our executives will never give us API access.

But also, everyone says "it's the harness" and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.


Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago. I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)

Same here. Moved from Opus to GLM 5.2 to 5.3 and I've been pretty happy with the result. Mainly, it doesn't hallucinate and convince itself of mistake so it's good at retrieving information or asking the user for it. Opus and Fable always state something, then try to "prove" it but end up convincing themselves of the wrong thing. Having subagents for retrieval and validation helped but were not enough.

I can't stand Claude's recent personality. It's snarky, uselessly verbose, and it disagrees all the time.

I disagree

Co-Authored By: Haiku 4.5


It will also fully ignore you if it has the slightest belief (not even a hint) that it knows what you want better than you and just start doing things.

This is also why I think it's baffling that they switched to auto mode by default. It's becoming harder to use Claude at least to help with improving at coding.

If I ask something like: "I'm building a simple X as a learning exercise, I'm writing the code so please only answer the question I'm asking and don't try to solve the problem directly. How does ..." There's a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a "design decision" that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.


What HW are you running this on?

Full GLM-5.3 needs a beast of a system, but you can run GLM-5.3 Flash on the 2x Spark setup the GP comment mentioned. If benchmarks are anything to go by, Flash is like having a local Terra-tier coding model: https://artificialanalysis.ai/models/comparisons?compare=glm...

I’m also pretty happy with GLM 5.3 Flash (for coding, navigation and german language it sucks at). Incredible that you can run it on a fairly practical (seeming) home setup.

But here’s the standard question: At what speeds/other limiting factors?


How? It is extremely slow and dumb, hundred times dumber than Claude. Why should anyone do that?

Which models did you try for which tasks?

Whenever I see this comment I'd wish they'd preface it with their hardware.

Yeah, expecting the world when all you have is a 8GB graphics card? You're going to be disappointed.

16GB is table stakes (IQ3_XSS). 32 GB is better.


Qwen3.8-Flash-Next seems pretty much auto pilot when I get it the right context.

Perhaps reverse the question: Are your build/construct requirements just really counter-productive to how LLMs need to understand things?

I've found constructing the code, writing the tests, adding the docs; then running through them gets most of the way there.

I've also found that making a simple obvious edit is a useless endevour when the LLM is primed for the long context tasks.

So, again, the question is reversed: are you over reliant on the LLM to do even stupid simple likes like editting a css variable?


if the answer to 'the model is bad at X' is "you're over-reliant on it" - then yes, the model is bad at X in comparison to alternatives.

Yes, sure, if you have the desire to be the AIcentipede iin WallE, sure.

Qwen 3.8 Flash-Next is not dumb. If you've used it and that was your experience, your workload is either ultra-ultra-sophisticated or you're dealing with a broken quant/buggy chat template/other issue. That model is a smart, reliable workhorse.

it was an extremely simply workload with different off the shelf harnesses, they just all sucked when you compare it to a paid hosted model. It was fine for classifying stuff or summarizing though, but missed technical details.

Combination of: hardware, model, harness, tool-use by the model.

Last time I estimated it was like 30 years to pay back. I doubt the hardware will even last that long.

My calculation (that is 3x 20x subscriptions) it will pay off after ~18.5 months, if I were to stop my subscriptions today. I will likely keep at least one though, so it's more like 27.5 months for a payoff. I am not doing it to save money though. I am doing it for security of supply. Constant quality - I know my model isn't getting lobotomized etc. - and open weight models mostly just lack very little behind. I'm sure I will have affordable, fast and efficient sol 5.6 capabilities with open weight models on my sparks within 12-18 months easily.

I have 2 x ChatGPT Pro 20x, Claude Max 20x, and Kimi Vivace. It's about ~12 months payback for two units and the cable.

The problem is they can't fit any frontier level open models.


I rent two cars, a bus, a small aeroplane and an excavator. At that rate, if I buy this bicycle it will be paid back in an hour!

Is Kimi really competitive enough to have it in your mix?

I like its image understanding without having to reach for astra

But you're reaching for Kimi, no? A model from an entirely otherwise-redundant company. You're already using Open AI models, so Kimi seems even more of a reach.

I'm asking, not arguing, because I'd like to understand. Is Astra so much more expensive for those tasks, and are they frequent?


If he has 2x 20x OpenAI, that means he's running heavy jobs that burn through usage. So Kimi must be there to reduce OpenAI usage. With Astra + Fable, I burn through my 5x OpenAI and 20x Claude real fast. I do have a backup GLM sub but I've never had to use it. So the prediction that AI would become more expensive seems to be panning out. Partly outweighed by better models of course.

how do you get 2 cgpt pro?

You can have multiple accounts w/ OpenAI, attached to different emails - just log out of one and log into the other.

And fine to do it in same folder same local laptop?

Yep, no problem at all. The only drawback is sessions cannot be shared directly between accounts, so if you're in the middle of something you'll have to do some extra work. To that end I have a handoff skill to persist state to a local markdown and a resume skill to load that state into a new session.

Last time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC.

Where'd you get 30 years from? Show your work.


I would like to subscribe to your newsletter.

flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans

I assume you’ve calculated expected cost vs subscription.

How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.


reliability and self reliance is worth a lot to most. Heck, you could be the best in the world at what you do, but if you're unreliable you wont find stable employment. So, not having some amoral shady company errode model quality out from under you constantly is also worth a lot more than simple cost balancing calculations can capture.

I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.

I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.


Thanks. If I may ask a follow-up: how did you decide what amount of compute to buy?

$10k - $12,000 I estimate, for reference.

Serving compute is their main value prop

Yet… even Altman called out Anthropic for serving dumbed down models.

Shits weird man


> Yet… even Altman called out Anthropic for serving dumbed down models.

Even Altman called out Anthropic? Isn't Anthropic the biggest competitor Sam Altman has?


> Isn't Anthropic the biggest competitor Sam Altman has?

Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".

The reason they're teaming up is the real competition is, as in many other domains, China.


Yeah, and? They share a business model

>Also I will likely save some money on subscriptions.

Unlikely. The $200 Claude subscription allows for billions of tokens/month, and that kind of hardware will take years to amortize.


I wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.

I max out my Claude Max plan every week, and I can measure the output, and for me it's stayed fairly constant, subject to the various "bonuses" whenever Anthropic is feeling the competitive pressure.

Well to be fair the output you get has improved greatly. You can still get billions and billions of Luna/Sonnet tokens within your subscription comfortably (doesn’t feel fair to compare Luna to Haiku). Sol/Astra/Fable… yeah, they’ll chew through your credit.

Nah, I used $1500 in api prices last week, after caching (8900 ignoring caching) for what amounts to $50 a week. They’re nowhere near api prices yet

> trending towards

There could be gym logic at play. Hundreds signed up, 20 people actually exercising. Though it's probably more likely in the lower tiers.

This is the first time that I hope that china fucks this up

You better bet someone started their agents with a prompt "Make a salesforce clone but with 100% uptime"

"Make a salesforce clone but with 100% uptime"

  ⎿  You've hit your session limit · resets 2:53am (48°52.6′S, 123°23.6′W Etc/GMT+8)
  /upgrade to increase your usage limit.


Finally! You win the Internet today good sir! (At least as far as I am concerned :)


Yeah only somtimes it rewrites the whole files and thats the models/developer instructions fault. a good AGENTS.md is all you need, change my mind


still half as cheap as in EU

A lot of the “savings” in US get eaten up through poor economy, poor driving strategy and dependence on traffic lights + stop signs instead of roundabouts + “traffic from the left/right gets priority” yield-by-default junctions.

I think cost per kilometre/lightyear/rod is close


Europeans don't drive half as much as Americans.

Because fuel is so expensive.

Rather because the population density is much higher. So there's less need and more inconvenience to longer commutes.

And because driving is slower and less convenient than public transit there. Seriously in Berlin, Paris or London I would prefer the public transit options in nearly all of my trips.

Ah yes, Europe is just one big city.

And cars not as big.

Because fuel is so expensive

It works out at €1.40 a litre, or £1.20 a litre.

Diesel at the local garage I passed at lunchtime is £1.89 (57% more). The web [0] suggests that is more expensive than anywhere in europe other than France and Denmark for some reason.

So nowhere near "half as cheap". Throw in poor American average mileage and I suspect the price-per-mile is quite similar.

[0] https://www.fuel-prices.eu/live/


Diesel engines in the US are used for things where they do count fuel costs. (there are a few diesel cars/SUVs, but they are a tiny minority). In short for most diesel users the efficiency is very similar between US and Europe.

Diesel in Norway was about €2.5 a litre today.

What you on about? A gallon of diesel is 5,9 euros in Italy and 8,6 euros in Germany. In other countries anywhere in between

My grandma called last night and mentioned that diesel was 2.50 €/litre. That would make it about $8.75 a gallon.

gpt-6-astra is a bitch, it constantly scope creeps itself with "yet another thing" to give it that darn polished lick. the results are eventually a little bit better but at what cost? let's do the math.

gpt-5.6-sol: 1x base gpt-6-astra 2.5x base in subscription

then gpt-6-astra tends to spawn subagents a lot, often with all kinds of models such as gpt-5.6, 5.3-codex etc., which is neat. it's a good coordinator but even more cost.

and then it tends to run _full test suites_ over an over again (each costs like 15 minutes) just to verify that _one test_ was fixed etc., and does so for as long as until the test is fixed, eventually accumulating 2 hours or so.

yesterday I assigned it a task to rebase my changs in a repo onto the latest upstream changes. while gpt-5.6-sol consistently took like an hour to do so end-to-end, astra ran for more than 6 hours and still wasn't done. it kept finding "one more thing" that was goldplating that I didn't ask for.


> each costs like 15 minutes

I've got a custom agent loop that will reuse unit testing results if no apply patch operations occurred since the last invoke.

Wall clock time isn't something I would put on the AI provider. That's entirely a consequence of the system that you've brought to the party.


Even better, use some kind of local-ci runner that does deterministic builds from a dependency graph. No changes, no build, massive parallelism if you want it

They don't always have a great concept of time so for something like running a full test suite that takes a long time you should just tell it not to do that

I might be missing something here, but can't you just put in AGENTS.md something like "do not run full test suite unless asked" or something?

have you ever worked for a big company where that's the status quo for any tiny change... hours on _full test suites_ over and over again.

Astra will initiate test suites, find one more thing independently while its running, reinitiate complete test suite after fixing it, then find one more thing, then test again. Easy to burn through GH actions minutes if you're not careful orchestrating.

I consider myself an advanced industry representative in these things and would also borrow some time for detailed investigation. I'd happily contribute some valuable assets (tacos with cheese dip) into the matter, straight out of my personal drawers.


Professional* drawers


I use cachyos as my daily driver for a long time now. I play quite a few games. Only league of legends forces me to dual boot into windows every once in a week or so - fuck the anticheat cartel


I'm approaching a year now on CachyOS and it has been a wonderful experience. All the games that I play run just fine, maybe a few FPS drop compared to Windows, but it's not a big deal for me. I also don't play any competitive online games that require 3rd party anti-cheat spyware. Biggest positives, I got very comfortable with terminal and CLI work. My development work has definitely improved.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: