Perfect, I will incorporate this as default as well as the commit from the other guy into my own codex fork https://github.com/AmazingTurtle/codex btw. I'm rebasing on 0.156.0 right now
> Could there be a benefit to releasing a new model, slowly dumbing it down over a couple months, then releasing a new model that’s marginally if at all better than the original to create a perceived improvement when in reality there isn’t really one?
Exactly what I am saying for months now. And it's exactly the reason why I am shifting to open weight models now. Just bought myself a 2x DGX Spark Cluster. Will run Qwen3.8 Flash Next on it, maybe Qwen4 when it comes out.
Not only do I have full control over quantization and inference, but also will I experience a constant level of quality. It won't be frontier. But it will be stable, and that's enough reason for me to switch. Also I will likely save some money on subscriptions.
I don’t know what people do with the open models but having tried a lot of them I just can’t make it make sense. they’re too dumb and it effectively makes them useless (to me). it’s probably worth being honest about the low ceiling here.
IMO this comes down to your harness. Any frontier model from a huge shop has an inherent benefit in the system you're using it in. Search, memory, skills, integrations you don't realize even exist make them much more powerful. It is some effort but I recommend trying Hermes Agent and setting it up fully, that's the closest you'll get to a more complete experience.
Some of us are stuck on subscriptions and our executives will never give us API access.
But also, everyone says "it's the harness" and almost nobody ever gives good examples, it gets a bit tiring to read everywhere, as if everyone wants to sell a harness to us.
Really!? Glm5.3 is my daily driver and I feel im having the most productive experience with agentic collaborations so far, by a lot. Using pi with tons of custom extensions, that to be fair I developed since making the jump off of codex and claude about 12 weeks ago.
I primarily do not write code for a living. I do a lot of modeling and commercial analysis and a lot of math (related to differentiable simulation)
Same here. Moved from Opus to GLM 5.2 to 5.3 and I've been pretty happy with the result. Mainly, it doesn't hallucinate and convince itself of mistake so it's good at retrieving information or asking the user for it. Opus and Fable always state something, then try to "prove" it but end up convincing themselves of the wrong thing. Having subagents for retrieval and validation helped but were not enough.
It will also fully ignore you if it has the slightest belief (not even a hint) that it knows what you want better than you and just start doing things.
This is also why I think it's baffling that they switched to auto mode by default. It's becoming harder to use Claude at least to help with improving at coding.
If I ask something like:
"I'm building a simple X as a learning exercise, I'm writing the code so please only answer the question I'm asking and don't try to solve the problem directly. How does ..."
There's a 30% chance it starts reading and writing code immediately and a 20% chance it argues with a "design decision" that will bite me in the non-existent future of my learning exercise. If I ask a follow up question, naively assuming that the context from my original question still stands without repeating, it will almost assuredly start making modifications to my code.
Full GLM-5.3 needs a beast of a system, but you can run GLM-5.3 Flash on the 2x Spark setup the GP comment mentioned. If benchmarks are anything to go by, Flash is like having a local Terra-tier coding model: https://artificialanalysis.ai/models/comparisons?compare=glm...
I’m also pretty happy with GLM 5.3 Flash (for coding, navigation and german language it sucks at). Incredible that you can run it on a fairly practical (seeming) home setup.
But here’s the standard question: At what speeds/other limiting factors?
Qwen 3.8 Flash-Next is not dumb. If you've used it and that was your experience, your workload is either ultra-ultra-sophisticated or you're dealing with a broken quant/buggy chat template/other issue. That model is a smart, reliable workhorse.
it was an extremely simply workload with different off the shelf harnesses, they just all sucked when you compare it to a paid hosted model. It was fine for classifying stuff or summarizing though, but missed technical details.
My calculation (that is 3x 20x subscriptions) it will pay off after ~18.5 months, if I were to stop my subscriptions today. I will likely keep at least one though, so it's more like 27.5 months for a payoff. I am not doing it to save money though. I am doing it for security of supply. Constant quality - I know my model isn't getting lobotomized etc. - and open weight models mostly just lack very little behind. I'm sure I will have affordable, fast and efficient sol 5.6 capabilities with open weight models on my sparks within 12-18 months easily.
But you're reaching for Kimi, no? A model from an entirely otherwise-redundant company. You're already using Open AI models, so Kimi seems even more of a reach.
I'm asking, not arguing, because I'd like to understand. Is Astra so much more expensive for those tasks, and are they frequent?
If he has 2x 20x OpenAI, that means he's running heavy jobs that burn through usage. So Kimi must be there to reduce OpenAI usage. With Astra + Fable, I burn through my 5x OpenAI and 20x Claude real fast. I do have a backup GLM sub but I've never had to use it. So the prediction that AI would become more expensive seems to be panning out. Partly outweighed by better models of course.
Yep, no problem at all. The only drawback is sessions cannot be shared directly between accounts, so if you're in the middle of something you'll have to do some extra work. To that end I have a handoff skill to persist state to a local markdown and a resume skill to load that state into a new session.
Last time I estimated, it would only take 3 months to pay back because the 1TB Mac Mini running Qwen RSIingly developed ASI and made infinity dollars off of crypto and I got put in jail by the SEC.
flash next is good, I've been running it for like 2 weeks now and it's pretty solid, hope you like it and it meets your needs. I still lean on Claude and codex a fair bit for harder stuff, but I'm rapidly moving towards 2x $20 plans instead of 2x $200 plans
I assume you’ve calculated expected cost vs subscription.
How do the numbers pan out? Ack that it isn’t always juts about cost, so even if it is pricier to self/host it might still be better for you for other reasons.
reliability and self reliance is worth a lot to most. Heck, you could be the best in the world at what you do, but if you're unreliable you wont find stable employment. So, not having some amoral shady company errode model quality out from under you constantly is also worth a lot more than simple cost balancing calculations can capture.
I'm getting really sick of the constant rot and "magic breakthrough" cycle, so im going full local, at expense on paper but being able to trust something which I need to understand the reliability of is priceless.
I like predictable. I'll take slightly less capable over unreliably capable, since with reliable i can calibrate my expectations and learn what aspects of my workflows to entrust and trust it will work. You simply cannot do that with models you don't control and in my experience they will all errode after the initual marketing wave passes, likely you eventually get fed heavily quantized versions and are expected to accept degraded service when what convinced you to pay was a far superior product. No such issues with local.
> Isn't Anthropic the biggest competitor Sam Altman has?
Not by a mile. They're even (probably illegally and there's apparently a class action lawsuit oncoming: at least something to that extent was posted on HN today) teaming up, as a duopoly, to push for the same bullshit regulations / "we need to slow down AI research".
The reason they're teaming up is the real competition is, as in many other domains, China.
I wouldn't be so sure. The generosity of the subscription plans has declined GREATLY over the past 6 months or so. They are likely trending towards api pricing parity. In which case, having your own hardware makes sense if you can utilize it well.
I max out my Claude Max plan every week, and I can measure the output, and for me it's stayed fairly constant, subject to the various "bonuses" whenever Anthropic is feeling the competitive pressure.
Well to be fair the output you get has improved greatly. You can still get billions and billions of Luna/Sonnet tokens within your subscription comfortably (doesn’t feel fair to compare Luna to Haiku). Sol/Astra/Fable… yeah, they’ll chew through your credit.
A lot of the “savings” in US get eaten up through poor economy, poor driving strategy and dependence on traffic lights + stop signs instead of roundabouts + “traffic from the left/right gets priority” yield-by-default junctions.
And because driving is slower and less convenient than public transit there. Seriously in Berlin, Paris or London I would prefer the public transit options in nearly all of my trips.
Diesel at the local garage I passed at lunchtime is £1.89 (57% more). The web [0] suggests that is more expensive than anywhere in europe other than France and Denmark for some reason.
So nowhere near "half as cheap". Throw in poor American average mileage and I suspect the price-per-mile is quite similar.
Diesel engines in the US are used for things where they do count fuel costs. (there are a few diesel cars/SUVs, but they are a tiny minority). In short for most diesel users the efficiency is very similar between US and Europe.
gpt-6-astra is a bitch, it constantly scope creeps itself with "yet another thing" to give it that darn polished lick. the results are eventually a little bit better but at what cost? let's do the math.
gpt-5.6-sol: 1x base
gpt-6-astra 2.5x base in subscription
then gpt-6-astra tends to spawn subagents a lot, often with all kinds of models such as gpt-5.6, 5.3-codex etc., which is neat. it's a good coordinator but even more cost.
and then it tends to run _full test suites_ over an over again (each costs like 15 minutes) just to verify that _one test_ was fixed etc., and does so for as long as until the test is fixed, eventually accumulating 2 hours or so.
yesterday I assigned it a task to rebase my changs in a repo onto the latest upstream changes. while gpt-5.6-sol consistently took like an hour to do so end-to-end, astra ran for more than 6 hours and still wasn't done. it kept finding "one more thing" that was goldplating that I didn't ask for.
Even better, use some kind of local-ci runner that does deterministic builds from a dependency graph. No changes, no build, massive parallelism if you want it
They don't always have a great concept of time so for something like running a full test suite that takes a long time you should just tell it not to do that
Astra will initiate test suites, find one more thing independently while its running, reinitiate complete test suite after fixing it, then find one more thing, then test again. Easy to burn through GH actions minutes if you're not careful orchestrating.
I consider myself an advanced industry representative in these things and would also borrow some time for detailed investigation. I'd happily contribute some valuable assets (tacos with cheese dip) into the matter, straight out of my personal drawers.
I use cachyos as my daily driver for a long time now. I play quite a few games. Only league of legends forces me to dual boot into windows every once in a week or so - fuck the anticheat cartel
I'm approaching a year now on CachyOS and it has been a wonderful experience. All the games that I play run just fine, maybe a few FPS drop compared to Windows, but it's not a big deal for me. I also don't play any competitive online games that require 3rd party anti-cheat spyware. Biggest positives, I got very comfortable with terminal and CLI work. My development work has definitely improved.
reply