Gemini 3.8 Flash looks like its better than v7 Luna/Sol on DeepSWE v1.1 while at $0.75 per million input tokens and $3.75 per million output tokens. Luna is much cheaper, but Flash has nearly Astra's performance for under the price of Sol ($2/$10).
I use it for golang, and it is fantastic. Incredibly fast. It seems the llm and I “understand” each other. I have to be less careful in my exact phrasing. It kind of just does what I want and expect.
When I ask for an explanation it adds the right amount of detail. Of course, some of the material is new to me so subtle errors are hard to spot. But at least I’ve caught Terra and Sol on inconsistent messaging.
Also I’ve found 3.8 flash to circle back to root issues even at the conceptual level like problem fit and conceptual solution direction or architecture when I wasn’t achieving my goals. It flat out said I was attempting to use the wrong tool. Whereas Sol and Astra kept rabbit holing and looking for tiny implementation errors. Even after prompting them specifically to look at it broader.
If you don’t have a plan yet, log in to antigravity. There will be a button “upgrade plan” somewhere. Sometimes it pops up and otherwise lookup in settings > account. There should be some button that says upgrade. Clicking that brought me to the google studio ai page which offered the 20-something plan for €5/month.
TBH: I really like how fast 3.8 Flash is... Once I have clear plan, I feel quite confident in delegating large parts of implementation to Flash and Luna
Maddening for a bit - I've got problems that Flash is better on, and some Luna is better on, but I generally don't know until one has wasted time/tokens. Then I switch to the other one and... it's often just... bam - done. Correctly. I can't find the patterns ahead of time to determine what model I should be using first. :/ That said, I've been alternating between both the last month or so and they've both been pretty good compared to earlier models.
Gemini 3.8 Flash and 3.1 Pro are pure rubbish. Very little thinking, mediocre and usually incorrect results. They cannot be compared to frontier models.
I suspect that the people saying this haven't used Luna.
It's also weird that anyone uses it outside of an enterprise. They force you to use Googles inferior harness on the plans and I doubt any mere mortal is paying that much, for so little usage, with the worst harness on the market.
I prefer 5.6 Luna while a coworker prefers 3.8 Flash. The difference seems to be that they chat with Flash (with code context) while I just ask Luna to directly modify the code. I was already very impressed with 5.6 Luna so I am looking forward to running 6.0 Luna all day tomorrow to see how it compares.
I do like Typst, but Typst isn't really ideal for diagrams... you could use it that way but I don't think it's ideal. Mermaid or PlantUML are better/cleaner options.
Agreed. My resume and slides are all templated with Typst, but I can't be bothered trying UML languages to make diagrams that'll probably be referenced once or twice a year, if ever.
This is the coolest thing on HN that I've seen in a minute. The whole thing is a perfect blend of ideas with the end result of something that feels... magical. Honestly this is the highest inspiration to me as a builder as I just want to create magical experiences, even if they are just small oddities.
That's the e-ink monopoly for you. Though you could probably use a slightly smaller b&w eink display with a second hand raspberry 4 to shave a lot from that number.
It's been around in slang for a few years with a few variants, e.g. "I haven't seen her in a hot minute." I think it's just filtered into wider cultural vernacular.
I had not heard it before (or if I did I did not interpret it correctly, eg if one told me "see you in a minute" I would interpret it as "see you in a bit".
But there is this podcast discussing it in 2021, and it comes from black community slang from 70s (which is where most slang I encounter comes from). And these terms take a while to catch up usually, but it seems it was already circulating more broadly since 2000s.
As a founder of a few software startups, I agree with this. If I can give advice to other founders, remember than software is usually a means to an end, and for most users, they just want the bare bone essentials to work and work well. Everything else is nice-to-have-fluff that won't drive 99% of sales.
The apps I pay for are simple and focus on a super basic interface. BookFusion and Libro.fm are great examples.
It’s interesting because my experience is inverse (b2b software) - the bloated shiny solution what actually doesnt work at all performs better in the market than the sober, unsexy solution that just works.
Companies like Oracle or Microsoft are so massive not because they sell simple and focused products that work well; quite the opposite
Might be different when you’re selling to endusers
Oracle (and SAP and IBM) are not in the business of building software, they are in the business of never complete software, instead stretch the project as far as you can and bill to death.
Runs the 3bit model faster than the 2bit one runs on my old-ass card. Can’t vouch for its intelligence yet, but i suspect whatever loss in smarts it takes is made up for by the extra resolution.
I've seen some product prototypes using Iroh and it appears pretty good for "just get edge device / node connectivity to work" when you don't want end users to bother with networking shenanigans but can guarantee Internet connectivity.
Basically Tailscale but embedded into the app without the hassle of requiring users to setup accounts.
I really wish Moonlight would use Iroh. I use Tailscale to make it work, but it would be so much easier to just pair a machine and have it always work.
What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.
They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.
Only relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.
These models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
reply