Hacker Newsnew | past | comments | ask | show | jobs | submit | jadbox's commentslogin

Gemini 3.8 Flash looks like its better than v7 Luna/Sol on DeepSWE v1.1 while at $0.75 per million input tokens and $3.75 per million output tokens. Luna is much cheaper, but Flash has nearly Astra's performance for under the price of Sol ($2/$10).

Flash thinks much more so it’s pretty much line with Sol for performance. That said I like flash coding style much more than OpenAi models.

out of curiosity, what type of code/language do you usually use flash to write?

I use it for golang, and it is fantastic. Incredibly fast. It seems the llm and I “understand” each other. I have to be less careful in my exact phrasing. It kind of just does what I want and expect.

When I ask for an explanation it adds the right amount of detail. Of course, some of the material is new to me so subtle errors are hard to spot. But at least I’ve caught Terra and Sol on inconsistent messaging.

Also I’ve found 3.8 flash to circle back to root issues even at the conceptual level like problem fit and conceptual solution direction or architecture when I wasn’t achieving my goals. It flat out said I was attempting to use the wrong tool. Whereas Sol and Astra kept rabbit holing and looking for tiny implementation errors. Even after prompting them specifically to look at it broader.


what harness or plan are you using 3.8 flash with?

I’m using antigravity. I’m still on the AI Pro plan for the promotional $5/month.

where is this promotion?

If you don’t have a plan yet, log in to antigravity. There will be a button “upgrade plan” somewhere. Sometimes it pops up and otherwise lookup in settings > account. There should be some button that says upgrade. Clicking that brought me to the google studio ai page which offered the 20-something plan for €5/month.

anti-gravity with gemini 3.8 or gtfo

TBH: I really like how fast 3.8 Flash is... Once I have clear plan, I feel quite confident in delegating large parts of implementation to Flash and Luna

Maddening for a bit - I've got problems that Flash is better on, and some Luna is better on, but I generally don't know until one has wasted time/tokens. Then I switch to the other one and... it's often just... bam - done. Correctly. I can't find the patterns ahead of time to determine what model I should be using first. :/ That said, I've been alternating between both the last month or so and they've both been pretty good compared to earlier models.

codex seems pretty solid on review, flash is fast on basics but makes more mistakes/errors, Claude is just to picky for me

Gemini 3.8 Flash and 3.1 Pro are pure rubbish. Very little thinking, mediocre and usually incorrect results. They cannot be compared to frontier models.

This is my experience as well. I am surprised that lot of people find it much better than Luna.

I suspect that the people saying this haven't used Luna.

It's also weird that anyone uses it outside of an enterprise. They force you to use Googles inferior harness on the plans and I doubt any mere mortal is paying that much, for so little usage, with the worst harness on the market.


I prefer 5.6 Luna while a coworker prefers 3.8 Flash. The difference seems to be that they chat with Flash (with code context) while I just ask Luna to directly modify the code. I was already very impressed with 5.6 Luna so I am looking forward to running 6.0 Luna all day tomorrow to see how it compares.

look at token use, 3.8 flash is a huge token hog compared to openai models

I do like Typst, but Typst isn't really ideal for diagrams... you could use it that way but I don't think it's ideal. Mermaid or PlantUML are better/cleaner options.

Agreed. My resume and slides are all templated with Typst, but I can't be bothered trying UML languages to make diagrams that'll probably be referenced once or twice a year, if ever.

This is the coolest thing on HN that I've seen in a minute. The whole thing is a perfect blend of ideas with the end result of something that feels... magical. Honestly this is the highest inspiration to me as a builder as I just want to create magical experiences, even if they are just small oddities.

Really cool, except the price. Around 500 Euro for the full thing... I miss the times you could make amazing things with a Pi under 100 Euro.

That's the e-ink monopoly for you. Though you could probably use a slightly smaller b&w eink display with a second hand raspberry 4 to shave a lot from that number.

To be fair even if they could build a 10x larger e-ink panel factory it’s uncertain if they could shave costs by more than maybe 40% to 50% per panel.

And it’s very questionable whether there is genuinely 10x more latent annual demand for e-ink panels at that still not cheap price point.


If you have a Kobo (~ US$ 120 or even less used) you can run this using Cobalt https://github.com/BandarLabs/Cobalt/tree/beta/apps/birds

It seems like Kobo eReaders don't come with a mic.

You could probably use an ESP32-S3-PhotoPainter (which comes with a mic) to do this for even cheaper: https://www.waveshare.com/esp32-s3-photopainter.htm


Yes absolutely, the current Kobo one uses the Mac's mic. I work near a window on my Mac so this is handy without any hardware setup.

Agree, this was really a beautiful thing to see!

EDIT: no Ai for pictures - cool!

> This is the coolest thing on HN that I've seen in a minute.

Off-topic, but what is up with the increased use of the phrase "in a minute", presumably to mean "in a long time", lately?

I've only started encountering it in the past year.

Did it get popularized by some celebrity, tv show, influencers, etc?


seems like you know what you're talking about 5k% increase starting in July this year

https://trends.google.com/explore?q=I%27ve%20seen%20in%20a%2...


I think the trend is probably misleading. A popular youtuber may have released a video in August with a title containing "in a minute".

It's been around in slang for a few years with a few variants, e.g. "I haven't seen her in a hot minute." I think it's just filtered into wider cultural vernacular.

people have been saying "it's been a minute" to mean a long time since at least the 90s in NY

Yeah, I’ve been saying hot minutes for decades myself!

I had not heard it before (or if I did I did not interpret it correctly, eg if one told me "see you in a minute" I would interpret it as "see you in a bit".

But there is this podcast discussing it in 2021, and it comes from black community slang from 70s (which is where most slang I encounter comes from). And these terms take a while to catch up usually, but it seems it was already circulating more broadly since 2000s.

https://waywordradio.org/its-been-a-minute/


Scarlett Johansson says, “See you in a minute" in Avengers: Endgame.

I say "see you in a minute" to mean "see you in a bit". What did Scarlett Johansson mean by that in that movie?

This phrase has been around for years, I know because it has always infuriated me.

They took a well-defined unit of time, which is relatively short, and made it mean “some unknown but very long period of time”.

So frustrating. /oldmanyellsatcloud


it's been around. I more commonly see it as "been a minute!" when you see someone you haven't seen in a good while.

It's been a phrase for decades.

[flagged]


BirdNET has great coverage for the US/North America. Part of its main training set is the Macaulay Library, which is also used by Merlin/e-bird.

https://birdnet.cornell.edu


Also: please use a credit union and use mutually-owned insurance agencies. As a general statement, you'll be in way better hands.

As a founder of a few software startups, I agree with this. If I can give advice to other founders, remember than software is usually a means to an end, and for most users, they just want the bare bone essentials to work and work well. Everything else is nice-to-have-fluff that won't drive 99% of sales.

The apps I pay for are simple and focus on a super basic interface. BookFusion and Libro.fm are great examples.


It’s interesting because my experience is inverse (b2b software) - the bloated shiny solution what actually doesnt work at all performs better in the market than the sober, unsexy solution that just works.

Companies like Oracle or Microsoft are so massive not because they sell simple and focused products that work well; quite the opposite

Might be different when you’re selling to endusers


Oracle (and SAP and IBM) are not in the business of building software, they are in the business of never complete software, instead stretch the project as far as you can and bill to death.


Oracle and Microsoft are in the contracting business, and software is just an excuse for them to get contracts.


Q3 XL and Q3 XS are the two I'm trying to decide on


You might want to test this new dynamic GGUF:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

(I don’t know much about it, just saw a YouTube video about it last night)


Another one to try:

https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-GGUF

Runs the 3bit model faster than the 2bit one runs on my old-ass card. Can’t vouch for its intelligence yet, but i suspect whatever loss in smarts it takes is made up for by the extra resolution.


who uses Iroh?


We're using in Radicle for artifact distribution: https://radicle.network/nodes/iris.radicle.network/rad%3Az4V...

And also working to replace the Radicle's networking stack with Iroh: https://radicle.zulipchat.com/#narrow/channel/369274-General...


I've seen some product prototypes using Iroh and it appears pretty good for "just get edge device / node connectivity to work" when you don't want end users to bother with networking shenanigans but can guarantee Internet connectivity.

Basically Tailscale but embedded into the app without the hassle of requiring users to setup accounts.


I really wish Moonlight would use Iroh. I use Tailscale to make it work, but it would be so much easier to just pair a machine and have it always work.


isn't it brand new?


developers, not end users


Hello Hacker (M)ews furry. Tell me a little about your journey. What got you into furries and tech? What kind of tech are you most interested in?


To what end?


The new IQ4XS has been working pretty well so far on 4090 16gb.


What size context are you able to squeeze in with less than 2gb of headroom? I have had some luck using a quantized kv cache but i fear that also decreases overall quality.


Try the llama.cpp fork by thetom. It's called turboquant after the technique


I need someone to run actual benchmarks between the two.


Benchmarks are the BMI of model evaluation.

They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.


> Benchmarks are the BMI of model evaluation.

That is such an elegant way to put it.


Only relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.


These models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.


Hi Gertlabs, my inference company has actually started offering Ornith1.5 today for 9B and 35B A3B.

If you are still interested, send me a message on X and I'll help you get started! https://x.com/romulushill

https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-35b-a3b/

https://scalattice.com/developers/


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: