Hacker Newsnew | past | comments | ask | show | jobs | submit | CMay's commentslogin

People are animals and animals can be trained. Dogs can be trained, elephants can be trained, tigers can be trained. This fact is hated, but has been known for thousands of years.

Foreign adversaries, political extremist groups and so on are all clawing away at every facet of society that can contribute to training people to be and do what they want. For all of our information, language and intellect, those strengths are just an open vulnerability if they aren't protected.

You have to train yourself sufficiently on the grounded basics of reality, the truth and the nature so that when adversarial training against you and others strikes you can be wise enough to push back like this. You have to be able to surround yourself with people who disagree and stand your ground, because they're all being pumped with garbage from everywhere and don't know any better.

It is one of Russia, China, Cuba's and even Iran's way of attacking the US at a distance by funding these groups and amplifying the dumbest stuff you can imagine until these minority voices reach majority saturation in some of the information spaces where they can fester. Taking over education, journalism, social media and so on.

If the idea of the open society is going to survive into the next century, we're either going to need AI to ground society if we can prevent AI from becoming too centralized, or we're going to need to rapidly change education to teach people to think even if the reason we haven't been doing so up to this point is that it fragments society too much. If the alternative is a fragmentation that our enemies control, then a new enlightenment may be the better way.


> It is one of Russia, China, Cuba's and even Iran's way of attacking the US

It's the USA eating itself and trying to attack Europe, if anything. It's American fascist billionaires who own, buy, and steer these social networks.


This type of content is mostly on reddit, which leans heavily leftist. They censor right-wing views aggressively.

Owning the social networks doesn't mean you control all the information inherently, because it can easily get out of control. It wasn't that long ago that hundreds of thousands of bots had to be removed after running rampant for a long time which were manipulating trends that cause peaks on journalist's social media dashboards that influences what stories they cover.

Yes, social media bots impacting statistical trend services which some news organizations use to decide what stories to cover.

Communists caused World War 2 and yet love to play the victim. Not everybody who is against communism is a "fascist" and most people don't even know what that really means.

We're not attacking Europe, we're making Europe stronger so it can stand more independently so it doesn't require as much US intervention in the case World War 3 does start and a new front opens there beyond Ukraine. That way the US can invest more of its resources in the pacific. NATO understands this.

The reason we had to do it rudely, is because we've been trying to get Europe to do this for decades nicely and that hasn't work. So you get the rude approach.


> Communists caused World War 2 and yet love to play the victim.

This is an unconventional position. What's your rationale?


My best guess is something like this:

The Soviets gave the Germans a free hand to begin the war by splitting Poland via the Molotov pact. Without that agreement, and without the preceding years of Soviet resource shipments and technical exchanges, it is unclear if the European war would have happened.


This is about as far as I could guess as well. Unfortunately by this logic you could just as easily implicate France and the UK for the Munich agreement, as well as the failure to excercise their defense pact with Poland. So essentially every party with exception of Germany and the United States is at fault for WW2. This also leaves the Pacific theatre unaccounted for as you implied. So I'm very curious what the original commenter intends.

"ackchually, facism good" is what he intends.

It doesn't merit more response, you can just look at the reality of actual fascist states. Economic stagnation, the most extreme corruption and incompetence, and cruelty beyond imagination. True masterminds.


The rationale is that Communist International (Comintern) based in Russia at the time was openly aiming for total global communist revolution. These weren't simply activists (though there was that too), but actual armed militant Marxist groups. They were in China, Japan, Spain, Germany, France, Britain, the US you name it. It was a pervasive problem in the early 1900s.

The US and Germany were both helping fight back against the communists in China and Spain. Everyone was basically on board with being against the communists, they just weren't on board with the extremes to which Japan and Germany were going. Japan was passing laws to suppress Marxist revolution in their country in the 1920s and Germany went to many extremes, routing out communists anywhere.

In Germany you even saw education being changed to fortify new students as anti-communist, because preventing communist revolution from taking over Germany was existential for them. The teachers knew what they were teaching, but they were afraid enough of communism that it seemed a reasonable defensive mechanism. There had already been examples of communist revolution upending countries, so it wasn't as if it was a fake threat or concern.

It was very real, and people who dismiss the pacts signed between Japan, Germany and Italy to cautify the fact that they explicitly saw communism as a threat, do so in ignorance. It's often argued that it was some kind of fake political front to justify their actions, but on close inspection that seems to absolutely not be the case. These people legitimately believed communism was an existential threat.

This is further evidenced in the years after WW2, where there were many more communist conflicts in other countries. Korea, Vietnam, Cuba, Cambodia, etc. If you put yourself in their shoes, simply imagining being any of these countries, would it be unreasonable to feel an existential threat? No. Just, flat out, logically, no.

So why did they end up invading all these other countries? What was that all about?

The communists love to control the resources and in terms of a total global communist revolution, you've heard things like domino theory and you could understand it in numerous ways. One is that, if you expand to control a tipping point of resources, then you can leverage those resources to dominate the rest of the world and enforce your revolution.

Japan and Germany wanted to route out communism in other countries, but they also wanted to leverage the resources. Germany working with Russia to take Poland was largely strategic, not a sign of actual good faith towards Russia. Both Japan and Germany explicitly intended as their end goal, to conquer Russia and destroy Comintern.

This is part of why Stalin dismantled the Comintern, to deflate Japanese and Germany aims while simultaneously fostering military support from the allies.

Another critical point to support my argument, is that even if you don't think communism was the primary issue, why the heck would Japan and Germany join forces, causing it to become a World War? You could certainly argue that it doesn't become a World War until that happens. After all, people often make the mistake of saying the war didn't even start until something random in Europe happened and they can never ever convincingly explain how that caused the war to spin out of control. It makes no sense, because it is BS.

The war was already raging in the Pacific long before anything began in Europe. So why were the things in the Pacific happening? The thing that unifies both of these conflicts is communism.

People often deflect and just say oh Hitler was a racist and it was just this and that. They never answer the question: so what about Japan?

If they were fascist, what is a fascist? Many things, but most fundamentally they are a reaction to the communist threat. Communism is underpinned by a philosophy that all action is social action, so every person must be an activist. This puts society out of order rapidly, because anyone not involved gets overrun if it gains traction.

So fascists develop psychological counter-strategies to communism. That's why the Nazi flag was red. That's why they emphasize race instead of class, to unify people rather than divide them. That's why they uplift beautiful art and collected precious artifacts that communists hate and try to destroy. Everything was an intentional counter-strategy to communist revolution. It's easy to look back on it and say it looks insane, because it is and it is patterned as the inverse of communist revolution which is also itself insane. They wanted to immunize their country as a cultural powerhouse to fend off communism for all time.

And if the revolution was going to take over sufficient resources to steamroll other countries, then Japan and Germany were reasoning that they would need to expand to counter it.

Comintern wanting total global communism was logically inept and wrong for them to do. Germany and Japan going nuts and killing a ton of people was wrong for them to do.

It wouldn't have happened if the communists had not created a global existential threat. Remember too, this is before the internet, before pervasive communications and such. People still largely got their news from word of mouth, newspapers and radio so I'm sure there was plenty of miscalculation to go around, but the slow verifiable things were undeniable and that's what they were reacting to.

The communists have been threatening to start World War 3 and that is also what the Trump administration is reacting to today. Between Cuba, Venezuela and Iran, both Russia and China have lost essential partners in rapid succession. We've been helping Ukraine target Russian oil infrastructure which has dual purpose, since it both weakens Russia and also demonstrates to China that not only is their Venezuelan and Iranian oil not secure, but not even the oil they get from Russia is reliable. China built the largest strategic oil reserve in the world which is something like double the rest of the entire world combined and they were going to use it to sanction proof themselves in the case of engaging with Taiwan. Now they are drawing it down, which reduces the chance of WW3.


What reading materials/authors do you recommend? The history here and in many of your comments is orthogonal to my mainstream understanding.

There is whole genre of authors (Ernst Nolte, Viktor Suvorov, David Irving) who argue WW2 was all preemptive strike, that middle classes felt under threat and that soviets were preparing wide actions to take over europe.

It's revisionist history. Constructed arguments that quickly fall apart - like Yes Hitler was afraid of Bolsheviks but that doesn't explain why he invaded all the other countries France, Poland, Denmark, Norway... many more.

The repeat of the same trick on current events with Iran/Venezuela/Cuba is very telling. I doubt anybody believes that US attacks Iran because of communism. But yes all the US conflicts are because of communism.

More likely explanation is that big business backs fascism to crush social causes. Fascism exists to save capitalism from the people.


> Not everybody who is against communism is a "fascist"

The man that bought one of the most important social networks and has substantially changed its algorithm is nearly openly a nazi. It's not just the sieg hail the day of the presidential swear-in, it's everything he does to the world and the USA. And so do all the others.

> So you get the rude approach.

You're digging your own grave. Your president is handing over your empire piecewise to your adversaries for a fistful of gold. You must understand this: the only thing worse than communism, is fascism, because it's just corruption as a form of government.


What makes someone a nazi? What specific worldviews? Holocaust sympathy? Aryan supremacy?

german here. just adding some soup to the meal, not offending.

for the definition what a nazi is, the movie "american history x" - that imo shows exactly what a nazi is: the main character is not german and not even in germany, but shows the nazi-mindset exactly! He glorifies hitler's thinking about others, fighting for the wrong idealogy for "own" country. ... Thats making a nazi. Neglection/sympathy for holocaust, supremacy thinking, .., too.

As I was told, the word nazi was invented by U.S. for the hitler and his followers so its a short word for national socialists (the party, its memembers and sympathisants) the US was fighting ("name your enemy"). Your question is about/implies nazi as "an individual person with a special characteristic", which makes that someone a nazi.

Its a "nation first + everything for the nation" thinking. For me a nazi is someone who glorifies the nazi time, saying it was all well done then, wanting those times back - like some elderly tend to say "earlier, everything was better". They dream of powerfull country with a straight people-body. Nazi think in races. And always say others are guilty for one's own failure, argument with "they" and "us" when they talk about own country/others.

So basically, patriotism, skinheads and far right are subsets of "nazi", but not full nazi :)

And may be, if one wants to be a REAL full nazi, its a must to shave the head and to have some on topic tattoos with being caucassian at the same time. This is how we differentiate them in germany - you see: "we" and "them". But i'm not nazi :)

funny, in US there was a nazi gang a few years ago, with the group's "leader" having the room full of nazi-stuff, flags and propaganda - as it came out later, he was jewish by himself and did not tell ..

Frank Meeik, Daniel Burros, but the one i mean was just like a joke of a nazi. But still dangerous thinking.


Has there ever been a prosperous and successful country that has not put its own nation first?

Exactly, which is what makes the Nazi/fascist/autarkist/autocratic position even more absurd. As we're seeing play out in the US, calls to "put the country first" are rooted in hate for the country that actually exists, and end up being quite destructive to the country's actual interests!

There would seem to be a common pattern whereby preaching of values tends to fill in for actually living out those values (eg religion, performative virtue signalling, so-called "rightist" libertarians), and the whole "put our country first!" or "make our country great again!" seem to be strong examples of this.


Has there ever been a prosperous and successful fascist country?

No, they stagnate economically in corruption.


No, no, answer the question posed--which country has prospered by putting other countries first?

> Its a "nation first + everything for the nation" thinking.

If you think this means "our nation over other nations" then you misunderstood. It means "nation above all", e.g. "everything for the government, above citizens". That is, forget the individual, there is a greater good for which we must endure some hardship, trust me.

FYI, the promised land never materialized, historically.


> If you think this means "our nation over other nations" then you misunderstood. It means "nation above all".

A distinction without a difference, the second automatically leads to the first, no exceptions.

> "everything for the government, above citizens". That is, forget the individual, there is a greater good for which we must endure some hardship, trust me.

Right, in addition to being a code for "we over the others", "nation above all" also provides the excuse for the social mobilization necessary for wars. Nationalism is a tool of war, it may start in some covert form but the end result is the same.


Fair enough--so then, which successful nation has existed that prioritizes the individual over the nation's laws in any capacity?

Which successful nation has eschewed the idea of the greater good?

(And to your point--this absolutely has been perverted in various cases to justify atrocious behavior! It is necessary but not sufficient in order to govern well and righteously and be a successful country.)


It's not necessarily a binary choice, is it? There's a spectrum between true democracy that isn't simply "rule of majority", and authoritarianism.

There is a weakness in democracies too, that if they are too tolerant and allow too much freedom, this can paradoxically undermine democracy itself and lead to a collapse.

So in general, western democracies tend to respect human rights, sacrificing the greater good, because this in itself is considered a greater good.


No, _you_ must understand this.

Originally we had The League of Nations. The idea was a sort of collective defense. Europe failed. They did not come to each other's aid.

After WW2, that was the second time the US had to come stop the wars. We weren't simply going to join a defense pact if Europe couldn't demonstrate that its own collective defense had teeth. So the establishment of the Western Union was an attempt to demonstrate to the US that Europe was finally ready. Then NATO was formed.

Now, Europe has been underfunding its defense for decades and not meeting their NATO commitments, so that original question was back on the table. Can Europe even defend itself? Is it overly reliant on the US?

So, no. We will not apologize for being tough with you, because the world is changing and you no longer have the option to just stagnate. We're not losing any allies, that's just rhetoric and propaganda. You're not a serious person if you think that somehow western countries aren't going to help each other.

Making you stronger means you can help yourself or even help us later. Fundamentally, if our allies are stronger it also changes the Chinese calculus on how much of a threat from Europe they have to factor in before possibly starting World War 3.

The bottom line is, get off your asses and light a fire under your government to invest in defense spending, or be a small part of the cause of future millions of deaths, because your feelings were hurt.

If you don't even have enough compassion for fellow humans to want to stop World War 3, then you have no moral ground to stand on.


Your underpinnings for your arguments are the classic infant/adult fallacy where the people you're not sympathetic to are adults with incredible agency and power, and the people you have sympathy for are helpless children with no choices.

Putin wasn't forced to invade Ukraine. What was he even putatively trying to avoid? NATO's not belligerent (you can't simultaneously argue that Europe's military is pathetically weak, but also that it was so muscular and menacing that Putin had to respond proactively).

Hitler chose to start WW2. He wasn't forced to invade France to stop communism. He wasn't forced to attempt over and over to invade Great Britain to stop communism.

China, Russia, etc aren't looking for openings to start WW3. It would be devastating to them--just the war with Ukraine has had generational consequences for Russia. They also aren't innocent nations who, after intolerable saber rattling by the West, will have no choice but to start WW3. It's further not the case that the only deterrent to these innocent nations regrettably starting an insane conflict is a heavily armed Europe. The US is pretty heavily armed! France and GB are nuclear powers!


The way I see it, World War 2 was kind of a perfect intersection of the great depression creating a perfect opening for communist revolution as countries had seen during and after World War I and the protectionist contraction of trade during that time made matters worse. So you get countries that are simultaneously starving economically while risking being eaten alive by communism.

What we see today is different. Russia and China are projecting a shift to their hemisphere in terms of world influence. The president of China has openly stated that he thinks Karl Marx was right and has been re-aligning compliance away from reform. With such a large military build-up, Taiwan could simply be the first domino to fall to complete their historical communist revolution.

The CCP was directly established by Comintern, and even though the current thinking is that global communist expansion was abandoned, the way Russia, China, Cuba and Venezuela have been behaving make communist or "post-communist" states a resurgent threat.

War would be devastating, but the idea is that if it was a World War, then military forces would be distributed to fight on many fronts and this would limit the west's capacity to interfere with China. This is part of why we have supported Ukraine, but not overinvested. It's part of why we captured Maduro, who was threatening to invade his neighbor. It's also part of why we're sorting out Iran, because we need to take that trigger button out of China's toolchest. Imagine it going off while we're busy with China.

With the state of the world now, if China invaded Taiwan and tripped the Japanese tripwire, we would largely be free to focus on China. That is I think what Chinese analysts see now.


> The way I see it, World War 2 was kind of a perfect intersection of the great depression creating a perfect opening for communist revolution as countries had seen during and after World War I and the protectionist contraction of trade during that time made matters worse.

I think when economies falter governments falter, and their philosophies are discredited. So it makes sense that governments were on thin ice in the wake of WW1 and the Great Depression, and that more radical parties from the left and the right were ascendant. I think something similar is happening now, if of lower magnitude. But, no way was this a door only communists walked through, anarchists, fascists, etc were all in the mix.

> The president of China has openly stated that he thinks Karl Marx was right and has been re-aligning compliance away from reform. With such a large military build-up, Taiwan could simply be the first domino to fall to complete their historical communist revolution.

China is far from a Marxist state, they're much closer to Leninist/Stalinist/Maoist statism. If you want Marxism you should look to the EU, specifically Spain, Italy, and Denmark (lots of worker-owned coops).

Second, China's military build up is likely mainly to fortify its hegemonic aspirations, controlling the region, shipping routes, providing security to allies, etc. An invasion of Taiwan would be extremely costly. It's far more likely they blockade it and manipulate the internal political situation as long as they need to until Taiwan "chooses" reunification. My guess is we'd try to prevent this because we still need TSMC.

> then military forces would be distributed to fight on many fronts and this would limit the west's capacity to interfere with China

You are ascribing way too much strategy to the Trump admin. They deposed Maduro because they could and for oil, all of which was dumb, and they tried to decapitate Iran because Trump wanted that on his resume (also to try and out-do Obama). Neither Iran nor Venezuela had any capability to strike the US or meaningfully strike its allies, definitely not on any long-term basis.

Finally, I don't exactly know what you mean by "focus on China." The West's relationship with China is complex, a lot of people are focused on it. The likelihood of direct confrontation with them is quite low, as it'd be disastrous for both sides.


You have huge gaps in your understanding of history.


The reason that many people don't understand how dangerous AI can be, is that listing the real dangers now becomes like a laundry list for less clever people to follow. It's highly unlikely you've ever seen publicly mentioned the real risks AI poses, because the vast majority of people are simply not clever enough to produce them and the few that are have no interest in spreading it.

If you go to the various CEO blogs or misc people within this sphere and peruse their lists, they don't scratch the surface. It's all pretty vanilla stuff.


A real problem is that you might not even need to give these internet access. If your neighbor gives their TV internet access and it can establish a mesh connection then it might be able to exfiltrate data that way. I don't know of any TVs that do this, but it's always been a possibility.

Another possibility is even if you don't give it access to your wifi, if you eventually give your TV away or sell it, the new user could connect it to wifi and if it has any persistent storage then it could upload all of its stored data about you at that point.

Companies that do this need to end their entire brand, because they're just helping the CCP.


> A real problem is that you might not even need to give these internet access. If your neighbor gives their TV internet access and it can establish a mesh connection then it might be able to exfiltrate data that way.

No TV has ever been shown to do this. Please stop spreading this rumor.


I didn't say a TV had been, if you read. I will absolutely not stop spreading the possibility, because it is in fact a possibility. This is more true now than it has ever been, because even if a manufacturer doesn't do this we're now in a world where AI could potentially exploit devices and make hops like this. The risk is higher in areas with higher population density though, like multi-story apartments.

We know that device connection sharing is a thing that has occurred and it's not unreasonable to see the risk of it coming to TVs.

Why, do you work for a TV company?


Partly it annoys me because the answer to the question "how do I protect myself?" is simply "don't connect it to wifi" and that's it. These rumors make it all complicated to answer when it's really not that deep. This isn't deep NSA spy craft, it's shitty software uploading logs and half-assed data collection to the Internet. That's it.

But if I'm honest, it's mostly because I'm disappointed that they don't do this. It'd be a fascinating story. Imagine the braindead balls it'd take to make a feature that insanely abusive, and how good the outrage out it would be to read. That'd be a killer story for someone to break, it'd be like the xz backdoor or the Snowden NSA leaks. Instead we get these dumb "Mew under a truck"-quality grade school rumors in comments sections and it's just so stupid and disappointing.

(And honestly-honestly a tiny part of it is trying to goad one of these commenters into actually doing the work to find someone who really is doing this, so I can read that story.)


The story is both a little overblown and a little underblown, just different parts of it. More people need to be thinking about the feature level of devices and understand that "off" doesn't really mean it most of the time. Does the TV need a camera? Does it really need a microphone? Does it need to be smart at all? The matter of whether it is abused is increasingly less important than whether it can be abused. There have been all kinds of clever abuses we've been lucky not to see over the past 3 decades, but I'm afraid that relative quiet may be over.

I still don't understand why there wasn't an absolute outrage about internet providers putting their own routers inside your network and blocking other routers by default unless you call them. Suddenly all these providers have managed hardware inside your network and if they wanted to, they could scoop up all of the LAN traffic.

On top of that, some of them even portion out wireless to people outside your network from the router you're paying for while increasing attack surface.

The big tech companies, the providers, the government and foreign adversaries are winning bigtime. It does feel like there's a little bit of hope that people are wising up about AI exploding privacy and security risks, so more people are thinking about the dumbest things which have become the norms.


The guy can't read, he only knows how to keep saying "there's no evidence TVs do this" over and over, even when you already said there wasn't evidence yet.


Does this mean that TVs can never be recycled or resold now, adding to the garbage problem? All you have to do is sell the TV to someone who does eventually connect it to the internet and it will upload all the data it acquired while the previous owner had it?


This is not a bold claim, this is just how it is.

If you ask any average gamer, "if you buy a game on Steam or on a CD from Walmart, do you operate on the understanding that you are now legally allowed to make as many copies of it as you want and sell those copies?"

The answer will unanimously be no. If they owned it, the answer would be yes. They might say they own it, but they will clearly and reasonably understand that they do not own it in legal terms, because they understand what they cannot do with it.

Some of them will understand that they can legally make copies for backups, but why would you need the law to tell you that it's legal for you to make a backup of something you own? You wouldn't.

You can buy a hard drive, but buying it does not give you the IP for all the technology that went into it. No reasonable person believes that would be the case, either. You can buy a car, but you can't then copy all the parts and start mass producing your own copies of that car. Do any of you go through the McDonalds drive thru and believe you now own the burgers, fries and all the packaging that goes with it to the extent that you can start up your own McDonalds with logo and all?

Whether it's physical or digital, even if people have contradictions in their head since they aren't lawyers, they understand enough about how things work to conclude that what they understood when they pressed the purchase button equates to not obtaining total ownership of all aspects.

It is simply true. This is so broadly understood that I don't even think you would need to use a jury. A judge could simply throw the case out at this point on that alone, if it hadn't already been settled in past legal precedent, which it has.


> if you run one model, run glm-5.3

That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.

Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.

Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.


The point I'm making is that most models are good enough for most tasks, so choose on speed/cost.

Definitely benchmark on your own tasks. GLM-5.3 is the winner on mine, on yours maybe not. I am not trying to be a universal benchmark, as these serve no one but the person doing the benchmark.

My previous post sank like a stone, but all the evidence + code to run this + what you have todo to adapt it for your own use cases is all here https://github.com/ed-is-ai/featherbench

Encourage everyone to eval like the devil


I think they said they were keeping the old UD 2.0 quant for the larger sizes? So maybe they kept those the same and simply reuploaded them. They said the newer UD 3.0 quant performed worse on some things for the higher quants. So now it's a mix of UD 3.0 and UD 2.0.


For me that moment was Gemma 4 12B QAT. You're not suddenly going to start throwing your hardest programming problems at Gemma 4 12B QAT, it is still 15B parameters less. It's more that, aside from pelican art which isn't what local models are for, I didn't see anything on Simon's post that it couldn't assist with or largely succeed at.

It can run 80-100t/s on a laptop, can understand images natively and do bounding boxes, read tiny text, understands audio natively as well and can transcribe or translate anything you say, can do accurate long context retrieval with pretty large context windows, tool calling, excellent reasoning and is very token efficient.

It's only 7GB including the mmproj or 8GB with MTP. The Qwen 3.8 27B model Simon was using is ~18GB with MTP+mmproj, rather than 17GB alone. The point is not really that you compare these models directly, but that Gemma 4 12B QAT was really a special moment in model releases deserving of a similar reaction relative to its size, but was mutilated by Google themselves, Unsloth and Llama.cpp.

The overall appreciation I think we're seeing this year in particular is that people are easily surprised when multiple things are improving simultaneously which produce seemingly exponential changes. It isn't just that models are getting smaller, or that reasoning is getting better, or that speculative decoding is becoming mainstream, or that models can understand audio and images better now, or that they can reliably call tools which expands their capabilities, or that context windows are getting larger, or that accurate retrieval is improved, or that.... and so on. It's all of them narrowing in at once that is starting to make local models incredible and truly useful for far more use cases on the existing hardware people already have.


> It's only 7GB including the mmproj or 8GB with MTP.

Even more impressively it doesn't have a separate mmproj at all — it is fully integrated, and the vision encoder doesn't speak words into the LLM, as it were —- it is directly integrated into the model's weights.

I have banged on about this model here enough but I really agree that Gemma 4 12B is a candidate for the most impressive LLM of the year. It is remarkable, and I think because it is a small model that isn't apparently excellent for long-context agentic coding, it has been largely ignored.

It is, actually, quite good at coding jobs. (Though its grasp of nuance is a bit weaker. For example, it doesn't know that closures created inside PHP objects have implicit access to the object as $this, and always seems to need reminding.)

If you instead treat it as a prediction of what consumer on-device AI may very soon be able to do, or even as a possible future into a sort of lower-ratio MoE, or the basis of a modest private offline educational LLM model, it's very interesting indeed.

I've learned a lot from it — the fact that it performs so well at such a small size really does help you assess claims made for much larger models, and it's quick enough on my M1 Max to just muck about with.

I do think the release of these models was somewhat fluffed up, and I don't think it helps that the 31B model uses global attention so it underperforms on the kind of older GPUs that are on a lot of desks; it's no better on those than it is on my M1 Max, where other attention schemes seem to be radically better.

Now that tool-calling is mostly fixed, it's well worth playing with them.


Well in my case I'm using llama.cpp and the mmproj is required, but I think it is just an extracted part of the original model file. Even with audio, yes it technically supports them natively and they're "encoder-free", but in practice that doesn't mean no translation or processing is required before it goes into the model. It does require much less processing though, which reduces latency.

As for coding, for sure there are many important details that a model needs to know in order to produce correctness and the smaller a model is the more it ends up training out. If there's a task you do consistently enough though, often times you can simply provide a pile of essential context so it has good enough reference to not need the extra training data.


My issue with Gemma 4 is that any task fails to complete after any compaction event. It often ends up in a loop that keeps compacting and showing the same compaction output. Qwen3.8-27B-IQ4_XS was a massive improvement. It's tasks survive compaction and actually get completed. I switched to Qwen3.8-27B-UD-Q3_K_XL for better performance and its working just as well.

Gemma4 screwed up a proxmox install I had. I booted to a SystemRescue install and tried to get gemma4 to fix it. It just could not do it and kept having issues where it dropped a linux command into the local powershell because it did not ssh into systemRescue or killed the ssh connection somehow so the text landed on the wrong system.

I told qwen3.8 to investigate fixing the partition. It said information was lost, but displayed enough info that it was easy to tell it was right. I told it to install fresh proxmox and gave a short rundown on settings and partition sizes I wanted. It made a plan and told me I had to manually installed proxmox by booting the iso. I responded with something like "there are other ways to install promox without human interaction so use one of those". That was it. I woke up to the system having booted to a new proxmox install with my previous ssh keys restored and my existing zfs pool already mounted.

I don't see how any model that is limited to a single context window in a single session would be viable for coding. I want something that can manage the entire project and not just individual files or inline suggestions. I need to be able to feed it all the info I would use to make coding decisions and then have it at least make a working project that it can launch and test successfully. You want it to ask as many questions up front to enable continuous work without stopping for human input.


Out of the loop here. What did Google and unsloth and llama do to mutilate Gemma? I can understand Google shenanigans but llama and gunsmith is kind of surprising.


Google provided incorrect settings and an imperfect template.

Unsloth modified the template and then finetuned their own version of the model to optimize for some benchmarks as a means of validating quants.

Google and Llama.cpp then adopt template changes by default, so anyone downloading the new model or even using the original model will now automatically be using it incorrectly.

Llama.cpp also uses the same inference setting defaults regardless which version of the model you use and some settings are simply defaults it uses for all models.

Then even if you account for all of these, you have to be using Gemma 4 itself correctly, which many people do not.

All of these little changes and inconsistencies hurt some of the model's original capabilities. Even if you go directly to Google's repo and download the full float 16 weights with the template they have there now, you cannot simply assume you're getting the best results.


I am very much a beginner to local LLM stuff and I find it incredibly hard to figure out how to run models optimally with the correct settings for my hardware. The number of different variations of the same model and how each quant work is super confusing as well.

When I tried to run llama.cpp directly I was getting max 9tk/s on qwen3.5-9B, then I tried LM Studio with the same model and got 77tk/s. I haven't figured out yet how to get MTP working properly in either.


If you are on Mac, have a look at the Llama-macOS app. They claim sensible settings for the linked model downloads. I'd expect the authors of Llama.cpp and the Huggingface folks to know this stuff.

https://news.ycombinator.com/item?id=49328008


I have had luck telling the free chatgpt my graphics card brand/vram and asking it to recommend latest qwen3.8 or gemma4 model variants. Then I pasted in my server command and ask it to optimize it. I also pasted in the token per second logs to get further tweaks. If you paste the token per second info log info back to chatgpt, you can iterate with the free chatgpt to get better settings.


If your package manager / configurator isn’t claude code or codex, you’re wasting time.


Unless, of course, your goal is to actually understand what is going on, regardless of whether that is difficult.


Some of us prefer to avoid Anthropic/OpenAI


Your funny.


And what's the right way to use Gemma? Where can I find the correct template and settings if those aren't the ones provided by Google, Unsloth, and aren't built into llama.cpp? I discarded using Gemma 4 because it got into weird loops when tool calling


Some weeks ago a new official Gemma 4 release was posted that corrected some of the chat template problems. So the official release files on hugging face should be the way to go.


The updated version will handle tool calling better by default, but the reasoning quality is no longer preserved and is mutilated quite badly.


So then how do you run it unmutilated?


Download the original model with the original template, not updated versions of the model or finetuned versions of the model.

Then create your own reasoning tests to verify that it is working correctly. You can set a specific seed value to make sure the generation is the same every time, that way you can identify any tokens that are different.

Afterwards, try making small incremental changes to the template and validate your tests each time in order to try to adopt the improvements from the newer templates. If the reasoning quality degrades, undo your changes and try again or test alternative solutions.


Log all the calls and run on a periodic cadence (cron or ever N turns) a larger model (like Opus) to read samples of the traces and edit the template to fix observed problems. There are some signs that help find interesting things to look at, errors of course, but also overly long responses, prefix cache misses, tool call errors, etc.


I run llama.cpp and specialized forks on 64GB of HBM and I still cannot figure out where to find the final correct guidance on using the Gemma 4 models.

Would appreciate any kind of pointer to the latest!


> It can run 80-100t/s on a laptop

That is a lot, what is your laptop hardware?

One issue I have with Gemma is that they seem to use old architectures that rely on full attention, requiring a lot of RAM for context and quickly degrading speeds as context is filled.

Qwen 3.5+ is much better in that regard with its super efficient context. Even on Macs, speeds take degrade much more slowly.


> One issue I have with Gemma is that they seem to use old architectures that rely on full attention, requiring a lot of RAM for context and quickly degrading speeds as context is filled.

Yes, this is something I hope they will change. Gemma 4 31B is much slower on pre-Blackwell GPUs as a result, which is a bit of a shame for local model experimentation.


Even Muse Glimmer (as did GPT-OSS I think) does ~4 sliding window attention layers + 1 full attention layer (like Gemma 4). I’m assuming both labs have good reason to think that gated delta nets are not optimal.

Of course it’s possible the labs just stick with the optimal architecture for large models and GDN is best for smaller models.


Thanks for the reply. Sooo much I have to learn.


I've found that the Gemma series of models are made for someone entirely different than myself. They fail at even the most basic questions I throw at them, like 12B just now failed at answering how `XGrabKey` from Xlib is used. It hallucinated the entire API and made up an entire flow of code based on it, for no particular reason. It could've even decided to research this via web search because I have a tool specifically set up for that, but it "chose" not to, relying instead on completely made up information.

This isn't an isolated incident, really, I find myself always having these issues with the Gemma series. I'm sure they can do useful things for someone else, but for the things I want to use LLMs for (very small code generation, quick questions, code review) they always seem to disappoint me. I'm sure it's because of the stuff that I do and use, but it's a very consistent red thread with these models for me.

Edit:

The same question for Qwen3.6-35B-A3B produces a pretty concise and correct answer that would be useful to the questioner, without even going to the web. I don't know what Gemma models are trained on, but it's not the stuff that's relevant to me.


> transcribe or translate anything you say

Is it multimodal? How do you do transcription with it?


Gemma 4 E2B, E4B and 12B unified accept audio - here's a recipe using MLX that can use it for transcription: https://simonwillison.net/2026/Apr/12/mlx-audio/

Only up to 30s though, and the larger 26B A4B and 31B models are text and image only.


Or you can use parlor to chat with it directly https://github.com/fikrikarim/parlor/


Credit where it's due. Qwen 3.8 27B is only the second local model after Gemma 4 that managed to correctly reason through one of my private benchmarks. It took 5x as many tokens to do it and 12m30s with MTP enabled, but it did do it.

Gemma 4 reasoned through it more implicitly, while Qwen 3.8 reasoned more explicitly. Laguna and Muse Glimmer failed hard on it, though they're useful for other tasks.

The VRAM usage seems way less efficient than Gemma 4 or Glimmer though, with 32K of context taking 2.5GB of VRAM. With those, even with MTP or a DFlash model loaded, you could still fit 256k-768k of context. With Qwen 3.8 27B I can't even fit 128k if I quantize V to Q4_0. Maybe with some trial and error I can find some settings that perform well enough with a larger context window that it's still useful for longer tasks.

Lots more testing to do, though I was getting some decent results out of Muse Glimmer which was more than twice as fast and supported huge context windows, managing to solve some bugs that Gemma 4 struggled with. I can't even begin to throw that task at Qwen, because just the prompt alone would use the entire context window and then it would reason for probably that same amount.

If you've got a 32GB card, it should be a decent model even if it really is memory hungry.

EDIT: Tried a few kv cache quantization settings, but it failed with those. I designed this benchmark to be pretty brutal in the face of KLD and any reasoning quality loss, so it's not too surprising. Gemma 4's QAT held up pretty well, at least and could consistently complete it.


Can you please tell which Gemma 4 variant managed to correctly reason through your private benchmarks? Was is Gemma 4 31B?

What quantizations and context lengths did you use for Gemma 4 and Qwen 3.8 27B?

I am asking because I can't even load Gemma 4 31B on my GPU with any reasonable quantization (even with small context), while I can run Qwen 3.8 27B with large context and good quantization...


Gemma 4 12B, Gemma 4 12B QAT, Gemma 4 31B, Gemma 4 31B QAT

Gemma 4 26BA4B would get close, but not quite and sometimes even get stuck in loops despite a repeat penalty.

Do not use any newer updated templates or Unsloth fixes. Use older official templates that released with the models on the huggingface repo. The template here worked: https://huggingface.co/google/gemma-4-12B-it/tree/657684fef0...

llama-server --model "model.gguf" -fa on -np 1 --jinja --ctx-size 262144 -b 768 -ub 768 --cache-type-k f16 --cache-type-v q4_0 --repeat-penalty 1.1 --chat-template-file "chat_template.jinja"

If you don't explicitly point to the template file, then llama.cpp will either use the template inside the model file or it will use its own template copy and your results may vary. Obviously some of the template fixes are useful to people, so it depends if you're having problems with tool calling or can't fix the tool calling in other ways for your scenario.

My experience with the QAT models was that quantizing v to q4_0 gave me better results than q8_0 or even f16. I think the Gemma QAT models may have been QAT trained to expect a q4_0 quantized v cache. If you're not using a QAT model, I would leave both at f16.

Another thing aside from using the QAT models and a Q4_0 v cache since you're having trouble fitting the models, is that you don't have to use the mmproj if you don't intend to use vision. If you need vision, but are hurting on VRAM, then you should be using --no-mmproj-offload. That will keep the mmproj loaded in system RAM instead of on your GPU. Loading images will be a little bit slower, but it can still be quite fast and you'll have more breathing room on your GPU. If you don't provide the mmproj file on the command line at all, then it won't load it anyway. If you're using some program like LM Studio, a simple thing you can do is move the mmproj and mtp files out of the directory for the model so LM studio can't find them and then it won't load them at all.

For Qwen 3.8 27B, doing any quantizing definitely hurt results a lot, so in my case I used: llama-server --model "Qwen3.8-27B-UD-Q4_K_XL.gguf" --spec-type draft-mtp --spec-draft-p-min 0.35 --spec-draft-n-max 2 -fa on -np 1 --jinja --ctx-size 65536 -b 768 -ub 768 --cache-type-k f16 --cache-type-v f16


Gemma 4 12B? This sounds really interesting with Q_4 (preferably QAT) this fits comfortably in 12 or 16 GB VRAM.

Could you elaborate on Gemma 4 12B capabilities from your experience and benchmarks?


Then you might be missing SWA. Gemma models are extremely memory hungry without


So long as they have flash attention enabled, Llama.cpp enables Sliding Window Attention by default for Gemma 4 models. Even if they're using Ollama or LM Studio I would expect those to mostly be doing the right things.


I would not expect Ollama to be doing the right thing fwiw.


"correctly reason through one of my private benchmarks"

i would also like to make one myself for my testing. could you give a rough idea or an outline or point in the general direction on what to do?


You can try https://www.vals.ai/vals-smith for this, saw it recently.


Qwen’s 3.6/3.8 27b actually has some algorithmic advantage when it comes to the kv cache size needed, so it actually needs less memory at equivalent context. My experience using both in vllm supports, with considerably more overhead in context size on these models than Gemma 4 31b, and better performance in most tasks I’ve tried on both models.


Qwen 3.x does have an advantage but it's relatively small (64KB/token vs 80KB/token) - Gemma4 actually has less % of full attention layers, but the largest geometry and has the biggest "fixed" state for it's non-global layers. Muse Glimmer actually has by far the lowest per-token cache usage for the competitive 30B-class dense models - it's at about 13KB/token - very aggressive GQA (32Q/2KV) and also by far the smallest QKV dimensions.

Actually perf (speed) is going to mostly on token output, and here Qwen 3.x historically tends to lose badly as it tends to overthink a lot. I'll be running evals on 3.8 myself this weekend to see how its reasoning levels perform.

I assume that AA will have 3.8 numbers soon and Intelligence Index vs Output Tokesn per Intelligence Index Task is a decent way to view that: https://artificialanalysis.ai/models/muse-glimmer?intelligen...


I was quite impressed by Muse Glimmer, and while I am sure people will observe that it is less good on benchmarks, my first experiences with this new 27B have been somewhat exasperating, whereas testing Muse Glimmer was rather fun. I have not tested either in an agentic context, mind you.


Yeah, Glimmer is excellent. You don't really test Glimmer with one-shots, because it's explicitly designed for multi-turn solution finding. The way I see it, if I've got a task that could be done either agentic or requires a lot of context (for example, dumping 600KB of API documentation and another 300KB of codebase for a project) then I would reach for Glimmer easy and it seems like it could get there most of the time.

Qwen might be useful to bring out for a second opinion on some more focused details that are largely information complete. Like, use Glimmer to bring together all the relevant critical data and evaluate what the actual problems are, then maybe prototype a solution. If it's still acting up, maybe throw the resulting context at Qwen and let it meditate on it.

I think there was some study done where ideally you would want to throw a bunch of different models at a problem since they don't all have the same perspective or diagnosis on what the problems or the solutions are.


That is exactly how this model has worked for me so far. Muse on a one-shot task will get to 80%. And if you even nudge it and say, "Hey, finish up," or "Review the syntax," boom, it's done. And I'm getting 20 t/s with Ollama on a MacBook M5 Pro with 48GB of RAM. It is a seriously impressive little model.


Glimmer works really well as an "explore" agent model (like in Opencode.) It seems to be extremely efficient at searching and collating that info, and executing commands.

From my testing so far, Qwen 3.8 is better at code but it tends to meander and take forever if it has to look in a lot of places. Glimmer will use like ~1k tokens to formulate a plan and Qwen 3.8 will routinely go over 10k


Have you tried turning down the new Qwen's reasoning effort level from xhigh, which it defaults at?

LM Studio isn't exposing a dropdown for this, at least with the unsloth build.

Unsloth Studio / Desktop does.


These templates actually fix the effort selection for LM Studio/3.8

https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates


Oh, that's cool, I saw those and I did wonder! Thank you.


I wouldn't find glimmer interesting except that it has much less memory usage per token of KV than Qwen. So I can get 24x concurrent glimmer on 2xRTXA6000 (with 128k context) where I can only get 6 Qwen 27b. This means I can get something like 4x the aggregate tokens/s out of glimmer.

For some usages that speedup more than makes up for it being inferior to Qwen intelligence wise.


Glimmer is fun because it's fast, tight, and doesn't wander or waffle. My favourite local model so far.


> The VRAM usage seems way less efficient than Gemma 4 or Glimmer though

Maybe it's implicit that you're using llama.cpp (although you don't mention GGUF), but it's hard to reach concrete conclusions about the model architecture based on one implementation in one runtime.


Aren't things like KV size inherent to the model?


There's a tiny bit of play, like sliding window attention. As tokens leave the sliding window you can keep them or discard them. If you keep them, you can freely truncate the context and resume generation from an earlier point. If you discard them, you have to recompute the KV cache up to that point.

Llama.cpp checkpoints and moves snapshots of the cache to main RAM for faster resumption after truncation.


In my experience MTP's speed increase doesn't seem to justify the apparent loss of success at the edge, it would have to be at least 4x faster to meaningfully churn through the first 3 failures in the time it would have taken to do it once without

What exactly are you doing that the prompt is eating an entire 65536 window? Surely it would be better to let it use any number of approaches that call tools to access that in parts and reason/summarize to an output file as it works through the whole thing. This allows it to keep the initial instructions in the start of the window and toss out the middle as it goes. IME many people who have written off local models entirely are, for lack of a better term "not holding them right" and consider them worthless.

Definitely also try IQ4_NL for K/V if you haven't. Because it's non linear it's far more hit/miss from model to model and especially quant to quant, I've found generally that it shines brightest when you start with a Q6K+ quant that you otherwise wouldn't bother with because of its size, which it then makes up for in both inference speed and often a larger context.

I do agree about Glimmer, though. It is quite good, far better than the benchmarks let on, especially in heavily agentic cases where it needs to rampage around the OS and utilize many different utilities to zero in on things. It is especially good at being told to try something itself, and if/when it fails, try Qwen, and if Qwen can't do it, call out to Deepseek.


> In my experience MTP's speed increase doesn't seem to justify the apparent loss of success at the edge

If you set manual MTP settings, you'll override dynamic adjustments the inference engine will try to do. Sometimes the dynamic adjustments aren't optimal. With the settings I use, MTP is always a net win.

> What exactly are you doing that the prompt is eating an entire 65536 window?

I'm not using the full context window.

> Surely it would be better to let it use any number of approaches that call tools to access that in parts and reason/summarize to an output file as it works through the whole thing.

Tools would not help.


> In my experience MTP's speed increase doesn't seem to justify the apparent loss of success at the edge, it would have to be at least 4x faster to meaningfully churn through the first 3 failures in the time it would have taken to do it once without

Speaking in terms of wall clock, the expensive part of decode is fetching the weights from memory. Predicting and validating a bunch of tokens using the already fetched weights is insignificant in comparison. Even if you have a poor acceptance rate for predictions, you won't really see a slowdown vs not using MTP.


Have you tried Muse 30B yet? I have been impressed with it. I have Qwen 3.8 27B hammering away right now against Muse. And Muse is doing a little bit better.


Vibes


> correctly reason through one of my private benchmarks

Want to say more about these private benchmarks? :)


seems to me like "private" is a good descriptor - I also have a set of "private" test cases - and they are kept private on purpose so they aren't scraped and fine-tuned on.


I have a sneaking suspicion that Qwen is fine-tuned on youtuber test cases (like Luke's Dev Lab, where Qwen 3.8 27B has just done almost eerily well).

Part of my suspicion is drawn from the thinking trace I got when I tested the car wash problem. That really does seem to have been post-trained; it's too good.

e.g. Gemma 4 26B solves this concisely without adding any filler about fuel economy or how long it will take, but it generally gets there by breaking down the problem in the thinking trace the way you'd expect.

Qwen 3.8 27B is just a little too certain right off the bat in low reasoning mode.


That clarifies it.


> Qwen 3.8 27B is only the second local model after Gemma 4 that managed to correctly reason through one of my private benchmarks.

I don't expect you to blab publicly about your private benchmark, but what sorts of reasoning does it require?


Well, I will say:

#1: it does not require deep world knowledge, because that's not what local models are for.

#2: it directly attacks drive-by understanding, overly linear processing training, poor attention mechanisms, poor reasoning patterns or lazy assumptions that ignore very easy low hanging fruit.

#3: it requires solid instruction following in the face of errors. a lot of models will run into errors and then fall back into some kind of error recovery process that bypasses instruction following.

#4: does not require prompt fine tuning to tweak to each individual model. they all seem to understand.

#5: not unfair. almost every model demonstrates in their reasoning that they have the necessary information that if reasoned about appropriately, could arrive at the correct answer.

#6: not designed to add unnecessary complication that it is intended to exhaust reasoning budgets of any sort, so it is not inherently unfair to models that reason a little more or less. for example, it does not require unnecessary reasoning soaks (ie: hiding the prompt inside base-64 encoding)

#7: has real world use and is probably applicable to overall ability to generalize.

#8: can be scaled up as models get better.

#9: is a very good indicator of how bad a model is falling apart under various inference settings.


I especially like #8. If you have some free time (don't we all have so much of that?) it would be really interesting to run a binary search on each model you have, to see at what size/complexity level it manages to solve the problem, say, 50% of the time.


Are you willing to share this benchmark’s internals? Kinda weird to expect folks to take you at your word without the ability to “trust but verify”


The nature of LLM benchmarking is that they seem to saturate public benchmarks so quick, they are a uniquely efficient case of https://en.wikipedia.org/wiki/Goodhart%27s_law

I'm not asking anyone to take my word, they can believe or not and in practice people should be taking signals from a variety of places and doing their own testing to see how models behave in their own use cases. What I'm measuring and why I'm measuring it may not be the most important metric for your specific use case.

Most other models are simply failing at these tasks. I think the tasks are relevant to overall model capability, but they are not the only metric. You don't give a jellyfish a tool and expect it to produce wonders, so the other capabilities of the model matter.


How much time did you invest in creating this benchmark? Any recommendations/resources you could give on how to do it?


It writes turing complete Beauty and the Beast fanfic.


I laughed so hard at that, thanks


:D, just upvoting this in case of someone downvotes


this is efficiency im looking for. artificialanalysis.ai model review not up. so considering output token per intelligence, do u think is it better than muse glimmer or no?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: