Which has nothing to do with the stalking of the adtech industry, unless you are suggesting that the EU might sell access to the keys to the likes of OpenAI?
[though you are not wrong that encryption backdoors are a backwards step on individuals rights to privacy]
What can the Kremlin do that is worse than the tech bros? Leverage data to deliver people additive divisive content until it all boils over into the "real world" and splits society? Oh wait...
It kind of does. It's like "I want an all powerful regulatory agency that can impose rules by dictat, but I only want them to do this things I agree with"
In the end, US doesn't need encryption backdoors because most chat and email protocols are not end-to-end encrypted and the TLS data streams are decrypted in the datacenter owned by US companies.
Chat control is merely their newest attempt at this, carefully navigating around the reasons it was opposed last year.
They keep trying to legislate encryption backdoors by any means, and once they manage to get their foot into the door every other country will use that as a precedent to also mandate it.
HBM isn't dramatically better in every dimension. You get huge bandwidth and good energy efficiency per bit transferred, but not necessarily a meaningful latency improvement and capacity expansion becomes tied to the package
It's not that there's a blocker. It's that it takes roughly 3x the manufacturing capacity to produce an HBM package at the same storage capacity as DRAM. We are sacrificing total bytes for bandwidth.
I don't fully understand the source of the "total bytes" constraint, but a major factor may be because HBM4 / HBM4E can only make use of the footprint directly above the processor/logic die (or in direct vicinity of its interconnect), while traditional DRAM can be placed further away where there's lots of real estate on the motherboard.
I gather a practical max ceiling today is a stack of 16 chips in height yielding 64GB?
These chips have a massive bus size of 2048 bits, instead of the 64 or 128 bits (dual channel) used by DDR5. That's what gives them their order-of-magnitude bandwidth speedup. But even though they technically pack in more capacity per square millimeter of motherboard, I gather they take up more space than older technologies once you account for the vias and interconnects to route all those signals.
Thanks for that, just went down an interesting rabbit hole. Many of us were hoping this re-tooling would eventually trickle some fast RAM down to DRAM-exhausted PCs, but given it would require a rearchitecture of the motherboard it's unlikely.
HBM4 has over 2048 signals to the processor’s PHY with tight signal integrity requirements that require the HBM stack to be < 0.5 mm from the processor die. That’s why HBM integration is done with interposers (soldered on the package). So, it’d be the CPU package that integrates it. Motherboard is too far away.
How so? FEOL is pretty much the same, BEOL is almost the same save TSVs, the packaging tech is different and more advanced, but not exactly 1:1 comparable. Do TSVs really occupy 3x the area of DDR IO's?
It doesn't really say that it needs to use 3x more area, but that 3x more gets consumed due to "advanced packaging and manfuacturing complexity". Which doesn't properly explain why it consumes 3x more and could simply mean they have a bad yield and 2/3 produced is garbage.
This might not be far off the mark. You are irreversibly linking the fates of these devices after a certain stage of manufacturing. If something goes wrong at final packaging time, you lose all dies instead of one.
The advantage of HBM over regular non-stacked DRAM is memory bandwidth, which also requires a super-wide memory bus - 2048 bits wide for HBM4. Compare that to the 128 bit wide bus of a modern CPU.
So to take advantage of it on the desktop, or anywhere else, you need that 2048 bit wide bus, and a processor capable of consuming 2-3 TB of data per second!
These are not normal requirements, other than for a GPU.
This will blow your mind, but it actually is pretty close to being unbounded. :)
Consider the 'MMA N matrices' primitive modern CPUs are starting to support. For the current generation of CPUs, N is a constant like 16 or 32, but there's nothing preventing it from being 1024 or larger if we have more memory bandwidth.
There is a world where scale and experience means it's about the same though, is the thing.
And memory is already expensive. It's downright hard to even get it though - you frequently would prefer not what's cheapest, but whatever is in largest scale production.
Scale and experience almost entirely share between HBM and normal memory. And they're both in large-enough scale production to not have a big difference on availability; if you're willing to pay HBM prices you should find even more sellers of DDR.
The only way I see HBM becoming competitive for consumer CPUs is if they solve the yield issues. Or if AI crashes so hard that people are putting those GPUs on fire sale and salvaging mass quantities of HBM off of them.
I think people might prefer the lower idle power consumption from lpddr over the better bandwidth in hbm in battery powered stuff. That said right now the price is definitely preventing us from finding out.
That hasn't been true since early HBM2 days, before the controllers standardized on power/voltage management and did things like leave them in P0 to ship on time
LPDDR/DDR/GDDR generally still win over HBM when there’s no data being transferred. HBM is primarily better in watts/byte transferred. Consumer electronics spend most of their time idle.
Right, but HBM idle is significantly worse than LPDDR idle or even DDR idle for that matter. That matters a lot precisely because the device is idle most of the time. Your idle power draw dominates.
Mainly bus width. Afaik HBM is like 1024 bits vs DDRs 64 so you need lots of transfers in parallel to saturate the bus, and CPUs kinda want 64 bytes of data as that's the size of a cache line ASAP. So you need a ton of in flight transfers which isn't a thing CPUs provide, maybe multicore workloads.
Buy the way you win with CPUs is with latency, and not bandwidth, which is why Apple M series actually uses DDR with lower latency because of the stacking.
Even if the cost were the same of the RAM itself, you'd need much bigger and more expensive CPU to deal with it.
Running 17 chrome tabs doesn't benefit at all from that HBM and all the additional hardware+software complexities that come with it. You want a specialized coprocessor to handle specialized workloads. The GPU exists separately from the CPU for a reason.
If they're talking about production capacity, that is some product of die area and process steps, right? It doesn't have to be 3x die area, just 3x lower factory throughput for the same number of functioning memory bits.
In one sense, nothing, in another, everything. It is DRAM, but the bandwidth requirements mean it’s paired to a processor, i.e. no DIMMs. Not 100% sure but things like MacBooks and the Framework tower, where you have fixed RAM for the device lifetime, have ~0 tradeoff.
Yup, the only reason Macs have higher memory bandwidth is because they use more memory channels, which gives them a wider bus. Both Intel and AMD only allow more than dual channel memory on server class processors these days.
The “Apple only does X better because they do Y” thing has been a meme for ages. I remember dismissals like “PowerPC is only faster at math because it has more integer units” or something along those lines, and thinking, uh, isn’t that a good thing?
There's a qualitative difference between "they're doing a different thing" and "they're doing the same thing, tuned differently". GP is saying that this is a case of "they just tuned it differently".
This distinction doesn't change what the performance numbers look like today, but it does inform what changes would be necessary for those numbers to look different tomorrow. E.g. Apple Silicon isn't fundamentally orders of magnitude more efficient than x86, they just used smaller features. Newer Intel and AMD chips made on equivalent processes _also_ get similar efficiency gains.
There are AMD and Intel devices on similar process (not talking about A20Pro or M6 which are set to ship later this week), and they do not get the same gains.
And honestly, they have historically had different markets.
When the design is for only one customer, you don't need to generalize things, and those things you generalize to give different customers different options has costs.
AMD will soon be a larger customer for TSMC than Apple (NVIDIA is already there) so Apple's pre-booking new processes is likely to be gone in the near future.
I am just saying it's not magic and x86 is capable of doing the same. Quad channel memory used to be more common in consumer hardware but now it looks like you don't even have the option anymore for desktops. AMD's Strix Halo was the first sign to reversing that (it has quad channel memory!) and hopefully we see more of that in the future.
AMD does market segmentation and limits consumer chips to dual-channel. You need to cough up the dough and get Threadripper for quad-channel. Or even more for Threadripper Pro to get octa-channel.
Kind of but not really. DDR5 splits your typical 64-bit channel into two 32-bit subchannels meaning the bus width is not increased. These subchannels are not always advertised because it's a just a feature of DDR5. Actually adding more channels increases bus width, which is what meaningfully improves memory bandwidth.
Keep in mind that non-Pro threadripper is still only 256 bits wide and Pro is 512. And the memory is 30% slower than with an M5. So an M5 Ultra has 3x the memory bandwidth of the best threadripper.
Moving memory around is a bottleneck. Non-unified memory usually means copying over the PCIe bus which is way slower than RAM (and way way slower than VRAM). Actually unified memory means you don't need to copy anything at all though which is the absolute best case for performance.
It isn't really, as long as you don't care about power consumption, physical constraints and money. Basically desktops <2025 (and hopefully >2027).
DDR is optimized for latency and stability at the cost of bandwidth whilst GDDR is optimized for bandwidth at the cost of latency and stability. GDDR is pushed so hard these days that a small percentage of errors is expected and corrected because this is still faster than running it slower but more accurate.
GDDR7 often has 10-20x the total bandwidth but 3x the latency of DDR5. Graphical workloads want as much bandwidth as possible but care relatively little for latency. Conversely, applications love low latency but don't really see any performance benefit from higher bandwidth.
So basicallyt you have workloads that are diametrically opposed and running unified memory forces you to compromise.
The price is that you essentially glue CPU and GPU together, which limits total compute, from a size and thermal perspective.
This is really not a limit because of unified memory -- in principle, PCIe GPUs could read/write main memory without the CPU. But it's a limit for /fast/ unified memory, because fast means close.
So unified memory is great as long as the integrated GPU is strong enough. Then it has two advantages:
a) probably faster transfer CPU<->GPU (but that's an implementation choice for the non-unified case
b) If you either need a lot of memory for the CPU or the GPU, but not for both at the same time, you pay for memory only once.
In theory perhaps but the benefit is not as relevant with a weak iGPU. In practice all PC's with performance ambitions had a dGPU until Strix Halo and Panther Lake.
More channels or soldered memory? Channels are basically RAID 0 so it depends what you're measuring. Soldering memory down was the only way to use LPDDR5X so if you wanted the best memory you had to solder it down. LPCAMM2 exists though so newer devices can use that instead of soldering them down, but not all devices would be able to fit the required LPCAMM2 slots.
That's tiny! But it still depends a lot on form factor. LPDDR5X is used in phones and making memory removable would have its compromises. You may even have compromises in laptops. Look at how tiny a MacBook Air's mainboard is and you'll see the RAM modules on the same package as the SoC. SOCAMM2 is too large for that but a variant with only two modules could possibly work.
The normal non-tech-savvy person already does this. They simply don't know that their ram is upgradeable or something else breaks first, before having to touch ram.
what is this too much RAM thing you mention? I thought the only valid RAM situation you could find yourself is not enough RAM. Too much? That's just fantasy
There is some truth (but not accuracy) to it, but not for this reason.
This article is about brainstem (aka hindbrain) vs the rest (midbrain + forebrain).
Evolutionally older animal lineages like fish, reptiles, birds all have midbrain + forebrain as well as the brainstem, so reptilian brain != brainstem.
Where there is some truth to our older "reptilian" brain is that the part of the forebrain called the pallium, that exists even in fish (& reptiles), has much expanded/evolved in mammals giving us our cortex and hippocampus, so you could kinda say that we have a reptilian brain + cortex. Note though that mammals didn't evolve from reptiles - they evolved in parallel from a common ancestor, so maybe better to say we have a fish brain + cortex (& hippocampus).
While the cortex is what gives mammals their intelligence, in birds the pallium evolved in a different way to give them theirs.
am no expert in brain development, still have read about evolution. I like aspects of the way you mentioned this, but seems that putting cortex separately from fish brain isn't quite right. As you mentioned, the cortex is simply a part of the fish brain that evolved much (much, much) larger than a fish's. Maybe:
Fish Brain = brain stem + small cortex (+ other parts of brain)
|
(evolved into)
|
V
Human Brain = brain stem + gigantic cortex (+ other parts of brain)
Sorry, know this is splitting hairs, but just thought it a useful distinction to make.
That's not quite right - a fish does not have a cortex.
The pallium of a fish and mammal are structured differently. A fish's (& bird's) pallium has a clustered (aka "nuclear") organization, while a mammal's has a mostly layered organization which is what is referred to as the cortex.
There is also more to a brain than just brainstem + pallium. In a vertebrate it's brainstem (aka hindbrain) + midbrain + forebrain, with the pallium being part of the forebrain.
Because we want people to write simple idiomatic code that runs fast today and also in ten years on alien hardware.
> Maybe I'm doing some rounding?
Compilers won't do this optimization when it is illegal to. If you disagree with the compiler's idea of what is legal, you can either write inline assembly, or put this code into an always inlined, but never optimized function.
It's the snake eating it's tail. For business profits to rise, people need to to have disposable income. If you replace enough people with AI, the average disposable income comes down to a point that businesses can't be profitable.
We are already seeing a localized version of this with restaurants in large cities.
For business profits to rise, people need
to to have disposable income
Does it matter, if the business pays employees who then use their salary to buy goods from the business or if the government collects taxes and uses that to pay for the goods? The government could also give the taxes to the people, so they have disposable income.
Imagine an island with two people on it and there is one tree that grows apples. And every other day, one of the two people climbs the tree and picks apples so they both have something to eat. Then the next day, the other person owes to climb on the tree because they are in debt to the first person.
Now, let's say one day, one of them discovers, hey, we can just give the tree a kick and the apples fall down on their own!
With the perspective you are proposing, the other one would have to say: Oh my God, don't do it! We will be out of work and starve!
Tax rates for wealthy folks are what, 10%? It makes a big difference if the majority of money that is made is distributed to the people via salaries, or whether it's profits and only 10% makes its way to the people.
Problem is there are more people who are not wealthy than there are wealthy people by a couple of orders of magnitude so even if you taxed the wealthy folks at 99%, you would not have enough to go around, if there isn't anyone to tax but wealthy folks. That is unless you think the slice of wealthy folks grows in size.
This is something the "eat the rich" crowd keeps missing.
Sure, you could take wealth away from billionaires, and sometimes there are good reasons to, but it's weird to pretend that it will cover more than a tiny tiny fraction of a government's budget.
Another thing some do is assume wealth accumulation happens on a yearly cycle. Like if you take Musk’s ~trillion today you’ll have another trillion to confiscate next year and the year after that, etc.
That's true locally but not globally - it means the economy will increasingly shift to serve those who still have a reason to bargain with each other, which will be the owners of the AI and other useful capital.
The version of windows you get for software development at a company that needs windows for software development is pretty good, no sign of any adware or bloatware, no phoning home, no sus telemetry or updates forced by microsoft.
I was laughing at the expected expectation of universally working UTF-8 just because it's 2026 when I read that. Soooo many things just do not work with UTF-8 as expected. I'm looking at MS Excel with heavy side eye
> Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
I'd classify mandatory encryption backdoors as an industry crisis rather than an annoyance.
reply