Rendered at 19:30:12 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
Traster 12 hours ago [-]
This semi-analysis article reads a lot more like an OpenAI press release than a real analysis. And to be honest some of the statements seem like just straight up lies - they initially claim they were invited to benchmark it, and then half way down switch to claiming that OpenAI provided all the numbers. This really kind of sucks, because I want to read actual detailed nuanced and credible analysis of what's happening in the industry and it doesn't seem like you can trust this as far as you can throw it.
automatic6131 11 hours ago [-]
Semianalysis is an AI hype organisation, not a serious, unbiased semiconductor reviewer/journalist like chipsandcheese nor a documentarian of the semiconductor industry like Asianonmetry (as it relates so strongly to the modern economies of Asia). If you see something from semianalysis, you can simply ignore it.
HarHarVeryFunny 3 hours ago [-]
Semianalysis seem to have useful information but increasingly crazy extrapolations of trends and future predictions.
It's helpful to realize that Dylan (Semianalysis), Dwarkesh, Aschebrenner (the Situational Awareness guy) and Sholto Douglas (Anthropic) all share a house in SF, so what you are getting from any of them is the SF AI scene view of the world, which is interesting to know, but probably not the best predictor of how things are going to pan out.
hermitShell 6 hours ago [-]
I tend to agree, but the Semianalysis + Dwarkesh side of reporting still does surface interesting information. You just have to take it all with a grain of salt, as indeed it is more hype focused, and look for real information hidden in the noise. And possibly to be a bit more entertained as you do so.
automatic6131 5 hours ago [-]
I suppose access journalism does have to get something out of the bargain, however small it's not zero. If that's the way you like to spend idle time, well it takes all kinds. I don't understand competitive scrabble either.
kimbler 7 hours ago [-]
As Jensen said, these guys are childish. They might have some technical chops somewhere in the organisation but their communication style of hubris + meme is really grating. I can't take a research organisation seriously when they so clearly want attention on X.
bwfan123 5 hours ago [-]
> some of the statements seem like just straight up lies
Why the angst ? I suspect this announcement punctured a lot of people's bubbles, and many are in disbelief and denial and hence the emotional reaction seen here. That a company which never designed chips could suddenly leapfrog the best in the industry. What many forget is that openAI and anthropic are in a unique position to own the end-user experience, and that provides them a distinct advantage. But, making announcements and actually delivering are two different things, and it remains to be seen if these are actually viable. In any case, it gives openai leverage over their vendors.
HarHarVeryFunny 4 hours ago [-]
A lot of the comparison is apples and oranges though - it seems that OpenAI's chip is targeting inference (FP8, FP4), while most of the chips it is being compared to are general purpose.
Notably the only one of the chips that also has a strong inference focus (but not only) is AMD's M1950X, which trounces OpenAI's chip (20 vs 3.4 FP8 PFLOPS, 40 vs 13.4 FP4 PFLOPS, 23 vs 15 TB/sec memory bandwidth), although it does use a lot more power (2500 vs 700W).
Google's TPU (now 8th generation) is glaringly absent from the performance comparison.
At the end of the day what really matters is cost not performance since you can always just run more chips. Google are full stack optimized from chip to data center, and might be expected to have an advantage.
fg137 9 hours ago [-]
The "better than" in the title should have given it away. You don't compare two sophisticated products that have different ecosystems and summarize your findings with such simplistic wording.
bjourne 9 hours ago [-]
It's very common in the industry today. They let some journos be the first to break the news. In return the journos write a glowing review. A symbiotic relationship.
mchusma 24 hours ago [-]
I think they talked about this being general purpose chip but I would think that Anthropic/OpenAI are at the scale now they could bake LLM weights into chips themselves.
For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.
I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
freakynit 16 hours ago [-]
In case anyone's interested in these niche startups like taalas, here are a few more:
Probably! But not viable yet; the chips would be about a year behind SOTA. Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing). We'll have to bootstrap our way there: AI is actively being used to get us closer to viable lead times for this.
Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a month, start to finish, for the physical processing.
Maybe once LLM improvements asymptote further?
kurthr 21 hours ago [-]
The metal masked ROM is basically only 2 metal/contact layers. It's not a full new design and tapeout. You could roll a new set of parameters every ~2-3months. It's not an architectural change. See statements below.
Part of the key is that by moving even from 6nm to 3-4nm one could embed a 20-30B model as part of a MoE (or only a subset of activated layers) on a single reticle die (note B300s are already multi-reticle), with a separate predictive/dispatch model controlling them each on a separate chip. This is without even stacking CiM ROM die. Moving the layer activations (and KV cache etc) between die requires relatively high speeds (and low latency), but distributed with multiple die in parallel might well be doable even with standard multilane PCIe. Of course KV cache prefill could also be handled by external GPUs. I'm sure AMD will make some reasonable choices.
MBCook 18 hours ago [-]
But that means your different chips all have different sets of weights and are different generations.
If none of that is baked into the chip as now then all the chips are running the latest weights every time.
Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.
kurthr 14 hours ago [-]
Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years. If you're still running H100s they draw >10x kWh/Mtoken as new designs. All of these systems become dated, but not all of them require entirely new infrastructure.
If a ROM rack running a near frontier agent model at >10ktoken/sec costs <$1M (rather than $4-8M for NVL72s) and draws only 10-20kW (rather than 100-200kW), and doesn't require a completely new cooling and power system every time you update? There'll be lots of demand for GLM5.3 in a year.
What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.
magicalist 3 hours ago [-]
> Since the current NVL72 are still at the ~15%/yr failure rate it's not clear your new data center is going to have half it's compute in 3 years
This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?
The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.
I'll note CoreWeave didn't commission the very first Blackwells until ~Jan 2025 so it hasn't been very long for RMAs. Furthermore, for training even when the NVL backplane can use the remaining GPUs a ~20% drop in performance means it isn't in the training cluster.
I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.
I don't want to out anyone, but these are similar comments:
Doesn't matter if the chip is 100x-1000x more efficient and faster, and you can just make a new one for new weights. Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful? Or Qwen 3.8 27B at 10k tokens/second. The super long thinking that makes qwen so effective would take a couple of seconds.
ben_w 7 hours ago [-]
> Imagine being able to run GPT Sol at 1k tokens/second a year from now, at a 100x lower cost per token than now. Would that be useful?
Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:
A similar trend is evident in the cost of models scoring above 50% on GPQA, a substantially more challenging benchmark than MMLU. There, inference costs declined from $15 per million tokens in May 2024 to $0.12 per million tokens by December 2024 (Phi 4).
15/0.12 -> factor of 125 cost reduction in 7 months.
But that may well be an extreme case. To show how broad the range is, another quote from the same publication:
Depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year.
WithinReason 6 hours ago [-]
This is from 2025, how about the last 6 months?
ben_w 6 hours ago [-]
You tell me. Most of these reports take that long to get published, or even longer. Sometimes I even see new-ish reports talking about 4o.
CrimsonRain 8 hours ago [-]
A model is not useless if it is not sota. Price, and speed are also important.
A 6 months old model that can run at 1/10th hardware and much faster too, can be much more capable than a sota model when you don't have unlimited budget.
geysersam 14 hours ago [-]
Why would it be useless in 3 months?
imtringued 13 hours ago [-]
Because on hacker news the only thing that matters is being in the current news cycle and not whether your business is profitable.
Certhas 14 hours ago [-]
Imagine Anthropic gives you Opus of 6 months ago but at much higher speeds and much lower cost (that they might or might not pass on).
Would you use it?
tyre 12 hours ago [-]
Yes, this would basically obviate Sonnet and Haiku. If you consider them 1 and 2 generations behind, respectively (that's not really what they are), you can still get a ton out of those older chips. Not to mention people still use older Opus versions happily. (In part because they don't like the new Opus but still, the cost effectiveness is a huge boon.)
user43928 9 hours ago [-]
GPT 5.6 Luna is already rather fast at 300 tokens per second, performs better than Opus 4.6 from what I know, and is very cheap.
I don't know why anyone would use Haiku.
6 hours ago [-]
vineyardmike 23 hours ago [-]
How much of that 16mo is design versus just production? If there was a “plug and play” chip where you just BYO weights, how long would it take?
The bigger issue seems to be that these chips can’t hold that many weights at the moment.
(I’m curious if chips with large weights in them would be more tolerant or less to yield issues. If you flip a few bits in the weights, does it really matter at scale?)
RealityVoid 21 hours ago [-]
Talaas, from what I understand is building stuff just like that. The infra is the same and the weights layer is all you need to change. I guess you could half etch the chips and then finish them with the weights only. I think their turnaround is 6-8 Weeks. The size of the models fitting on the chips at the moment is llama 3 I think?
derefr 21 hours ago [-]
> I guess you could half etch the chips and then finish them with the weights only.
It appears that to have working ASIC with the LLM baked into it we need to place and route macroblocks, and not a great variety of them. These macroblocks can be pre-placed-and-routed, available as masks already and shared between different LLMs.
Thus it appears that the tapeout delay can be substantially lower than a year.
tintor 17 hours ago [-]
They could etch the model architecture, without the weights into the chip.
This way newly post-trained model can be loaded and served the same day.
6 hours ago [-]
6 hours ago [-]
nerdsniper 16 hours ago [-]
> Maybe once LLM improvements asymptote further?
Maybe! But it also doesn't require the rate of improvement to slow down. As long as some current model is eventually "good enough" for general use, it could still be a market-killer at a very low marginal price thanks to ASIC. Even if slower, much more expensive models are 10x better, that doesn't actually diminish the utility of the ASIC model, as long as it's "good enough".
lelanthran 13 hours ago [-]
> Note the ~16 months that the article quotes as being insanely fast to get this chip to tape-out (read: start producing).
Yeah, with any luck it would put pressure on Nvidia to charge less, and not just to OpenAI. With a little more luck, we would see all the other players do the same thing, driving down the price of actual GPUs from GPU manufacturers.
kushie 24 hours ago [-]
tapeout could shrink but days per mask layer (DPML) does not have much margin..
smallmancontrov 18 hours ago [-]
I'm not in industry, is DPML (which I assume is the time required to make a mask?) set by electron beam scan time or something?
jeremyjh 22 hours ago [-]
I think Sol is already good enough though.
mcny 11 hours ago [-]
I don't think this is true at all. I use opus 5 to build a small project. It bothers me that people simply handwave away concerns so easily.
The context size is still not nearly as big as it needs to be to store all the code an enterprise needs. And apparently as you increase the context size, there are more defects. So none of this is a solved problem. We are still in early days and there is a lot to be done.
I'm sure there are marketing people who will say "coding is solved" and other such snake oil but none of this is done, far from it!
That being said there is still enormous value in older models especially with tool calling which will let them access the latest data. I feel like we need to be a little more careful and the tools should cache in a smart way to avoid rework but clearly if we could have opus 4.8 level of work for like a one time payment of a system for local LLM it will have value for years into the future.
So I agree in a weird way that sol is good enough for certain tasks but really there is a long road ahead.
basilgohar 19 hours ago [-]
"640k (token context) should be enough for anyone."
jerf 18 hours ago [-]
I know what you're saying, but modulo things like losing track of what year it is as time passes by, a current frontier model is going to continue to be useful for many tasks for many years, even moreso if it's 5-10x faster due to the chip architecture.
It's not that it would be the best forever, it's that it would be useful for plenty long enough to be worthwhile, even if there was better stuff available. In exactly the same way that this computer I'm typing this message on is not the latest and hottest cutting edge stuff. A 7 year old CPU, 7 year old Intel integrated graphics, an older NVMe disk, a mere 32GB of RAM... ok, that's one spec that's still pretty modern although it is slower RAM... but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge.
dgently7 16 hours ago [-]
exactly, but the "goes out of date" is bad when we talk about software.. but this isnt software, its hardware.
the youd have to buy a new one to get a better model is a FEATURE not a bug.
like if im apple... and i can put a sol level llm in an iphone, market it as privacy first you own your data personal assistant, integrate it all over the os... and then when there is a better model/siri make all the users buy a new phone... thats how they "win" ai.
the old standbys of better screens thinner cameras and batteries arent enough anymore. its basically tapped out. all modern phones are as thin as they need as big as they need as fast as they need and last all day on a battery...
apple needs a new number to up thing that people
can actually feel/see. model generations could be it... every year faster, smarter, more capbilities and integrations.
throwuxiytayq 17 hours ago [-]
> but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge
While it’s still too early to tell, I don’t think that’s how intelligence scales. Better models get you better solutions even to trivial problems. The ceiling for getting it done better is very high even if you’re not doing anything complicated. And difficulty isn’t uniformly distributed anyway - it seems to me that “mostly simple” tasks often have annoying 1% tails that low-intelligence models struggle with. I think we’ll see people chasing the top models for quite a while, or indefinitely - depending on the cost curve.
Marha01 10 hours ago [-]
Well, with proper compaction algorithm (like in Codex), it could indeed be useful for majority of tasks even in the future. The context length in Codex is just 272k, but it can reason well about much larger codebases due to good exploration and compaction algorithm.
I have yet to saturate the 1M context of Gemini, for example.
thoughtbefore 18 hours ago [-]
It may not matter. Think about why SOTA model companies are exploring chips. What do chips offer?
If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.
calebkaiser 17 hours ago [-]
Google rolled out TPUs in 2015. AWS released Inferentia and Trainium chips in 2020.
If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
edgyquant 6 hours ago [-]
Neither of those companies core business model was serving llms
calebkaiser 3 hours ago [-]
What? Both of those companies absolutely serve LLMs, and both of them would love for serving LLMs to be an even bigger part of their business. Not only that, AWS is Anthropic's primary compute partner for training and inference. They literally use the newest generation of the Trainium chips I mentioned before: https://www.anthropic.com/news/anthropic-amazon-compute
Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.
fooker 11 hours ago [-]
Computer Science history has taught us that this decision was almost always wrong.
Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.
aleph_minus_one 7 hours ago [-]
> Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.
This often happened, though not always.
An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned into modern GPUs.
Even today, GPUs live on as separate hardware in a computer instead of having become integrated into, say, the CPU. Intel's attempt to do something like this with the Larrabee project [1] was discontinued.
It didn't happen for the original fixed function graphics though, CPUs kept eating their lunch every year or so with fun rendering techniques. We still use some of these techniques for 2d graphics.
Only when they became more general with shaders, and then added support for GPGPU, did it truly take off.
I think this generality is the lesson here, not the fact that GPUs are not CPUs.
Aurornis 20 hours ago [-]
> I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
Taalas needed a giant chip (6nm) for an 8B model.
At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.
jubilanti 16 hours ago [-]
> Taalas needed a giant chip (6nm) for an 8B model.
You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon. Which is also not new, it goes back to the 1980s with fixed function digital signal processors and little linear regressions or hardware classifiers for industrial control systems, all are the same basic principle.
It's just usually not worth it to go super small process node, because most models people thought to turn into silicon were pretty small parameter sizes. We're talking 10-100 weight regression or at most 2-4k weight neural net, used in some instrument or factory equipment. You can do a decent MNIST OCR with a 4k weight neural net. For this, 180/130nm is fine.
Or you might think it's required with their special 4-bit as transistor thing (plausible). It's more that when you're experimenting and iterating, TSMC 6nm is their advertised path for rapid prototyping at cost for proof of concepts. And that's already in hot demand, while good luck if you're a startup trying to break in with 3/4nm as your first run.
Aurornis 14 hours ago [-]
> You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon.
It is.
As I said, they could have shrunk it with a smaller process node, but that's not at all close to what would be required for a GPT Sol size model.
greenknight 20 hours ago [-]
Nope. But we are hitting some pretty impressive levels with 128B models.
The other thing is, a lot of the time, model performance is improved with more 'thinking' time.
The thinking time is just more tokens... but instead of say 1000 tokens or 10,000 tokens worth of thinking its 1,000,000... how does that improve model performance? Could a 128B model hit levels of GPT Sol?
cherioo 20 hours ago [-]
Thinking generates a ton of tokens. These baked in chips tend to not have a lot of memory for context. I am not sure taalas supports Thinking at all.
The more problem like these they solve the more they will look like GPU.
nextaccountic 20 hours ago [-]
couldn't one just add some hundreds of GB of HBM?
kimixa 19 hours ago [-]
Yeah, but then there's the size of KV cache needing to be read through that HBM interface for each token, putting a hard limit on the tok/s based on the memory bandwidth.
On some models a large context can be a notable proportion of the size of the weights themselves.
For example, qwen 3.8 27b uses ~64kb/token for the kv cache - so for a 256k token context that's ~16gb of the kv cache for a ~54gb model (assuming 2 bytes-per-param/f16 for both).
So if the current non-baked-in chip is already memory bandwidth bound, as is often the case for current hardware and models, and the "only KV cache in HBM" chip has the same total memory bandwidth, it can only ever be (54/16)=~3.4x faster for the baked in-silicon model.
EDIT: I guess actually (54+16)/16=~4.3x faster, as the current implementation would need to read that KV cache too :)
andy_ppp 22 hours ago [-]
Yes, they could also sell me GPT Sol 5.6 or 5.7 on a chip and I’d probably buy it. It’s a really really useful model for me, I’m not sure how much better for coding I need it to be. For most things I find Sol good enough with a small amount of coaxing around my tastes.
structural 21 hours ago [-]
Keep in mind that what previous work has done on a single chip with weights baked in was on a 8b parameter model. Sol is likely something in the 5T parameter range, perhaps higher. Serving the whole thing at BF16 is on the order of $3m in hardware just to serve it at all, and closer to $1-1.5m of hardware if it was being served as NVFP4. And power draw starting at high tens to low hundreds of kilowatts.
Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower.
It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it.
dgently7 16 hours ago [-]
ill give you that the way we talk about this ppl seem to think wed do this tomorrow, but in the 70s a kb of ram took an entire rack and tons of power also. Its seems equally plausible that we could go into a cycle of iterative refinement of baked model hardware that would end up in "personal ai" just like we got to personal computing.
structural 2 hours ago [-]
Baked model hardware is not the next step in the chain here, in the next 2-3 years we might hopefully see some HBF (high bandwidth flash) hardware to try and get at same memory bandwidth today at a somewhat lower price point and much lower power dissipation.
Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.
Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.
nimchimpsky 21 hours ago [-]
[dead]
Caracas288 22 hours ago [-]
Man wouldn’t it be cool to be able to slot a massive ROM AI chip into the external AI drive of the pc…
pantelisk 18 hours ago [-]
It should look like a NES cartridge! That you have to blow on its end to clear any dust and it should do a satisfying click when it slots in.
Cooling might be an issue though...
xyzsparetimexyz 19 hours ago [-]
It'd just be pcie probably
porphyra 22 hours ago [-]
Also right now Sol 5.6 Max is super slow but if it were way faster on a chip (like Taalas' Llama 8b demo) then it would be an extreme value multiplier. But the model is so large that "baking it onto a chip" doesn't seem straightforward.
redox99 22 hours ago [-]
That'd be ungodly expensive.
mf_tomb 20 hours ago [-]
"Baking in" a model into a chip is a bad idea because chips take 2 years to tape out and then you're stuck doing inference on llama 3 in 2026 when fable/sol are available. Every accelerator is a tradeoff between flexibility and performance and GPUs are already pareto-optimal
guhcampos 18 hours ago [-]
There are other options. I worked for a startup called NVXL and we were programming DNNs into FPGAs using OpenCL, on custom boards we built to plug into NVME. It worked great, but it wasn't fast enough at the time to compete with Nvidia, or even Intel AVX512. Ultimately the company failed, but maybe some hybrid like that could work for LLMs? I haven't been up to date on how DNNs and LLMs look like under the hood these days, but there's got to be someone doing something similar.
twobitshifter 20 hours ago [-]
It depends when the good enough level hits. Pretty sure we are almost there for most common applications of AI.
dgacmu 19 hours ago [-]
That's only half the problem. OpenAI is contractually obligated, if you will, to believe that models will continue improving at an impressive rate for the foreseeable future (otherwise their valuation makes no sense).
If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to end up serving that Sol-class model several times more expensively than it will take someone who comes along a few months later. (such as what recently happened with DeepSeek's update.)
And under that assumption of continuing advancement, baking things in doesn't make sense in general - it's a play you'd make if you think things are slowing down a lot. Which may be right but it's not OpenAI or anthropic's play.
imtringued 13 hours ago [-]
If a strange and quirky architecture can increase their margins, they could start becoming so wildly profitable that they don't care about their investor driven valuation anymore.
adventured 18 hours ago [-]
Assume a $800 billion valuation. $100 billion ad network. $30 billion op income. 26x price to op income ratio. It's right there for them to grab, or someone else to grab.
Their valuation does make sense if you believe: 1) they can retain a massive user base and 2) a massive user base can be monetized. Future value is almost always pulled forward these days for high growth tech companies.
An LLM the size of Google search in users is even more valuable than Google search. The ad market for LLMs will be even larger than search was (no matter what HN prefers).
The monetization part is the easier part. Silicon Valley understands extraordinarily well how to build ad networks. If OpenAI maintain their gigantic user base, a $100 billion ad network is a given bolt-on. They'd have to screw that up in an epic way to not get there.
Facebook - Insta - WhatsApp is an absolute dogshit tandem with a gigantic user base. $228 billion in ad sales and still expanding 10% per year.
Google knows this is what's happening, that's why they don't care about chasing Anthropic very much. They're busy completely remaking how their core search business works.
Gigachad 17 hours ago [-]
Good enough will hit when the tech stops advancing quickly. You could have a "good enough" model but in 2 years if the general purpose chip can run it just as fast, there is no point having the single purpose one.
usef- 19 hours ago [-]
The whole point is that it's supposed to be more efficient. But models are also still getting absurdly more efficient every year, so you're likely nullifying much/most of the advantage. 18 months is a long time right now (and 18 is only time to tape out, not operational in data centers).
Even if the balance was net positive, you would also not be able to train them against new tools/harnesses or knowledge. How many years do you expect to keep using them?
adventured 19 hours ago [-]
The good enough level isn't ever arriving. We're in the first or second inning for LLMs. They will rapidly subdivide in complexity, they will not stagnate in the next decade.
Beyond the model, when would you freeze processor performance, such that it was good enough? Because that's exactly what freezing on Talaas is premised around.
The semiconductor technology will also continue to improve. You lose twice. Talaas is one of the dumbest ideas I've seen in semiconductors in decades.
murderfs 11 hours ago [-]
Closer to 2 weeks, as long as the new model fits: your model bits are entirely in mask rom, so you can change them with a metal only ECO that only touches two metal layers. You probably don't even need to run any timing analysis, etc. since all of the changes are going to be isolated in a gigantic square that's isolated from everything else.
pantalaimon 19 hours ago [-]
Well we'll see those surplus chips being repurposed for toys then.
Who wouldn't want a new Furby that can actually hold a conversation.
vunderba 19 hours ago [-]
I've seen several attempts even on HN of the LLM meets Teddy Ruxpin (or more accurately AG Talking Bear) but most of them offloaded the AI to some off-site servers.
I’d like to think that most parents would be weary of handing their children what basically amounts to a tape recorder that siphons all the data off to a large corporation.
OTOH, a completely local one (LLM + VAD + Speech Rec) would be a fun little thing to build.
Yeah this feels like Altman not knowing engineering well enough to realize where the focus should be.
Like he is optimizing to keep providing a vanilla token factory when weighted chips are coming and local models will supplement.
My head canon is savvy chip execs will be etching architecture his OpenAI pioneered into their flagship products while trying to minimize how much foothold he can get in hardware. Murica done offshored it. Not ours to control.
sebzim4500 24 hours ago [-]
My guess is we only see this once they start saturating computer use benchmarks. That's a use case which would be extremely valuable at the right costs/speed, but the current models just aren't there yet.
miki123211 12 hours ago [-]
I really want this for Whisper, particularly in some kind of a power-efficient, portable form factor.
I think people underestimate how much of a revolution having an always-on, privacy-preserving personal notetaker / secretary would be.
mike_hearn 11 hours ago [-]
AI is bigger than just LLMs. The real place model-hardwired chips will find value is in robotics. There you need very local, low latency inferencing with relatively stable models to handle motor control and navigation tasks. Higher level reasoning can be delegated to LLMs and run asynchronously.
walthamstow 10 hours ago [-]
I don't think any underestimates how powerful it would be. It's the social aspect that is difficult. I wouldn't want to be sat in a pub with you and your always-on notetaker.
16 hours ago [-]
andsoitis 18 hours ago [-]
> For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
but you trade updatability, which I don't think is worth it yet.
mchusma 17 hours ago [-]
Maybe! (1) Would SOL level intelligence be useful 3 years from now? 5 years? (2) would dedicated chips be the most affordable way to run this model in 3-5 years?
I suspect the answer to both of these questions is yes right now, but I agree it’s borderline.
andsoitis 17 hours ago [-]
3 years is an eternity.
imtringued 13 hours ago [-]
LoRAs are a parameter efficient way to update models. The fine tuning stopgap only has to work long enough to extend the lifespan of a model until the next chip comes out.
lqstuart 17 hours ago [-]
Eventually, someone is going to do this in Minecraft
bigfishrunning 6 hours ago [-]
Lol i'm surprised this hasn't happened already. Minecraft is a great platform for scripting as long as you think every single other platform is too fast and convenient.
raincole 17 hours ago [-]
It won't happen until IPO. If they do it now it'd be signaling that AI isn't improving fast.
fl0id 22 hours ago [-]
isn't that what they are doing with cerebras?
mkl 22 hours ago [-]
No, Cerebras holds the weights in SRAM - they are changeable, not baked in.
htrp 23 hours ago [-]
etched tried this.... it didn't go very well
anukin 21 hours ago [-]
I would assume asic based llm would work really well. Why did it not go well?
That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
mchusma 17 hours ago [-]
You are correct. I think this is the bull case. It seems like this would be useful right now for some things (eg moderation).
meowers1 11 hours ago [-]
[dead]
g00afthrowaway 14 hours ago [-]
Funny semi analysis has credibility here of all places. The founder is well-known in the hardware circle to be a black-market information trader.
It works like this:
1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles
2. Founder extracts insider information out of these interns
3. Founder sells this information to companies paying "consulting" fees
duchef 13 hours ago [-]
So you're saying the information is reliable?
amoss 11 hours ago [-]
Some of it is well sourced but the remainder is well sauced.
g00afthrowaway 12 hours ago [-]
Their negative reviews are more reliable than their positive reviews.
villgax 13 hours ago [-]
Basically the EvLeaks of AI lol
theideaofcoffee 4 hours ago [-]
So this is like the government attacking journalists for surfacing damning information, when in reality it's the government doing bad shit? Don't go after the interns doing the leaking, it's just easier to malign the messenger.
neevans 13 hours ago [-]
[dead]
epistasis 1 days ago [-]
It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
nxtfari 1 days ago [-]
Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.
kroaton 11 hours ago [-]
But Bonsai is garbage.
jacquesm 24 hours ago [-]
Ternary?
Razengan 8 hours ago [-]
Maybe we have to do what quaternions did for complex numbers and jump straight from 2 to 4
jeffbee 22 hours ago [-]
Knuth's base-e proposal enters the chat.
They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.
jacquesm 22 hours ago [-]
I can totally see how ternary would work from a physical implementation perspective but I have a really hard time visualizing anything using base-e, can you explain how such a thing would work in practice?
jeffbee 21 hours ago [-]
No it's impossible. But it would be optimal!
jacquesm 20 hours ago [-]
Ah, the spherical cow of number bases :) Thanks for the response, that saved me a sleepless night.
iFire 13 hours ago [-]
So out of all the inference only chips which ones can I buy?
The ASUS Store price for the ugen300-usb-8g costs $365.00 Canadian dollars.
corford 23 hours ago [-]
These nascent inference chip efforts are reminding me of the early 3dfx / riva / mach / powervr days. Will be interesting to see if inference chips are here to stay and, if so, who the eventual dominant player(s) will be
ehnto 23 hours ago [-]
Which in turn reminds me of Soundblaster audio cards! I suspect inference chips are closer to the GPU story than the Soundblaster story though.
I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb.
Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as graphics innovations did.
bayindirh 22 hours ago [-]
EAX was very powerful in its heyday, but it has died because of a thousand cuts.
First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features.
Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity.
Then Microsoft changed the Windows driver model, cutting the driver's direct access to the card. All of the timing sensitive effects were gone in an instant. I remember installing the new drivers and getting literally nothing. Sound Blaster was the only card with an hardware mixer, and Microsoft didn't feel like enabling them. Mixing at the DirectX layer killed the cards.
Soundblaster's very closed stance didn't help them either. None of the cards after Audigy2 worked with Linux when I had my desktop system.
After my Audigy2ZS, I moved to Asus Xonar D2X. Its positional audio capabilities were nice, but I mostly bought it for its Linux support and sound quality, and that was top notch in that regards.
Then sound cards became commodity. Everybody stopped making good cards. Musicians moved to audio interfaces, audiophiles moved to DACs.
Just looked to the SoundBlaster website. Internal cards are very limited. One DAC, one DTS enabled 7.1 sound card for PC cinema systems, three game oriented lower end cards, nothing else.
jonathanlydall 7 hours ago [-]
Creative Labs also turned itself into a brand I actively avoid for various reasons.
In early 2000s for example, friend with something like a Sound Blaster Live, but couldn't use it anymore as they lost the drivers, their website only had downloads for driver updates, requiring you to still have a driver CD, so no more CD meant no way to get the drivers.
They had not done very much meaningful stuff since the release of EAX and used patents to prevent anyone else from competing with them. Onboard sound cards were by and large indistinguishable from a quality perspective as Creative Labs ones, but cheaper which Creative Labs combated mostly with lawyers as opposed to upping their game.
Some guy after being frustrated with a long-standing bug in drivers for their Creative Labs sound card, dug into the binaries and made a fix for it. During which they also discovered that you could simply flip a switch in the driver to unlock features only meant to be available on more expensive hardware. Creative Labs of course went straight to lawyers to shut them down.
My brother bought the Creative Labs WoW headset which would have its mic get progressively softer until he would leave and re-join the call/voice chat room. They never released a driver update to fix this.
By the time that Microsoft announced no more "hardware acceleration" for sound cards, I had zero sympathy for Creative Labs, I was already convinced that they made pretty shoddy hardware/software and were mostly riding on their reputation from the 90s and some patents they managed to get.
thedougd 18 hours ago [-]
The DSPs that could double as a sound card and “soft” modem ruined their market in short order.
Razengan 8 hours ago [-]
Fucking Microsoft killed off a lot of cool shit with potential during the 1990s
jonathanlydall 7 hours ago [-]
As you mention the 90's you're probably referring to their embrace, extend, extinguish strategy. This situation with the sound card hardware mixing was quite different and a lot later.
Windows Vista moved away from kernel mode drivers where possible to help address the issue of BSODs which were not uncommon on Windows XP. It turns out that Microsoft was incorrectly getting the blame for BSODs when it was actually the fault of buggy video and sound card drivers.
While Vista was regarded as a "bad" Windows, it laid most of the foundations which allowed Windows 7 to be regarded as "really good" (by Windows standards).
From a hardware perspective, by the time Windows 7 arrived pretty much all drivers had been updated for Vista and had their kinks worked out (mostly, I would still have my NVidia drivers crash on occasion, but because they were user mode now, instead of a BSOD the screen would go black for a few seconds after which my desktop would come back with a Windows pop-up saying something to the effect of "the video drivers crashed and had to be restarted", the real perpetrator now being blamed!). Windows 7 also did optimizations to make things less resource intensive and it also helped that PCs had more RAM compared to when Vista came out.
On the UAC front Windows 7 was also much better, they calibrated the UAC prompts to come up less often, but what also happened is a lot of the 3rd party software which was needlessly requiring admin rights for no good reason (except that it was badly written), had finally been fixed by the time Windows 7 was released.
ryukoposting 16 hours ago [-]
Eeeeeh idk about the Sound Blaster comparison. Creative earned their place in the early-mid 90s solely because they were the one company making a sound card with drivers that actually worked properly.
It wasn't really the cool reverb effects or wave tables, though those were a nice bonus. It was just "I can tell my computer to make sound and it actually makes sound without days of troubleshooting."
Granted, similar things could be said about 3dfx. It's was a 3D card with drivers that actually worked.
And then there's the obvious "sound blasters and voodoos go in my computer, jalapeno goes in someone else's computer" thing.
noir_lord 22 hours ago [-]
On board got "good enough" and the separate cards died away.
In fairness on board (depending on the board but on the whole) is pretty good.
rubzah 5 hours ago [-]
I recently heard, for the first time, what Space Quest sounded like with a Roland board attached. It was mindblowingly amazing. And all most anyone ever experienced was bleep, bloop.
wmf 22 hours ago [-]
Every company is designing their own chips so the dominant players will be one level down: Broadcom, TSMC, Hynix/Samsung/Micron, etc.
adventured 18 hours ago [-]
There is drastically more power and profit in the software ultimately.
Apple is in the software first, the hardware second. Everyone at Apple has been trained to understand this for decades, and Jobs pointed it out endlessly. Apple's real moat is software (services, iOS, experience, MacOS).
Windows, Office, Azure, et al. Microsoft accumulated approximately one zillion dollars in profit on the back of software. It's a vastly superior business to anything hardware has traditionally seen. Nvidia is the first true juggernaut hardware profit machine, and the AI boom in extended hardware (RAM, storage) will prove temporary (even if there is a feast during that time). Microsoft's advantage and moat was Windows-Office for decades. It was a far better business than Intel's chip biz.
Google is a software company first. Every aspect of what made them and maintains them is software first, hardware second. They're a $400 billion software company. Their ad machine is software. Search is software.
Facebook is software. Instagram is software. WhatsApp is software. A $200 billion software company. They're not selling hardware, they're selling ads via software, they're monetizing users that use their software.
AWS is at least half software as an entity in terms of complexity, competitive advantage, et al. That's a two trillion dollar business.
LLMs can run successfully with various hardware approaches. The software is the value at the end of this, regardless of the hardware under it. The sole exception so far that may be sustainable is Nvidia, and we'll see if the bottom falls out from under that margin monster (China, specialized AI chips, whatever it happens to be that cuts under them massively).
Hardware always gets its margin squeezed eventually because it's a manufactured good (with inventory, fabs, etc). Software is hyper margin by default, you have to layer a lot of garbage on top of it to kill the margin. Nvidia is 33 years old, they have had a rich business for three years, that's it.
The AI boom is the sole reason anything in hardware has looked great in the past 20 years. Check the margins & op income for the top 20 hardware companies, from TI to AMD to Intel to Nvidia to Micron to Sandisk to Samsung to TSMC to ASML, prior to the AI boom of the past couple years. It won't last indefinitely. And after the return to a more normal environment happens, the hyper margins in software will persist.
thimabi 21 hours ago [-]
> Will be interesting to see if inference chips are here to stay
To me, the efficiency gains of inference chips are so significant that they are certainly here to stay — barring a revolution of sorts that leads to a world devoid of AI as we know it.
conradfr 10 hours ago [-]
Also the PhysX cards. Not sure how long they were useful.
ignoramous 22 hours ago [-]
> who the eventual dominant player(s) will be
This couldn't have been easy. The team at OpenAI has worked a miracle.
For example, Meta and Microsoft’s AI ASIC programs not getting off the ground despite being at it for much longer shows that cost is only one part of the equation.
rustystump 21 hours ago [-]
I bet cost is of no issue with the capx where it is at. It is almost certainly organizational. Meta throws money at every problem and it never seems to workout for them.
madaxe_again 12 hours ago [-]
I’d go with “whoever owns the entire vertically integrated ecosystem” - which is where the straight GPU comparison falls flat.
fraboniface 1 days ago [-]
I hadn't seen the token/Joules comparison with human speech before. Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
Phemist 24 hours ago [-]
The 20W number includes EVERYTHING else the brain does. The chips/models are literally only producing tokens. Let's see an LLM drive a robot harness and have the robot produce speech, as well as move through 3D space, keep track of metabolic needs, etc. etc. etc. before we compare efficiencies. That is even assuming the tokens are of equal quality. This comparison is currently Apples and Oranges.
phoghed 23 hours ago [-]
kind of a moot point if you can't get your brain to not do everything else. I think it's a fun comparison, even if it's not a 100% equivalence.
cmrdporcupine 21 hours ago [-]
Right, I can do the talked about ~3 tok/sec output and drive a car, hold my bladder, and eat chips at the same time.
Take that, Jalapeno!
falcor84 19 hours ago [-]
For what it's worth, LLMs don't really suffer from incontinence, so at least that part is pretty much a solved problem.
pantalaimon 19 hours ago [-]
They sometimes leak their system prompt
undersuit 18 hours ago [-]
So why are their water cooling systems filled with leak detectors? /s
perching_aix 11 hours ago [-]
> At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model.
This is about the same rate you get out of Sol Ultraspeed.
Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not a single user thing anymore, which is exactly what they have in the cooker with Astra.
Phemist 6 hours ago [-]
They are not. If the robot speech is a tool call, then for a fair comparison we need to take the tool call scaffolding (and probably the reasoning too) into account. So rather than a sentence of 10 tokens worth of speech being the output, the raw token output would be maybe 10x or 100x that. Even more if we consider the management of other aspects of the robot embodiment (or we reduce the brain's 20W number to whatever is actually required to produce coherent speech, sadly it is all rather entangled so this is not so easy).
perching_aix 3 hours ago [-]
But there are already voice models that do a reasonable job at a fraction of the throughput available?
The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much.
I half expect Boston Dynamics to show something like this off in Q4 or whatever.
Phemist 2 hours ago [-]
I am not arguing that there are perhaps other models that can run at the same quality, can coordinate between the different modalities, but are way less power hungry. My point is exactly about the comparison between the token output of the LLM running on the jalapeno chip, and sneaking in the power "usage" of the brain in the "token output" of human speech.
CooCooCaCha 23 hours ago [-]
And the brain is literally only producing electrochemical signals.
I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
Phemist 22 hours ago [-]
> I don’t see how tokens can’t produce speech or track metabolic needs.
It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.
DoctorOetker 22 hours ago [-]
I couldn't source the parameters from the screenshot or the nearby graphs, but from the nearby graphs you can see that at concurrency C=1, tokens/Joule (vertical axis) has totally plummeted, and obviously concurrent inference is much more efficient by batching. Divide the memory by the bandwidth and thats how long it takes to dump the full RAM contents through the chip. Do you want to do this once per token for a single conversation, or do you want to progress multiple conversations if you're going through all the weights anyway? The peak in the graphs is easily 22x more efficient than the low bottom right part on the graphs. So in batched mode its already more efficient than human speech.
nojs 20 hours ago [-]
> Humans are still 22x more efficient, which is not that far considering the rate of progress in this area.
Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison
xyzsparetimexyz 19 hours ago [-]
I do believe that this is the trade off. We are more efficient but slower in terms of thinking (at the same level of intelligence). Some animals go much further in terms of that trade off, see https://en.wikipedia.org/wiki/Portia_(spider) for example.
tesnorindian 8 hours ago [-]
Our brain is more like a MoE model activating only a few neurons for specific activities making it more efficient unlike a dense model activating all the params.
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
Patience pays.
freakynit 16 hours ago [-]
Just checked wikipedia page... they have like 100K neurons only.. wtf!!! How can nature cramp all senses, including spatial, motion, life maintenance and general thinking into just 100K neurons?
levocardia 13 hours ago [-]
A lot of the low level stuff is outsourced to biochemistry: the physical properties of proteins, and the various self-regulating biochemical systems of an animal, can "encode" a lot of intelligence, easing up on the computational demands of the brain proper.
freakynit 12 hours ago [-]
Hmm... is there something that can estimate how many bits of information in portia, each neuron+synapse combination might be encoding?
madaxe_again 11 hours ago [-]
100k is already a lot. Plenty of insects get by with a couple of thousand, and pack a whole bunch of complex behaviour, including flight, into that.
falcor84 19 hours ago [-]
What exactly are you questioning?
nojs 19 hours ago [-]
The claim that tok/s independent of quality is a useful comparison (I can get thousands of tok/s on a suitable small model), and secondarily that humans can’t output “tokens” faster than than in some sense, which I am less confident about
sinuhe69 14 hours ago [-]
22 times more efficient is not like 22 times more powerful. It’s extremely harder to close the gap in power efficiency than in raw power. Simply because the power efficiency we see is the result of billion years evolution optimization.
But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.
plasticchris 24 hours ago [-]
Probably not when you consider the training cost and upkeep expenses, not to mention the depreciation…
jstummbillig 22 hours ago [-]
At just inference! Which both a human and a model can not do without training, but while training rounds to zero for the model, for humans it scales linearly.
I am relatively certain we have already squarely been beaten in net efficiency at scale.
walrus01 18 hours ago [-]
Fairly amazing when you think about it, like human intellect can run on a bowl of rice and a chicken yakitori skewer.
saagarjha 21 hours ago [-]
You’re missing the factor for intelligence/token.
kemiller 23 hours ago [-]
I wonder how that stacks up if you consider all the time you have to keep the body alive when it’s not actively producing “tokens”.
jdiff 20 hours ago [-]
Careful, let's not put the whole matrix into stasis outside of business hours.
Productivity is not the only reason to let these meatbags burn oxygen.
danishanish 24 hours ago [-]
I mean, surely when quality is accounted for the difference is significantly higher
GaggiX 24 hours ago [-]
Or maybe significantly lower.
jimmySixDOF 1 days ago [-]
I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
tmp10423288442 1 days ago [-]
SemiAnalysis’ founder was roommates with Anthropic people, not OpenAI, so he may be slightly (very slightly) more objective here.
LogicFailsMe 22 hours ago [-]
Along with Leopold Aschenbrenner so maybe not so much.
rustystump 21 hours ago [-]
The guy that was part of FTX, fired from openai for alleged theft, got billions in a hedge fund somehow then lost billions. Why are all these people so scummy? It is like voting Trump three times in a row.
onion2k 17 hours ago [-]
They're very intelligent people who do very clever things at a young age, which draws the attention of very rich people who can exploit them to get richer, and no one tells the young person they're being exploited. They're heaped with praise and 'wealth' (millions, but crumbs compared to what they're making for other people), and told they're geniuses who can do no wrong, mostly by the media that happens to be owned by the rich.
Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.
And the cycle continues.
thelastgallon 12 hours ago [-]
The only effective altruism is making the mega-billionaires richer.
Bots do that because other platforms remove or hide posts with bad words
onion2k 17 hours ago [-]
Humans do it because they've been raised not to swear.
TiredOfLife 12 hours ago [-]
People that have been raised to not swear either do not swear or replace words.
onion2k 3 hours ago [-]
I f**ing do actually.
adabovehuman 13 hours ago [-]
> because they've been raised to grant advertisers and their vile spawn more rights than humans
xyzsparetimexyz 1 days ago [-]
s** posting? sex posting?
1 days ago [-]
msh 1 days ago [-]
shit posting
minimaltom 24 hours ago [-]
Thats what I thought too but then it would be s**?
jareklupinski 23 hours ago [-]
i see 'hunter2'
madspindel 24 hours ago [-]
s**?
Edit: OK, hn is removing one *
yjftsjthsd-h 24 hours ago [-]
If it's trying to convert it to italics, you may have to use a backslash to escape them
masfuerte 24 hours ago [-]
Or double them up: s****** gives s***.
2001zhaozhao 22 hours ago [-]
I like that to type s****** you had to type s************.
yjftsjthsd-h 20 hours ago [-]
Or escape them;)
18 hours ago [-]
subtlejellyfish 21 hours ago [-]
The "industry news and research" part of the AI industry feels very... suspect to me. My intuition is telling me that it's a bunch of people with influencer-y type social media skills and no actual credentials just grifting because there's so much money floating around.
senordevnyc 20 hours ago [-]
What credentials do you need to write a substack about an industry so it’s not grifting?
doctorpangloss 23 hours ago [-]
The semianalysis people have scripts which incorrectly count their numerators and denominators all the time. All their benchmarks are flawed. It is such a slipshod operation and they charge exorbitant amounts of money for it.
ShrigmaMale 21 hours ago [-]
Say more about this please
doctorpangloss 3 hours ago [-]
for example, their people think that GB300s are twice as fast as B300s, when really their benchmarks just incorrectly divide GB300 instances on azure by 8 instead of 4, since they don't read or verify any of the code that executes their benchmarks.
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
FrustratedMonky 1 days ago [-]
"not cut from the same cloth as Gartner McKinsey et al"
I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
latchkey 4 hours ago [-]
It isn't about bunk takes or not. It is about the motivation behind doing something.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
7thpower 13 hours ago [-]
This would have been far more effective with 1/10 as many words.
latchkey 4 hours ago [-]
Not really. It is a long story.
A_D_E_P_T 20 hours ago [-]
> McKinsey
lol. lmao even.
Have you seen the quality of their output? I'd take Claude or ChatGPT Free Tier over advice from McKinsey these days.
antonvs 1 days ago [-]
> I love how now you have to consider the possible s*** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods
I mean, previously you could have said something much the same except substitute "frat boys".
anthonypasq 1 days ago [-]
Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
kilroy123 1 days ago [-]
This is exactly what I see happening now.
Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.
CapsAdmin 14 hours ago [-]
Is this a normal thing now? I remember seeing this talked about as a surprising thing, but now I'm seeing posts about this as if it's normal.
(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)
sobellian 24 hours ago [-]
If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to the renaissance, mechanical work was much cheaper throughout the industrial revolution and remains cheaper to this day. We can still definitely say that the easier it is to produce tokens, the cheaper they will be.
vlovich123 17 hours ago [-]
All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies see with their customers (if they lower token prices consumers will use more tokens overall).
anthonypasq 1 days ago [-]
the total cost spent on tokens may go up, but i just cant imagine per token costs going up
jrflo 1 days ago [-]
Depends on compute capacity. If we become supply constrained on tokens, then prices will necessarily go up.
anthonypasq 24 hours ago [-]
no they dont because inference stacks are getting more efficient and models are getting more intelligent per parameter.
vlovich123 17 hours ago [-]
I would posit there’s no way in hell they’re getting sufficiently cheaper on a short enough time frame vs how much demand is sky rocketing. AI companies are seeing quarterly doubling of revenue if not more.
cactusplant7374 20 hours ago [-]
It is incredibly cheap now. What sectors are you thinking of?
holoduke 22 hours ago [-]
That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.
altmanaltman 24 hours ago [-]
I think you're reducing a very complex thing (the global economy) into a very simplistic model (Jevons' paradox) and thinking both are the same thing. This has no predictive power or rigor. You're just wishing things would happen as they did before, without considering that conditions and situations change significantly, and instead of Jevon's paradox, we look back at today 50 years from now and talk about Jensen's paradox.
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
goodmythical 24 hours ago [-]
[flagged]
senordevnyc 23 hours ago [-]
There is something counter-intuitive about the idea that making an engine that accomplishes the same amount of work with half the fuel will result in MORE fuel usage overall. You might expect it to be the same, or decline slightly, but the paradoxical element is that overall consumption goes up.
And you can say of course, it's so obvious, how could a dumdum not see that! But then there are lots of examples of things where increased efficiency results in less usage overall, because demand is inelastic, etc. Jevon's paradox doesn't apply to everything.
I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!
Imustaskforhelp 22 hours ago [-]
> I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!
Adding onto it, I feel as if this relates to some points regarding predictions of future in general. It is easier for us to look from the future to the past and think that it must be very obvious (as you also mention) but its also very counter-intuitive at the same time and there are just so so much nuance about basically any situation within it that its hard to really capture it all, and even then, be prepared for surprises and counter-intuitiveness.
I really like the Peter Drucker quote about it.
“The only thing we know about the future is that it will surprise us.” — Peter Drucker
and, “The future is fundamentally different from the past.”
— Frank Knight, Risk, Uncertainty and Profit (1921)
theobreuerweil 24 hours ago [-]
[dead]
dgellow 1 days ago [-]
There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.
dumberquestions 1 days ago [-]
The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.
jazzyjackson 24 hours ago [-]
Maybe 1000s of tokens per second unlocks realtime robotic decision making, and now every robot needs to continuously stream tokens to and from the cloud to operate. That could 1000x demand overnight, just to speculate :)
jacquesm 24 hours ago [-]
I would very much like it if anything that moves with appreciable mass is governed locally just in case the link drops and/or latency suddenly goes up. Motion is very unforgiving and accidents will happen if that's not taken into account.
dgellow 22 hours ago [-]
I think you just found what we will see in the S-1 prospectus of OpenAI
HDThoreaun 23 hours ago [-]
Seems unsafe to make locomotive decisions remotely
hypfer 24 hours ago [-]
Think about the agents buying computers for their agents. /s
simianwords 1 days ago [-]
The price has been going down for ages, its not clear what you are pointing at
phoghed 23 hours ago [-]
Pointing at the nay sayers who say tokens are heavily subsidized and it’s all going to come crashing down soon, surely any moment now
dgellow 22 hours ago [-]
I mean, it will obviously crash at some point. With so much pressure on token price to go down that means way less opportunity for margin for AI providers. OpenAI is in a pretty bad situation
simianwords 19 hours ago [-]
What does this have to do with margins? It can remain the same once prices go down
dgellow 22 hours ago [-]
At the price going down? And that it will continue to go down, even if the hardware improvements stop. Not sure what isn’t clear
stymaar 12 hours ago [-]
> continue to plummet.
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.
datakan 1 days ago [-]
Token prices coming down means nothing if the models keep wasting them
fg137 9 hours ago [-]
Counter point: Many AWS services barely decreased their prices (if at all) in the past decade despite advancement in hardware
m101 22 hours ago [-]
With the corollary that old hardware valuations will plummet with them.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
gwerbin 1 days ago [-]
Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
tmp10423288442 1 days ago [-]
Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
vlyan 1 days ago [-]
>polluting on-site generators
how much pollution do you believe modern gas-turbine engines to produce?
>Not to mention the water use controversy.
what percentage of US water usage do you believe is by AI data centers?
ilaksh 1 days ago [-]
Yeah but is it really even as good as Rubin? Seems just competitive.
mathisfun123 1 days ago [-]
this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.
anthonypasq 1 days ago [-]
OpenAI just dropped the price of Luna by 80% and Sol by 20-30%
mathisfun123 1 days ago [-]
and amazon shipping used to be free without prime, and uber used to be cheaper than taxis, and airbnb used to be cheaper than hotels.
you really don't get it?
simianwords 24 hours ago [-]
almost every pure tech commodity has gone down in price
- gpus
- retail computers
- laptops
- ~gpu~ appliances like washing machines
- cloud computing
i think you don't get how economy usually works in tech
zirkonit 24 hours ago [-]
I'm especially enjoying how RAM and SSDs are going down in price.
asveikau 18 hours ago [-]
GPUs and memory have gone up in price. It's more expensive to buy a 1-2 year old video card than it was at launch, sometimes by a shockingly large factor. Laptop vendors have recently shipped flagship models with less memory than the previous model, because they can't match price expectations for a laptop.
fer 21 hours ago [-]
I was checking laptops today for an upgrade from the model I bought back in 2019 and it's not gonna happen from how cheap they are.
thefreeman 23 hours ago [-]
listing gpu's here is crazy considering the current prices
fl4regun 24 hours ago [-]
GPUs and laptops and memory and storage are all crazy expensive
mathisfun123 17 hours ago [-]
i figured out why this comment is so confusing: this is actually a message from the past, around 2020. either that or simianwords is a time traveler that arrived today and hasn't read the news yet.
spacephysics 1 days ago [-]
We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss
So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.
anthonypasq 1 days ago [-]
> We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss.
what makes you think this?
RealityVoid 24 hours ago [-]
Because everyone keeps saying this so it must be true. Real "it is known" kind of vibe with these statements.
polski-g 20 hours ago [-]
He's a subscription truther. There's loads of them. OpenAI's profit increases with each subscription that is cancelled. Pretty soon they'll have more profit than God.
simianwords 1 days ago [-]
Yes, I can bet on this happening. If anything, this is a net gain for consumers as it is a competitive market.
mathisfun123 1 days ago [-]
go ahead and bet: alibaba is a publicly traded company
nimchimpsky 21 hours ago [-]
[dead]
kaveh_h 5 hours ago [-]
NVIDIA acquired Groq and have already integrated their technology in their solution. It seems it can really increase inference performance per energy used. I don’t believe
The benchmark compares Jalapenjo together with this solution.
https://nvidianews.nvidia.com/news/nvidia-groq-3-lpx-now-in-...
ChoosesBarbecue 1 days ago [-]
This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
manquer 19 hours ago [-]
ASICs always do better than general purpose chips. General purpose chips is turtles and turtles of virtualization and have to consider 4+ decades of backward compatible instructions set support.
ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.
General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.
New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
wmf 22 hours ago [-]
Existing CPUs have been extremely optimized by ~6 competing, well-funded teams. I expect AI to accelerate things somewhat but it's not clear that there is any low-hanging fruit available for AI to find.
tecoholic 20 hours ago [-]
The reliance on Deepseek and Kimi as the benchmarks from every chip maker from NVIDIA to OpenAI is a good tell of where things are heading. In the next couple of years, hopefully we will have systems at home for everyday use and corporations can buy bulk from providers.
chabons 14 hours ago [-]
As opposed to closed-source models? Benchmarks for GPT Sol wouldn’t be particularly meaningful, as no one else can run the benchmark, and we don’t know what the exact model specs are.
Picking the best open source models is really the best they can do.
aurareturn 9 hours ago [-]
One of the major advantages of Nvidia GPUs is that they can do both training and inference.
If you make an inference only chip, you better be damn sure that it's significantly better than Nvidia's GPUs at it.
Otherwise, it's better to buy Nvidia' GPUs because they're more flexible. You can do a big training run, then use them for inference right after.
thebeardisred 1 days ago [-]
All of these words spilled and no mention of the ISA.
2 hours ago [-]
dragandj 23 hours ago [-]
That's because it's AI-slopped.
saagarjha 21 hours ago [-]
I don’t think this is public?
kumarski 7 hours ago [-]
lamb-labs.com built some cool stuff.... I'm inclined to believe we're going to see static model chips everywhere, but I'm no expert in software.
I was early at efabless.com well now chipfoundry.io - they've done about 800 chip tape outs.
They've been doing open source silicon tape outs for a decade plus.
Founder recently built this: https://nativechips.ai --- not involved but I'm inclined to believe it's the future of where the market is going. I'm skeptical of many of the AI chip design startups and whether they've actually taped out chips and how many and at what scale.
lelanthran 24 hours ago [-]
This means that they're going to want to IPO soon - this is good news for investors + they need the capital.
rsync 23 hours ago [-]
No, this is because they want to IPO soon.
If the chips weren't this compelling they would have something different to announce.
These are paperclip maximizers who just happen to wear human skin - there is no underlying premise nor ideological goal.
Alien1Being 10 hours ago [-]
WARNING: AI VENDOR HYPE
blt 11 hours ago [-]
It can't come soon enough that AI workloads get their own specialized hardware to free up the general-purpose devices for general-purpose usage again.
throwaw12 1 days ago [-]
Competition is good for all of us, we will get better and faster chips.
Or at least Nvidia GPUs will become slightly cheaper for regular consumers again
WarmWash 24 hours ago [-]
That's if any datacenters are allowed to be built with them.
There is probably a ~50% chance that the next Dem candidate for presidency runs on a national datacenter moratorium or something equally as crippling.
bigyabai 19 hours ago [-]
If the populist campaign is to Make Affordable DRAM Again, then it's not a terrible solution.
The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?
zzzoom 17 hours ago [-]
No way to avoid the memory cartel, even if CXMT catches up.
einpoklum 23 hours ago [-]
If you think tanking Trillions in investments, warming the earth and increasion ocean water levels, creating water shortages and brown-outs is "good for all of us" - well, the rest of us beg to differ.
throwaw12 22 hours ago [-]
these GPUs make computation faster, I understand as of now maybe all the computation is used to generate yet another junk LinkedIn post or unnecessary RFC, but at some point this craze should settle and we will be left with powerful computation machines, which can be used for computing more useful things
cmrdporcupine 17 hours ago [-]
The GPUs being paid for w/ billions in investment will be obsolete and e-waste in a few short years same as a Cray-2 was just a decade after its release.
It's fine if you're one of the people selling shovels to gold miners for a while, but sucks to be building houses in the boom town?
throwaw12 9 hours ago [-]
I am not selling shovels, I just like to see when there is real competition on the market and I hate regulatory capture what Anthropic is trying to do for models, because they are scared of Chinese open weight models.
In terms of GPUs whole world with 8B people have only couple of viable options: Nvidia, AMD, Intel - and largest part of their inventory is going to enterprises to run those LLMs, and its impacting every consumer / hobby projects, like cheap phones, DIY electronics projects and so on.
I want to have more alternatives on the market
porridgeraisin 23 hours ago [-]
These are not replacing GPUs, they are entirely complementary. It's the same with cerebras, groq etc, they are all complementary to the GPU.
theandrewbailey 1 days ago [-]
The pricing of GPUs themselves aren't really the problem: it's the VRAM that comes with them.
m4rtink 20 hours ago [-]
So this will make GPUs and associated affordable for people, rigjt ?
jpollock 20 hours ago [-]
No, it's the wafer starts that are driving prices. Switching from Nvidia to custom doesn't change the constraints.
a2ff6eeb0 21 hours ago [-]
Sounds like a great way to get deals out of Nvidia.
chabons 14 hours ago [-]
Sure, but even heavily discounted Nvidia chips won’t be competitive for inference if they’re worse on perf/W.
redlewel 5 hours ago [-]
massive aura loss from a moronic name though
einpoklum 23 hours ago [-]
I hope the LLM wave will leave GPUs behind to go back to pursue more general-purpose computation rather than spending their die area on multiplying 4-bit-number matrices and such things.
acedTrex 22 hours ago [-]
Is that not literally the exect opposite of the direction asics for LLM inference is going?
einpoklum 8 hours ago [-]
My point is, that if companies develop ASICs for LLM work, then GPUs will stop being the go-to computation device for these workloads, and that will mean, hopefully, that their architectures will stop being warped so as to cater to LLM work.
bjourne 20 hours ago [-]
The article is a bit naive:
> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.
How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?
dist-epoch 9 hours ago [-]
In the slides on twitter you can see Jalapeno CAN do speculative decoding. In fact they explicitly mention how compute is disaggregated 3 ways now: prefill, predict, decode, and how a huge Jalapeno advantage is that it uses dark sillicon to switch between these without having to move the KV cache which remains local.
bjourne 9 hours ago [-]
Then the article contradicts the slides because it states that OpenAI choose not to disaggregate prefill and decode. Idk you men with "predict"---conventional LLM serving comprises only two phases.
dist-epoch 8 hours ago [-]
sorry, my mistake, I meant draft not predict
> it states that OpenAI choose not to disaggregate prefill and decode
They disaggregate INSIDE the chip, not by having separate machines for the 3 phases. the slides:
I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.
So, I started digging while waiting for various day-job inference calls to return, ha.
Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]
In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]
Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.
The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]
And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]
I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.
The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.
Not publicly acknowledging how misallocation of capital may be happening today shows he is corrupt. He’s not that dumb to not know it’s a major risk to the whole story, and is certainly financially incentivised to write as he does.
newyankee 18 hours ago [-]
If you research the origins of Dwarkesh, even more conspiracy level points emerge. I do not think even in the handful videos post Leopold fund collapse he has addressed it in any ways. That is the point of so called observers, they pretend to be impartial but everyone is a hustler in some way.
g00afthrowaway 14 hours ago [-]
See my other comment about this person. I'm surprised people take this site seriously.
mkw5053 21 hours ago [-]
I genuinely curious who’s downvoting me and why. I do not understand this forum sometimes.
empath75 1 days ago [-]
When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.
Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
impossiblefork 23 hours ago [-]
I don't agree. At the moment companies like NVIDIA take several times what it costs to make a chip. I think the fair split for the technology contribution is more like 50-50, maybe even 30-70 in favour of the manufacturer.
With competition we will actually have the fair split, whatever that is, and thus much lower prices.
At the moment, to have a big AI firm, or really AI firm at all, you need to be blessed by NVIDIA, in the form of receiving circular financing for your compute. They know that their prices aren't fair, or competitive.
Commoditization of inference is the end of that. The end of the mega-premium on inference hardware, and it's good not only for people who like running their LLMs, but it's the first step towards commoditization of training.
kubb 22 hours ago [-]
10-90 is the fair split.
impossiblefork 9 hours ago [-]
One could hope so, or it would be best for me if it were. I don't think I can count on that though. I don't think I can hope for anything better than 30-70 in my lifetime, and I think 10-90 almost requires you to buy the design firm.
simianwords 1 days ago [-]
I don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from.
If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.
skhameneh 22 hours ago [-]
> Its not like Steel which is more or less the same no matter where you purchase it from.
I’m not an expert in metallurgy by any means, but this seems really off. There are many recipes for steel and varied processes that also impact the final product.
lelanthran 1 days ago [-]
> I don't believe models will be commodified because each model is unique with strengths and weaknesses.
They are all converging.
zurfer 24 hours ago [-]
The same level of intelligence gets roughly 10x cheaper per year. So you might both be correct where a large part are commodity tasks but frontier is hard and valuable and not commodities.
airspresso 1 days ago [-]
This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.
imtringued 11 hours ago [-]
Steel has different varieties like carbon steel, alloy steel, stainless steel, and tool steel.
Each of those categories then has different grades of quality.
Tokens are a lot more like steel than oil, especially since a lot of tokens are used as structural material in the form of code.
calldacopsidgaf 21 hours ago [-]
Any article that features Sam's fucking creepy face should be marked with a jumpscare warning
philipwhiuk 9 hours ago [-]
In the near term , the real question isn't whether it's better - it's whether they stop buying as many NVIDIA chips as they can.
arrty88 20 hours ago [-]
Is this bad news for Cerebras?
villgax 4 hours ago [-]
obviously, Anthropic is gonna do their own chips as well just like openai, all hyperscalers already have their own chip stacks & neoclouds wouldnt even carry this as just as they are not carrying AMD stack lol
gbraad 10 hours ago [-]
... better than Blackwell in this specific case; which can also lead to OpenAI creating models that will only work on their own chips; vendor lock-in. When will they rename themselves?
7e 17 hours ago [-]
Once again we see the classic PR hype machine tactic of comparing a newer chip which is only available as an engineering sample to other chip designs which are widely available and much older.
They also fawn over the chip’s TDP when all other chips have to support 16 bit floating point and thus must run much hotter.
They make the classic mistake of equating max TDP with in-use-watts, and praise this magnificent (fictitious) performance per watt at FP8 with other chips’ max-TDP at FP16, which draw twice the power.
Evidence that the IPO can’t be far away.
Alien1Being 23 hours ago [-]
WARNING AI HYPE
dev1ycan 10 hours ago [-]
"human speaks at... tokens"
Yeah, stopped reading there, this is obviously some deranged sam altman paid blog post, I can't wait for the bubble to pop just so his newly launched chip falls flat on his face.
dkhid 16 hours ago [-]
R@6.....111
dkhid 16 hours ago [-]
hacker
crate_88 7 hours ago [-]
[dead]
shelldon42 8 hours ago [-]
[dead]
luciana1u 21 hours ago [-]
[flagged]
0xbadcafebee 1 days ago [-]
Story says they're power limited. That's half-true. Actually they're water-limited. To generate power, you need water. To cool chips, you need water. If you try to use less water on one side, you need more water on the other side (it's physics ya'll, making and using energy generates heat which requires dissipation). The world's freshwater is diminishing while also being consumed at an alarming rate. The future AI oligarchs are whoever controls the most water.
The other side of the conversation is the idea that large models in DCs on custom silicon is the future. Maybe for enterprise? But consumers will eventually (10 yrs) have affordable hardware designed to run crazy-good local models (more RAM + higher bandwidth). That will take pressure off of datacenters, but also reduce AI profits, and move that money to consumer chip/device makers. Apple is once again the biggest winner. Nvidia consumer chips might get cheaper, but nerfed, to encourage datacenter use where they make more money. I'm hoping AMD can stop being terrible at software so that when we finally have their better hardware we can actually use it.
minimaltom 24 hours ago [-]
For datacenters specifically I've never understood what specifically consumes the water. Arent the water-cooling loops closed, so the water just cycles around and around and around?
SirMaster 24 hours ago [-]
They evaporate the water which is what makes it cool so effeciently.
Ductapemaster 23 hours ago [-]
Evaporative cooling does not necessitate an open loop system
Eisenstein 22 hours ago [-]
The system which runs coolant over the chips can be closed but the part which uses an evaporative system to cool that is still open loop and vents water into the air, no?
Evaporated water is condensed, and in the process transfers its heat into another place that removes it. Another simple example is a pot of boiling water with a lid on it.
Eisenstein 18 hours ago [-]
The link you cited is not evaporative cooling and a pot of boiling water with a sealed lid on it is a pressure vessel which eventually explodes.
Ductapemaster 18 hours ago [-]
If you were to remove the heat at a sufficient rate by, say, turning the lid into a heat exchanger, you would have a stable system.
gorbypark 9 hours ago [-]
That's the problem, removing heat at a sufficient rate. Of course it can be done, but the most efficient way (in terms of cost) is just open loop evaporation.
I'm not a datacenter engineer, but I used to work in the ski industry. Snowmaking systems use vast quantities of compressed air. It works better if that air is cool. Blowing hot compressed air out of a snow cannon means the air temperature (wet bulb to be specific) needs to be colder to make snow.
Anyways, most air compression stations use water to cool the air, and then evaporative coolers to cool the water. The water is reused, but a ton (not sure of the percentage) is lost into the air. It's more or less a tower with a big fan on top, and water percolates down from the top, being cooled by the air as it goes. The water is then collected and pumped through the system again (but of course has to be always topped up to counteract what was lost to evaporation).
Anyways, long story short is it's most cost effective to just spray water into the air to cool water, as long as water is free/cheap.
Eisenstein 18 hours ago [-]
How is that different from not using evaporative cooling and just putting the heat exchanger on the burner?
Ductapemaster 5 hours ago [-]
Water is being used to get heat from one place to another. The idea is being able to separate the heat generation and the heat extraction
Eisenstein 5 hours ago [-]
You are adding an extra step though. Use a closed loop coolant and remove the heat from that coolant with the heat exchanger. Why would evaporation in between be more efficient?
justincormack 24 hours ago [-]
Yes they are for water cooling.
0xbadcafebee 17 hours ago [-]
At the datacenter side, it depends on the method of cooling. You can chill the air or the chips directly (or both), doesn't matter, you still need to cool, and that still needs water. The question is, where is the water being used?
- If they use either evaporative cooling or a liquid-cooled heat exchanger, that uses tons of water consistently. This requires less energy (it's mostly passive) so you use more water.
- If they use closed-loop water cooling and/or heat pumps/electric chillers, that uses much less water - at the DC. But it does require more energy to circulate the water, run fans, etc. If you are using more energy, where is the energy coming from? It's coming from power plants, which require... you guessed it... more water (e.g. thermoelectric, hydroelectric, geothermal, concentrated solar). They need water in order to generate the power, and lots of it. Coal, natural gas, nuclear, and concentrated solar, all use steam to generate energy. Nuclear also uses water to cool the reactor. And water is used extensively to extract coal, oil, and natural gas. Geothermal uses water in the ground.
You can't not use a ton of water in one fashion or another. It just depends what method, and on what end the water is used. And the crazy thing is, most new datacenters are being built in places with extremely little water. Guess how that's gonna work out as the planet gets hotter?
I don't know why I got downvoted to hell for stating facts every datacenter architect knows. HN be HN'in.
WarmWash 24 hours ago [-]
This only makes sense if you never looked at comparative water usage rates and available water.
simianwords 1 days ago [-]
How can OpenAI mass produce this chip at scale more economically than Nvidia which has experience in the supply chain and scale efficiencies to do it efficiently?
dpe82 1 days ago [-]
NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost.
One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.
vntok 24 hours ago [-]
> NVidia has enormous operating margins, so a competitive solution doesn't have to match or beat NVidia's scale efficiencies; it just has to beat delivered cost.
But then that means you have no actual moat against the behemot, right? Your competitor can move into the market as soon as they want to, at much better cost (so at slightly better price)... and Nvidia certainly can adapt much faster around hard hardware specs innovation than a new entrant ever could.
dpe82 21 hours ago [-]
Those are not OpenAI's concerns - they just need to scare NVidia enough to lower their prices more than they'd otherwise want.
acedTrex 22 hours ago [-]
But they are also PURCHASING from nvidia so any time nvidia lowers their prices they save money.
chris_money202 1 days ago [-]
In the short and medium term, it probably won't be more economical to produce for OpenAI. Where OpenAI is benefitting from their own chip is being able to tailor it to their models and workloads. When you buy off the shelf Nvidia, its not perfectly tailored and OpenAI has to spend marginally more to run off that chip. At the scale OpenAI is operating at and plans to operate at, that margin becomes pretty big $$
aurareturn 19 hours ago [-]
Because it's actually Broadcom that is doing most of the work.
airspresso 1 days ago [-]
By leveraging the experience Broadcom has in this area. Still remains to be seen how that goes when they want to scale production.
toasterlovin 1 days ago [-]
Replace OpenAI with Apple and Nvidia with Intel.
VirusNewbie 19 hours ago [-]
"How can OpenAI produce a LLM at scale more economically than Google, Amazon, and Microsoft which have experience in planet scale software and scale efficiencies unlike them".
One answer is they're quite good at poaching talent.
varispeed 1 days ago [-]
Why they don't research how to make their own RAM and they have to buy it from the common market?
They should GTFO with this crap.
Create barriers to computing for ordinary people while milking businesses for tokens.
petcat 1 days ago [-]
Building a custom-designed ASIC is much easier than producing state of the art memory chips.
There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.
chris_money202 24 hours ago [-]
Nvidia buys the memory it uses on its GPUs, same as all other ASICs.
To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
JV00 1 days ago [-]
Nvidia does not make RAM
varispeed 24 hours ago [-]
That doesn't excuse them from wrecking the market for ordinary person.
brcmthrowaway 1 days ago [-]
NVIDIA produces memory?
fc417fc802 1 days ago [-]
Fabless AFAIK. And that's the actual problem - drawing up CAD diagrams doesn't help if the factories are fully booked out.
Cyph0n 1 days ago [-]
A state of the art GPU is much harder to design & produce at scale and than an internal ASIC.
datakan 1 days ago [-]
People keep saying stuff like this without understanding what it takes to make RAM. It's one of, if not the most, heavily patented things in the world. The second you dip your toes into those waters the lawsuits begin.
If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.
Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.
chris_money202 1 days ago [-]
RAM chips are not hard to produce compared to many other types of semiconductors; Intel started in the memory game and left because the margins weren't great and they were going to fold. The failure rates on these chips are actually very tolerable; you can have a very bad yield and still have a viable chip due to things like ECC.
datakan 24 hours ago [-]
Intel entered the memory space because they partnered with Micron. They left the memory space when Micron pulled out of the partnership.
chris_money202 24 hours ago [-]
Intel started making DRAM in 1970, Micron was founded in 1978.
24 hours ago [-]
imtringued 11 hours ago [-]
Nah, DRAM is easy and very regular, it's a transistor and capacitor plus a massive decoder/encoder for addressing. The hard part is that it's a crushingly low margin business and without the added AI demand there were constant boom and bust cycles wiping out the manufacturers.
varispeed 24 hours ago [-]
Yes, it is difficult, but shafting working class is easy, therefor it is okay.
If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.
LarsDu88 1 days ago [-]
Well Sam Altman finally has built a moat against Chinese open weight AI. Well done. But what will this mean for Cerebras?
I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq
KaiserPro 23 hours ago [-]
> Well Sam Altman finally has built a moat against Chinese open weight AI
Hes got a press release.
The issue is, baking something to silicon requires discipline and about 2 years.
This isn't something you can just change your mind on halfway through. Trust me, I know. You need a clear vision of what you want to support, why and what bits of a chip you need to achieve that.
SV_BubbleTime 21 hours ago [-]
And yet, the top comment is about “hardcoding” weights into the silicon.
Man, if only someone made like, chips that could lots of different calculations all at the same time!
Eridrus 24 hours ago [-]
Cerebras is targeting a distinctly different point on the cost/latency curve. They are betting that there will be some high value applications where latency and not just throughput is super important.
porridgeraisin 23 hours ago [-]
It is being used as part of a combined system. For example AWS is pushing for Trainium + WSE 3. The WSE 3 does the decode and the Trainium does the prefill.
Even in nvidia land rubin + LPU does a similar thing.
It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.
Eridrus 21 hours ago [-]
AFAIK You can use WSE/LPU for prefill, it's just less efficient to do so.
porridgeraisin 11 hours ago [-]
Well ya, that efficiency is why it's split.
There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with
segmondy 21 hours ago [-]
I think the Chinese are going to be building their own chips aided with AI. DeepSeek, z.AI, MiniMax, Moonshot, etc, it's a race. The take off has really started.
epolanski 23 hours ago [-]
> and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies
That sounds quite like...nonsense?
Chip companies work on years-long cycles. They know today what are they launching 4-5 years from now.
It's helpful to realize that Dylan (Semianalysis), Dwarkesh, Aschebrenner (the Situational Awareness guy) and Sholto Douglas (Anthropic) all share a house in SF, so what you are getting from any of them is the SF AI scene view of the world, which is interesting to know, but probably not the best predictor of how things are going to pan out.
Why the angst ? I suspect this announcement punctured a lot of people's bubbles, and many are in disbelief and denial and hence the emotional reaction seen here. That a company which never designed chips could suddenly leapfrog the best in the industry. What many forget is that openAI and anthropic are in a unique position to own the end-user experience, and that provides them a distinct advantage. But, making announcements and actually delivering are two different things, and it remains to be seen if these are actually viable. In any case, it gives openai leverage over their vendors.
Notably the only one of the chips that also has a strong inference focus (but not only) is AMD's M1950X, which trounces OpenAI's chip (20 vs 3.4 FP8 PFLOPS, 40 vs 13.4 FP4 PFLOPS, 23 vs 15 TB/sec memory bandwidth), although it does use a lot more power (2500 vs 700W).
Google's TPU (now 8th generation) is glaringly absent from the performance comparison.
At the end of the day what really matters is cost not performance since you can always just run more chips. Google are full stack optimized from chip to data center, and might be expected to have an advantage.
For example, GPT Sol baked into a custom chip run for $100M that runs 10x as fast and 10x as cheap should pay for itself as long as the chip is useful for long enough.
While 2 years ago nothing was useful more than 1 year long, there are many older models in use now (e.g. Haiku 4.5, GPT-OSS 120b), and I expect this trend to continue.
I know this is what Taalas was doing (acquired by AMD), here was their demo, https://chatjimmy.ai/ which is based on Llama 3.1 8B. It feels like this should start to happen soon.
1. https://matx.com/
2. https://www.d-matrix.ai/
3. https://www.etched.com/
4. https://www.positron.ai/
5. https://hyperaccel.ai/
6. https://axelera.ai/
7. https://www.enchargeai.com/
8. https://furiosa.ai/
Unfortunately, there's some real physical constraints: IIRC, manufacturing a wafer takes on the order of a month, start to finish, for the physical processing.
Maybe once LLM improvements asymptote further?
https://www.eetimes.com/taalas-specializes-to-extremes-for-e...
https://www.turingpost.com/p/taalas
https://cambrian-ai.com/taalas-launches-hardcore-chip-with-i...
Part of the key is that by moving even from 6nm to 3-4nm one could embed a 20-30B model as part of a MoE (or only a subset of activated layers) on a single reticle die (note B300s are already multi-reticle), with a separate predictive/dispatch model controlling them each on a separate chip. This is without even stacking CiM ROM die. Moving the layer activations (and KV cache etc) between die requires relatively high speeds (and low latency), but distributed with multiple die in parallel might well be doable even with standard multilane PCIe. Of course KV cache prefill could also be handled by external GPUs. I'm sure AMD will make some reasonable choices.
If none of that is baked into the chip as now then all the chips are running the latest weights every time.
Even if you could ignore the stuff built into the chip when the time came, at that point you just wasted money on silicon that’s useless in 2-3 months.
If a ROM rack running a near frontier agent model at >10ktoken/sec costs <$1M (rather than $4-8M for NVL72s) and draws only 10-20kW (rather than 100-200kW), and doesn't require a completely new cooling and power system every time you update? There'll be lots of demand for GLM5.3 in a year.
What these don't do is TRAINING, they only do INFERENCE, but they could do it pretty well.
This would be a stupidly bad failure rate, basically the worst business decision you could make, especially if you're somehow on the hook for eating those losses (which seems to be the implication?). Is there a linkable source on this?
The only thing I could find is SemiAnalysis claims that 15% of Blackwells end up RMAed[1]. That appears to be a total failure rate, though, and if you're RMAing them, you're getting replacements. So that appears to be a pretty different state of things.
[1] https://www.dwarkesh.com/p/dylan-patel#:~:text=GPUs%20are%20...
I've personally heard this from several sources in the data centers (installers, training, network). It's not uncommon for 10% of racks to fail on delivery. I hear that's improved somewhat from GB200 to GB300, but the number of FW updates from the time they ship, until they're commissioned is >>10. If an HBM or GPU or backplane supply/cooling fails, it is basically not swappable or repairable. You have a "dead" rack, and deliveries are on allocation so you don't get a replacement for months (eg some "RMAs" for early delivered parts in late 2025 are still dead racks 9 months later). "Tray" swaps are technically possible, but still quite rare, perhaps because debugging takes as much time as commissioning a new rack.
I don't want to out anyone, but these are similar comments:
https://www.linkedin.com/posts/neelmaster1_aiinfrastructure-...
https://www.hostzealot.com/blog/news/nvidia-gb200-nvl72-is-n...
https://introl.com/blog/gb200-nvl72-deployment-72-gpu-liquid...
Given the current rate of change, it would be hard to guess either way. By some measures the cost at fixed quality score goes down vastly faster than that:
- https://hai.stanford.edu/assets/files/hai_ai-index-report-20...15/0.12 -> factor of 125 cost reduction in 7 months.
But that may well be an extreme case. To show how broad the range is, another quote from the same publication:
A 6 months old model that can run at 1/10th hardware and much faster too, can be much more capable than a sota model when you don't have unlimited budget.
Would you use it?
I don't know why anyone would use Haiku.
The bigger issue seems to be that these chips can’t hold that many weights at the moment.
(I’m curious if chips with large weights in them would be more tolerant or less to yield issues. If you flip a few bits in the weights, does it really matter at scale?)
Basically a https://en.wikipedia.org/wiki/Gate_array. (The non-field-programmable kind.)
It appears that to have working ASIC with the LLM baked into it we need to place and route macroblocks, and not a great variety of them. These macroblocks can be pre-placed-and-routed, available as masks already and shared between different LLMs.
Thus it appears that the tapeout delay can be substantially lower than a year.
This way newly post-trained model can be loaded and served the same day.
Maybe! But it also doesn't require the rate of improvement to slow down. As long as some current model is eventually "good enough" for general use, it could still be a market-killer at a very low marginal price thanks to ASIC. Even if slower, much more expensive models are 10x better, that doesn't actually diminish the utility of the ASIC model, as long as it's "good enough".
Yeah, with any luck it would put pressure on Nvidia to charge less, and not just to OpenAI. With a little more luck, we would see all the other players do the same thing, driving down the price of actual GPUs from GPU manufacturers.
The context size is still not nearly as big as it needs to be to store all the code an enterprise needs. And apparently as you increase the context size, there are more defects. So none of this is a solved problem. We are still in early days and there is a lot to be done.
I'm sure there are marketing people who will say "coding is solved" and other such snake oil but none of this is done, far from it!
That being said there is still enormous value in older models especially with tool calling which will let them access the latest data. I feel like we need to be a little more careful and the tools should cache in a smart way to avoid rework but clearly if we could have opus 4.8 level of work for like a one time payment of a system for local LLM it will have value for years into the future.
So I agree in a weird way that sol is good enough for certain tasks but really there is a long road ahead.
It's not that it would be the best forever, it's that it would be useful for plenty long enough to be worthwhile, even if there was better stuff available. In exactly the same way that this computer I'm typing this message on is not the latest and hottest cutting edge stuff. A 7 year old CPU, 7 year old Intel integrated graphics, an older NVMe disk, a mere 32GB of RAM... ok, that's one spec that's still pretty modern although it is slower RAM... but it's still plenty fast enough to comment on HN, even these seven years after it was cutting edge.
the youd have to buy a new one to get a better model is a FEATURE not a bug.
like if im apple... and i can put a sol level llm in an iphone, market it as privacy first you own your data personal assistant, integrate it all over the os... and then when there is a better model/siri make all the users buy a new phone... thats how they "win" ai.
the old standbys of better screens thinner cameras and batteries arent enough anymore. its basically tapped out. all modern phones are as thin as they need as big as they need as fast as they need and last all day on a battery...
apple needs a new number to up thing that people can actually feel/see. model generations could be it... every year faster, smarter, more capbilities and integrations.
While it’s still too early to tell, I don’t think that’s how intelligence scales. Better models get you better solutions even to trivial problems. The ceiling for getting it done better is very high even if you’re not doing anything complicated. And difficulty isn’t uniformly distributed anyway - it seems to me that “mostly simple” tasks often have annoying 1% tails that low-intelligence models struggle with. I think we’ll see people chasing the top models for quite a while, or indefinitely - depending on the cost curve.
I have yet to saturate the 1M context of Gemini, for example.
If SOTA models haven’t peaked, then the SOTA model companies would still be churning out better and better intelligence.
If companies working on ML-specific chips was evidence that large transformer models have fully saturated their potential, the field would have been done circa GPT-2.
Chips are another axis for improvements in training and inference. Orgs large enough to explore the space have been doing it for at least a decade now. This is just a silly line of reasoning based on the faulty assumption that somehow, looking for increases in efficiency in training/inference means teams have reached some theoretical limit in model capability.
Every time a company has spent resources doing this, a competitor innovated on the software and made the custom hardware irrelevant.
This often happened, though not always.
An important counterexample are 3D graphics cards, which basically put the OpenGL/DirectX fixed-function pipeline into silicon. It took a long time and many iterations to make the pipeline more programmable until the 3D graphics cards turned into modern GPUs.
Even today, GPUs live on as separate hardware in a computer instead of having become integrated into, say, the CPU. Intel's attempt to do something like this with the Larrabee project [1] was discontinued.
---
[1] https://en.wikipedia.org/w/index.php?title=Larrabee_(microar...
Only when they became more general with shaders, and then added support for GPGPU, did it truly take off.
I think this generality is the lesson here, not the fact that GPUs are not CPUs.
Taalas needed a giant chip (6nm) for an 8B model.
At best you could use a more advanced node to try to put a MoE model across several chips working together, but you can’t have GPT Sol size models on a single chip like that.
You're phrasing it like it was kind of an inherent technical limitation with this kind of burning weights into silicon. Which is also not new, it goes back to the 1980s with fixed function digital signal processors and little linear regressions or hardware classifiers for industrial control systems, all are the same basic principle.
It's just usually not worth it to go super small process node, because most models people thought to turn into silicon were pretty small parameter sizes. We're talking 10-100 weight regression or at most 2-4k weight neural net, used in some instrument or factory equipment. You can do a decent MNIST OCR with a 4k weight neural net. For this, 180/130nm is fine.
Or you might think it's required with their special 4-bit as transistor thing (plausible). It's more that when you're experimenting and iterating, TSMC 6nm is their advertised path for rapid prototyping at cost for proof of concepts. And that's already in hot demand, while good luck if you're a startup trying to break in with 3/4nm as your first run.
It is.
As I said, they could have shrunk it with a smaller process node, but that's not at all close to what would be required for a GPT Sol size model.
The other thing is, a lot of the time, model performance is improved with more 'thinking' time.
The thinking time is just more tokens... but instead of say 1000 tokens or 10,000 tokens worth of thinking its 1,000,000... how does that improve model performance? Could a 128B model hit levels of GPT Sol?
The more problem like these they solve the more they will look like GPU.
On some models a large context can be a notable proportion of the size of the weights themselves.
For example, qwen 3.8 27b uses ~64kb/token for the kv cache - so for a 256k token context that's ~16gb of the kv cache for a ~54gb model (assuming 2 bytes-per-param/f16 for both).
So if the current non-baked-in chip is already memory bandwidth bound, as is often the case for current hardware and models, and the "only KV cache in HBM" chip has the same total memory bandwidth, it can only ever be (54/16)=~3.4x faster for the baked in-silicon model.
EDIT: I guess actually (54+16)/16=~4.3x faster, as the current implementation would need to read that KV cache too :)
Let's say a magic set of chips comes along to host this. Maybe it's 2-3x more efficient in size and power. You're still talking a form factor that's a good chunk of a rack, draws tens of kilowatts, and could actually be sold at a similar if not higher price point because the OPEX is so much lower.
It may be useful but it's certainly uneconomic to spend >$1m to self host the model, plus ongoing power and maintenance costs, plus the cost to adapt whatever building you're in to be able to power it.
Something like the next iteration of Cerebras hardware paired with HBM for KV cache + HBF for weights could be incredibly strong here and much more likely to see away to make into a product with some lifetime compared to "let's bake a old model into a very, very large number of custom chips, design all the interconnects from scratch, and pray". Maybe in 10-15 years once this all matures.
Now if that works out that means in '29/'30 we could easily see a run on NAND that's even worse than the current DRAM price issues, on top of the current increases. Fun times if that happens.
Cooling might be an issue though...
If you believe that, then you should expect to get Sol-level performance out of a Luna-cost model within six months or a year. If you have a system with the weights baked in, that means you're going to end up serving that Sol-class model several times more expensively than it will take someone who comes along a few months later. (such as what recently happened with DeepSeek's update.)
And under that assumption of continuing advancement, baking things in doesn't make sense in general - it's a play you'd make if you think things are slowing down a lot. Which may be right but it's not OpenAI or anthropic's play.
Their valuation does make sense if you believe: 1) they can retain a massive user base and 2) a massive user base can be monetized. Future value is almost always pulled forward these days for high growth tech companies.
An LLM the size of Google search in users is even more valuable than Google search. The ad market for LLMs will be even larger than search was (no matter what HN prefers).
The monetization part is the easier part. Silicon Valley understands extraordinarily well how to build ad networks. If OpenAI maintain their gigantic user base, a $100 billion ad network is a given bolt-on. They'd have to screw that up in an epic way to not get there.
Facebook - Insta - WhatsApp is an absolute dogshit tandem with a gigantic user base. $228 billion in ad sales and still expanding 10% per year.
Google knows this is what's happening, that's why they don't care about chasing Anthropic very much. They're busy completely remaking how their core search business works.
Even if the balance was net positive, you would also not be able to train them against new tools/harnesses or knowledge. How many years do you expect to keep using them?
Beyond the model, when would you freeze processor performance, such that it was good enough? Because that's exactly what freezing on Talaas is premised around.
The semiconductor technology will also continue to improve. You lose twice. Talaas is one of the dumbest ideas I've seen in semiconductors in decades.
I’d like to think that most parents would be weary of handing their children what basically amounts to a tape recorder that siphons all the data off to a large corporation.
OTOH, a completely local one (LLM + VAD + Speech Rec) would be a fun little thing to build.
https://en.wikipedia.org/wiki/AG_Bear
Like he is optimizing to keep providing a vanilla token factory when weighted chips are coming and local models will supplement.
My head canon is savvy chip execs will be etching architecture his OpenAI pioneered into their flagship products while trying to minimize how much foothold he can get in hardware. Murica done offshored it. Not ours to control.
I think people underestimate how much of a revolution having an always-on, privacy-preserving personal notetaker / secretary would be.
but you trade updatability, which I don't think is worth it yet.
I suspect the answer to both of these questions is yes right now, but I agree it’s borderline.
That's quite a few parameters shy of today's trillion-weight behemoths, but it is fast.
It works like this:
1. Founder befriends undergrad interns/graduate student interns, buys them gifts, invite them to dinner/yacht/house/vc parties etc, or pays them to write articles 2. Founder extracts insider information out of these interns 3. Founder sells this information to companies paying "consulting" fees
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.
The only report is a smartnic fpga from Alibaba where we take an onnx design and write our own. https://essenceia.github.io/projects/alibaba_cloud_fpga/
On my M2 Pro Mac Mini the ANE only allows 2 gigabytes compared to the Metal GPU which can use the system ram.
Currently playing with https://www.asus.com/motherboards-components/ai-accelerator/... which is a 4bit, 8bit and 16 bit ai inference chip with 8 gigabytes of ram.
The UGen300 has the Hailo-10H chipset.
The ASUS Store price for the ugen300-usb-8g costs $365.00 Canadian dollars.
I remember one soundblaster card I bought came with a Lara Croft demo, that exploited the incredible immersion of real time dynamic reverb.
Genuinely I think game audio took a few steps back from that heady era, the innovation in audio likely didn't sell as many cards as graphics innovations did.
First we had to have the audio processor. Good EAX was available on top of the line cards, and they were not always cheap. Lower end chips got less features.
Then we had to have the speaker setup to have the greatest sound, or needed to get a real 5.1 headphones, which were bulky and never provided the same fidelity.
Then Microsoft changed the Windows driver model, cutting the driver's direct access to the card. All of the timing sensitive effects were gone in an instant. I remember installing the new drivers and getting literally nothing. Sound Blaster was the only card with an hardware mixer, and Microsoft didn't feel like enabling them. Mixing at the DirectX layer killed the cards.
Soundblaster's very closed stance didn't help them either. None of the cards after Audigy2 worked with Linux when I had my desktop system.
After my Audigy2ZS, I moved to Asus Xonar D2X. Its positional audio capabilities were nice, but I mostly bought it for its Linux support and sound quality, and that was top notch in that regards.
Then sound cards became commodity. Everybody stopped making good cards. Musicians moved to audio interfaces, audiophiles moved to DACs.
Just looked to the SoundBlaster website. Internal cards are very limited. One DAC, one DTS enabled 7.1 sound card for PC cinema systems, three game oriented lower end cards, nothing else.
In early 2000s for example, friend with something like a Sound Blaster Live, but couldn't use it anymore as they lost the drivers, their website only had downloads for driver updates, requiring you to still have a driver CD, so no more CD meant no way to get the drivers.
They had not done very much meaningful stuff since the release of EAX and used patents to prevent anyone else from competing with them. Onboard sound cards were by and large indistinguishable from a quality perspective as Creative Labs ones, but cheaper which Creative Labs combated mostly with lawyers as opposed to upping their game.
Some guy after being frustrated with a long-standing bug in drivers for their Creative Labs sound card, dug into the binaries and made a fix for it. During which they also discovered that you could simply flip a switch in the driver to unlock features only meant to be available on more expensive hardware. Creative Labs of course went straight to lawyers to shut them down.
My brother bought the Creative Labs WoW headset which would have its mic get progressively softer until he would leave and re-join the call/voice chat room. They never released a driver update to fix this.
By the time that Microsoft announced no more "hardware acceleration" for sound cards, I had zero sympathy for Creative Labs, I was already convinced that they made pretty shoddy hardware/software and were mostly riding on their reputation from the 90s and some patents they managed to get.
Windows Vista moved away from kernel mode drivers where possible to help address the issue of BSODs which were not uncommon on Windows XP. It turns out that Microsoft was incorrectly getting the blame for BSODs when it was actually the fault of buggy video and sound card drivers.
While Vista was regarded as a "bad" Windows, it laid most of the foundations which allowed Windows 7 to be regarded as "really good" (by Windows standards).
From a hardware perspective, by the time Windows 7 arrived pretty much all drivers had been updated for Vista and had their kinks worked out (mostly, I would still have my NVidia drivers crash on occasion, but because they were user mode now, instead of a BSOD the screen would go black for a few seconds after which my desktop would come back with a Windows pop-up saying something to the effect of "the video drivers crashed and had to be restarted", the real perpetrator now being blamed!). Windows 7 also did optimizations to make things less resource intensive and it also helped that PCs had more RAM compared to when Vista came out.
On the UAC front Windows 7 was also much better, they calibrated the UAC prompts to come up less often, but what also happened is a lot of the 3rd party software which was needlessly requiring admin rights for no good reason (except that it was badly written), had finally been fixed by the time Windows 7 was released.
It wasn't really the cool reverb effects or wave tables, though those were a nice bonus. It was just "I can tell my computer to make sound and it actually makes sound without days of troubleshooting."
Granted, similar things could be said about 3dfx. It's was a 3D card with drivers that actually worked.
And then there's the obvious "sound blasters and voodoos go in my computer, jalapeno goes in someone else's computer" thing.
In fairness on board (depending on the board but on the whole) is pretty good.
Apple is in the software first, the hardware second. Everyone at Apple has been trained to understand this for decades, and Jobs pointed it out endlessly. Apple's real moat is software (services, iOS, experience, MacOS).
Windows, Office, Azure, et al. Microsoft accumulated approximately one zillion dollars in profit on the back of software. It's a vastly superior business to anything hardware has traditionally seen. Nvidia is the first true juggernaut hardware profit machine, and the AI boom in extended hardware (RAM, storage) will prove temporary (even if there is a feast during that time). Microsoft's advantage and moat was Windows-Office for decades. It was a far better business than Intel's chip biz.
Google is a software company first. Every aspect of what made them and maintains them is software first, hardware second. They're a $400 billion software company. Their ad machine is software. Search is software.
Facebook is software. Instagram is software. WhatsApp is software. A $200 billion software company. They're not selling hardware, they're selling ads via software, they're monetizing users that use their software.
AWS is at least half software as an entity in terms of complexity, competitive advantage, et al. That's a two trillion dollar business.
LLMs can run successfully with various hardware approaches. The software is the value at the end of this, regardless of the hardware under it. The sole exception so far that may be sustainable is Nvidia, and we'll see if the bottom falls out from under that margin monster (China, specialized AI chips, whatever it happens to be that cuts under them massively).
Hardware always gets its margin squeezed eventually because it's a manufactured good (with inventory, fabs, etc). Software is hyper margin by default, you have to layer a lot of garbage on top of it to kill the margin. Nvidia is 33 years old, they have had a rich business for three years, that's it.
The AI boom is the sole reason anything in hardware has looked great in the past 20 years. Check the margins & op income for the top 20 hardware companies, from TI to AMD to Intel to Nvidia to Micron to Sandisk to Samsung to TSMC to ASML, prior to the AI boom of the past couple years. It won't last indefinitely. And after the return to a more normal environment happens, the hyper margins in software will persist.
To me, the efficiency gains of inference chips are so significant that they are certainly here to stay — barring a revolution of sorts that leads to a world devoid of AI as we know it.
This couldn't have been easy. The team at OpenAI has worked a miracle.
Take that, Jalapeno!
This is about the same rate you get out of Sol Ultraspeed.
Why do you think extra tool calls like that would be so unthinkable? It'd run circles around this, especially if the problem can be split up among a live-collaborating agent swarm, so that it's not a single user thing anymore, which is exactly what they have in the cooker with Astra.
The real question is how expensive it is to coordinate between these different modalities, and I really don't see why it'd be all that much.
I half expect Boston Dynamics to show something like this off in Q4 or whatever.
I don’t see how tokens can’t produce speech or track metabolic needs. You can talk to chatgpt can’t you? Or do you mean literally talking? Because that’s not a brain function, that’s the mouth, vocal chords, and lungs.
It probably could, but the point is this would require additional tokens, blowing up the comparison. The token output of LLMs and "token output" of speech are simply at different abstraction levels. Hence my comparison to the LLM brain driving the robot harness to produce speech etc. This would be more comparable, and also look significantly worse than "only" the 22x less efficient number.
Based on a human output rate of 3.3 tok/s, which seems questionable as a means of comparison
Also brain produces quality tokens @ 3.3 tps instead of fast generating hallucinated tokens by certain models. Thus MTP can produce low quality tokens at 2x speed.
Patience pays.
But the true number is IMO far bigger: orders of magnitude greater if we think in terms of equivalent performance.
I am relatively certain we have already squarely been beaten in net efficiency at scale.
Productivity is not the only reason to let these meatbags burn oxygen.
Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.
And the cycle continues.
https://fortune.com/2026/03/16/peter-thiel-giving-pledge-bil...
Edit: OK, hn is removing one *
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
Yeah, those guys aren't biased at all.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
lol. lmao even.
Have you seen the quality of their output? I'd take Claude or ChatGPT Free Tier over advice from McKinsey these days.
I mean, previously you could have said something much the same except substitute "frat boys".
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.
(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
And you can say of course, it's so obvious, how could a dumdum not see that! But then there are lots of examples of things where increased efficiency results in less usage overall, because demand is inelastic, etc. Jevon's paradox doesn't apply to everything.
I don't think we know yet what is going to happen as software development gets much cheaper. If in ten years we can produce software 1000x more cost effectively, will we need fewer software engineers, the same, or more? Guess we'll see!
Adding onto it, I feel as if this relates to some points regarding predictions of future in general. It is easier for us to look from the future to the past and think that it must be very obvious (as you also mention) but its also very counter-intuitive at the same time and there are just so so much nuance about basically any situation within it that its hard to really capture it all, and even then, be prepared for surprises and counter-intuitiveness.
I really like the Peter Drucker quote about it.
“The only thing we know about the future is that it will surprise us.” — Peter Drucker
and, “The future is fundamentally different from the past.” — Frank Knight, Risk, Uncertainty and Profit (1921)
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
how much pollution do you believe modern gas-turbine engines to produce?
>Not to mention the water use controversy.
what percentage of US water usage do you believe is by AI data centers?
you really don't get it?
- gpus
- retail computers
- laptops
- ~gpu~ appliances like washing machines
- cloud computing
i think you don't get how economy usually works in tech
So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.
what makes you think this?
ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.
General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.
New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
Picking the best open source models is really the best they can do.
If you make an inference only chip, you better be damn sure that it's significantly better than Nvidia's GPUs at it.
Otherwise, it's better to buy Nvidia' GPUs because they're more flexible. You can do a big training run, then use them for inference right after.
I was early at efabless.com well now chipfoundry.io - they've done about 800 chip tape outs.
They've been doing open source silicon tape outs for a decade plus.
Founder recently built this: https://nativechips.ai --- not involved but I'm inclined to believe it's the future of where the market is going. I'm skeptical of many of the AI chip design startups and whether they've actually taped out chips and how many and at what scale.
If the chips weren't this compelling they would have something different to announce.
These are paperclip maximizers who just happen to wear human skin - there is no underlying premise nor ideological goal.
Or at least Nvidia GPUs will become slightly cheaper for regular consumers again
There is probably a ~50% chance that the next Dem candidate for presidency runs on a national datacenter moratorium or something equally as crippling.
The current datacenter owners love a compute-bound world anyhow. A moratorium on new datacenters would increase their valuation, encourage efficiency and make computers cheap again. If Chinese labs can ship frontier models under 1T parameters, why not American labs too?
It's fine if you're one of the people selling shovels to gold miners for a while, but sucks to be building houses in the boom town?
In terms of GPUs whole world with 8B people have only couple of viable options: Nvidia, AMD, Intel - and largest part of their inventory is going to enterprises to run those LLMs, and its impacting every consumer / hobby projects, like cheap phones, DIY electronics projects and so on.
I want to have more alternatives on the market
> However, as previously mentioned, Jalapeño’s results are obtained without speculative decoding and Vera Rubin’s results use speculative decoding. Speculative decoding leads to a ~3-5x reduction in cost per token. When speculative decoding is implemented on Jalapeño, this will enable Jalapeño to serve tokens even more cost effectively.
How much speculative decoding improves throughput is workload-dependent. Yes, it can improve performance by 5x, but it can also slow down performance by 2x. It also requires a vastly more complex decode loop than the standard token-by-token decode. The tech journos didn't ask themselves if speculative decoding improves perf so much why wasn't it on by default? Perhaps, because speculative decoding is not optimal for OpenAI's ASIC? Perhaps that is also why they were only able to benchmark the not-very-representative single-user-decode case?
> it states that OpenAI choose not to disaggregate prefill and decode
They disaggregate INSIDE the chip, not by having separate machines for the 3 phases. the slides:
https://x.com/beffjezos/status/2092416851737518190
Maybe the money will still flow into this industry after all
I went down a rabbit hole after watching Dylan Patel on Dwarkesh today: https://www.youtube.com/watch?v=aV26V1UvkJw
I was initially just surprised by how bullish Dylan is on OpenAI/Anthropic and how bearish he is on China, despite Chinese labs getting closer to US SOTA while offering inference at dramatically lower prices.
So, I started digging while waiting for various day-job inference calls to return, ha.
Dylan says he spent years obsessively posting on hardware forums, moderating hardware subreddits, and running anonymous hardware blogs/videos before SemiAnalysis. But he also says most of that history is now gone, including from the Internet Archive, because he asked for it to be removed.[1]
In a 2024 interview he described his post-college job as “data science” around hurricane/earthquake/wildfire simulations for a financial company.[1] In a 2026 Sequoia interview he described himself as having been a “quant at a small quant risk firm” who generated $10M+ of “risk-free revenue.”[2] The Information reports that he declined to identify the employer and doesn’t list it on LinkedIn.[3]
Even harmless/silly stuff seems to drift. In February he said he kept bees for ~1.5 years. Today it was “few months, few months.”[4][5] I know, sort of silly and doesn't matter.
The Information reports that Patel owns stakes in ~20 startups in the same ecosystem SemiAnalysis covers, organized a $50M Fluidstack SPV, and is now targeting a $400M venture fund.[3][6]
And, in a 2022 HN discussion about SemiAnalysis disclosures, after saying his reports had moved smaller stocks by 20% in a day, Patel wrote: “If I thought I could move the stock, I'd make the position in the morning alongside my clients, and publish shortly after.”[7]
I don’t know that any of this is false or that anything improper happened (I’m definitely not claiming that). More that 1-2 of these things would just be odd. Taken together, though, they made me question how much trust I was putting in the broader story.
The dynamic of reminds me of crypto, WeWork, Theranos, Citron, etc. Once enough important people validate someone, things that would normally invite basic diligence somehow stop getting questioned.
[1] https://www.dwarkesh.com/p/dylan-jon
[2] https://sequoiacap.com/podcast/dylan-patel-of-semianalysis-w...
[3] https://www.theinformation.com/articles/dylan-patel-semianal...
[4] https://www.latent.space/p/dylanpatel-cooking
[5] https://www.dwarkesh.com/p/dylan-patel-3
[6] https://www.theinformation.com/briefings/exclusive-semianaly...
[7] https://news.ycombinator.com/item?id=31065646
Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
With competition we will actually have the fair split, whatever that is, and thus much lower prices.
At the moment, to have a big AI firm, or really AI firm at all, you need to be blessed by NVIDIA, in the form of receiving circular financing for your compute. They know that their prices aren't fair, or competitive.
Commoditization of inference is the end of that. The end of the mega-premium on inference hardware, and it's good not only for people who like running their LLMs, but it's the first step towards commoditization of training.
If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.
I’m not an expert in metallurgy by any means, but this seems really off. There are many recipes for steel and varied processes that also impact the final product.
They are all converging.
Each of those categories then has different grades of quality.
Tokens are a lot more like steel than oil, especially since a lot of tokens are used as structural material in the form of code.
They also fawn over the chip’s TDP when all other chips have to support 16 bit floating point and thus must run much hotter.
They make the classic mistake of equating max TDP with in-use-watts, and praise this magnificent (fictitious) performance per watt at FP8 with other chips’ max-TDP at FP16, which draw twice the power.
Evidence that the IPO can’t be far away.
Yeah, stopped reading there, this is obviously some deranged sam altman paid blog post, I can't wait for the bubble to pop just so his newly launched chip falls flat on his face.
The other side of the conversation is the idea that large models in DCs on custom silicon is the future. Maybe for enterprise? But consumers will eventually (10 yrs) have affordable hardware designed to run crazy-good local models (more RAM + higher bandwidth). That will take pressure off of datacenters, but also reduce AI profits, and move that money to consumer chip/device makers. Apple is once again the biggest winner. Nvidia consumer chips might get cheaper, but nerfed, to encourage datacenter use where they make more money. I'm hoping AMD can stop being terrible at software so that when we finally have their better hardware we can actually use it.
Example of such a system being used specifically for datacenters: https://blog.vantage-dc.com/2026/04/22/cooling-without-the-d...
Evaporated water is condensed, and in the process transfers its heat into another place that removes it. Another simple example is a pot of boiling water with a lid on it.
I'm not a datacenter engineer, but I used to work in the ski industry. Snowmaking systems use vast quantities of compressed air. It works better if that air is cool. Blowing hot compressed air out of a snow cannon means the air temperature (wet bulb to be specific) needs to be colder to make snow.
Anyways, most air compression stations use water to cool the air, and then evaporative coolers to cool the water. The water is reused, but a ton (not sure of the percentage) is lost into the air. It's more or less a tower with a big fan on top, and water percolates down from the top, being cooled by the air as it goes. The water is then collected and pumped through the system again (but of course has to be always topped up to counteract what was lost to evaporation).
Anyways, long story short is it's most cost effective to just spray water into the air to cool water, as long as water is free/cheap.
- If they use either evaporative cooling or a liquid-cooled heat exchanger, that uses tons of water consistently. This requires less energy (it's mostly passive) so you use more water.
- If they use closed-loop water cooling and/or heat pumps/electric chillers, that uses much less water - at the DC. But it does require more energy to circulate the water, run fans, etc. If you are using more energy, where is the energy coming from? It's coming from power plants, which require... you guessed it... more water (e.g. thermoelectric, hydroelectric, geothermal, concentrated solar). They need water in order to generate the power, and lots of it. Coal, natural gas, nuclear, and concentrated solar, all use steam to generate energy. Nuclear also uses water to cool the reactor. And water is used extensively to extract coal, oil, and natural gas. Geothermal uses water in the ground.
You can't not use a ton of water in one fashion or another. It just depends what method, and on what end the water is used. And the crazy thing is, most new datacenters are being built in places with extremely little water. Guess how that's gonna work out as the planet gets hotter?
I don't know why I got downvoted to hell for stating facts every datacenter architect knows. HN be HN'in.
One objective of the project might be simply to provide credible negotiating leverage when dealing with existing suppliers like NVidia. You don't have to deploy at scale for that to work, but you do have to look like you could if pushed hard enough.
But then that means you have no actual moat against the behemot, right? Your competitor can move into the market as soon as they want to, at much better cost (so at slightly better price)... and Nvidia certainly can adapt much faster around hard hardware specs innovation than a new entrant ever could.
One answer is they're quite good at poaching talent.
They should GTFO with this crap.
Create barriers to computing for ordinary people while milking businesses for tokens.
There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.
To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.
Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.
If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.
I remember when Tesla was building its own inference chips, and after about 2 years and billions spent, the whole effort was scuttled b/c they simply could not keep up with the iteration and R&D cycles of dedicated chip companies. I suspect the same will be the case with OpenAI vs Cerebras + Nvidia/Groq
Hes got a press release.
The issue is, baking something to silicon requires discipline and about 2 years.
This isn't something you can just change your mind on halfway through. Trust me, I know. You need a clear vision of what you want to support, why and what bits of a chip you need to achieve that.
Man, if only someone made like, chips that could lots of different calculations all at the same time!
Even in nvidia land rubin + LPU does a similar thing.
It has its downsides of course - if your traffic swings prefill heavy to decode heavy, you can't suddenly use your lpu for prefill. With GPUs they're totally interchangeable. Tradeoffs.
There is also the other idea where you run your attention layer on the GPU/TPU/Trainium and the FFN on the SRAM accelerator. Because KV cache is more difficult on cerebras etc, while MOE latency is easier to deal with
That sounds quite like...nonsense?
Chip companies work on years-long cycles. They know today what are they launching 4-5 years from now.