One aspect Paul Kedrosky mentioned recently is the concept of „duration mismatch“. The price per token goes down over time (either because the AI vendor reduces due to competition pressure, or because customers are now incentivized to use older cheaper models). But datacenters are financed through debt, with the assumption their revenue increases over time. Quoting him: „[AI vendors are] paying for a fixed cost with a depreciating commodity“[0].
So you have on one end the token revenue trending down, on the other end the training cost going up for the next frontier models, and you need to pay back your 10y debt.
"So you have on one end the token revenue trending down, on the other end the training cost going up for the next frontier models, and you need to pay back your 10y debt."
Not necessarily, the bond holders could simply take a massive hair cut and lose shitloads of money. On the topic of bubbles and exuberance, Jeff Bezos made the salient point that there was a massive over-invested biotech boom in the 1990s and tons of sophisticated investors ended up losing lots of money. But humanity still kept the medical advancements made by the boom. Stocks going down didn't un-research drugs, and it won't un-research new GPUs or un-build datacenters.
Drugs cost pennies to manufacture after they are researched and make their way through the approval pipeline. There are many generic drug manufacturers who can work off the existing formulas.
The more apt comparison is that LLMs won't be un-trained. Opus 4.8 now exists. Even if Anthropic somehow went bankrupt, that particular asset could, at the very least, be sold for proverbial pennies on the dollar to a "generic" inference provider.
Research does get lost over time. The whole point of the patent system is keeping that from happening; if the drug company goes bankrupt, even if they lose all their internal documentation in the process, hopefully the patents and other public paperwork provides enough information for an unrelated company -- either having acquired the patent rights, or after the patent period ends -- to reconstruct the processes with less investment then the original research.
If a bankrupt AI company maintains enough of a skeleton crew to consolidate and archive its intellectual property it could be sold off to another company, but there are also timelines where it all ends up digital dust in the wind.
> If a bankrupt AI company maintains enough of a skeleton crew to consolidate and archive its intellectual property it could be sold off to another company, but there are also timelines where it all ends up digital dust in the wind.
Only if that skeleton crew had deep deep pockets. If Anthropic closed their doors tomorrow because the market collectively saw that AI was not profitable and so open sourced everything, there wouldn't be any money to train Opus 5.0... it would then have to fall on governments to put money into the hat (which I can't see happening unless it was Europe)
Datacentres aren't the same as infrastructure or research though. All the hardware in them has a finite, useful lifespan. In 10 years time it'll be totally useless
Hardware fails, and also scales out in terms of efficacy to run it as more power efficient, modern hardware turns up. It requires constant investment to keep it useful, and cost efficient
When AI pops, we'll temporarily have some extra compute capacity that will be horrendously uneconomical to run due to the high grid load and low consumer demand, before they get shutdown. There's simply no real use for them at this scale
In order to not un-build the data centers, they at least have to make more than it costs to operate them, and also not have some attractive liquidation value (the land, maybe).
I could imagine something like “inference is done at home or in China, that’s the price to beat” and it’s not worth keeping all those GPUs cool out in Nevada.
But the parent comment was that one of the bigger costs in these data centers was the interest expense on the borrowed money. A restructuring removes or heavily reduces that amount.
The fiber laid during the dotcom bubble never paid back the investors or lenders, but it's still profitably connecting customers all these years later.
It’s true once built the data center can operate right up to a financed data center value of zero. The investors will loose money but the costs of AI will go down as they do
Those data centers are specifically for AI workloads. Let’s say everything crashes and we now have all the data centers, what do you do with them? GPU are pretty specialized hardware, without AI a data center full of outdated graphics cards isn’t really too valuable.
It’s really not obvious the infrastructure we are building for AI stuff is something that will benefit humanity over time.
Without talking about the fact that bubbles are extremely destructive. Bezos is obviously someone who came out ok from the dotcom bubble but we are talking about something that destroys a lot of value globally. That has real, direct consequences, not just investors losing some money. The US economy is currently only growing because of the AI bet
AI data centers are being already used at max capacity, aren't they? I have a hard time imagining people would suddenly use AI less than they do as of today, let alone collectively drop it altogether. So the worst case scenario is that they'd need to be auctioned off way under what they'd be worth now, but still for someone to use them for AI.
Inference is much cheaper than training a new model, so running them just for inference is a completely different thing than having to price in the fact that at the moment all of these companies need to compromise between compute for inference and compute for training new models. If no new models were to be trained, and all the compute was inference only, that would change everything when it comes to the overall compute cost of AI.
Dotcom infra buildup is a bad comparison, in that it wasn't even close to being all utilized. The infra was completely overproportional to the day to day usage.
AI data centers that exist and are operational are running at maximum capacity. That's why you see things like the tiny little data center run by xai showing up as a valuable resource to xai (on the sale side) and anthropic (buy side). It is "only" 300 megawatts and there's a 1.25 billion rent on it per month.
If all these other data centers were anywhere near coming on line, that 300mw data center would be a rounding error not a line item as it is right now.
So someone's signed contracts for way more and way larger data centers, someone's purchased billions in hardware for these not yet operational data centers. I'm wondering how depreciation's going to work on all these assets...
Anyhow, I'm not really sure what "max capacity" is here, nor am I really aware when they're going to be delivering the operational assets that are currently levered to their eyeballs and consuming 1/3rd of the memory made on the planet.
As far as inference vs training, have new gotten radically better than old models or only marginally (at the cost of 10x or more the training costs)?
I imagine the trend for AI usage will go up over the very long term (5-10yrs etc.), but short term how much usage is being propped up by employer's forcing their employees to use it? Or by user's being curious about the novelty but ultimately abandoning it if it doesn't do what they want? It'll be interesting to see what changes as tokenmaxxing disappears.
I would day that the dotcom was directionally correct but the timing was wrong. For instance you had pets.com in 1999 but in 2020 you had chewy.com. It's like you had broadcast.com in 2000 but by 2020 you had YouTube that was making more in ad revenue than the next 4 largest competitors.
> AI accelerators used in DC are not really "graphic cards" any more, you ain't running gaming on it
I think the lighter 40 series cards like L40 still have OK graphics features. But otherwise yeah, after the Ampere generation graphics features went down the drain. The A100 and A40 cards can do graphics well but it already makes no sense in terms of power-to-performance ratio.
AI GPUs have terrible graphical capabilities, if at all. They can run shaders, but they are lacking in texture units, rasterization, etc... huge bottleneck here.
These AI "GPUs" are worse for gaming than even the crappiest actual GPUs (with a G as in Graphics). Also, the display drivers won't support them, not officially at least.
Has there ever been a market for cloud gaming apart from middle class people with macbooks who casually want to play one particular game but not enough to pay for a whole PC or console?
I have a big beefy gaming PC. I still use cloud gaming from time to time. It means I don't need to juggle so many 100GB installs on my gaming handheld or cheap personal laptop, both of which can sometimes struggle to play actually demanding games. Battery life on those mobile computers are significantly better when cloud streaming a game instead of running computationally demanding games locally. It also makes the friction around trying out a game significantly lower, all I need to do is click play and the game is running instead of having to wait for it to download, play it a bit, decide I don't really like the game, and then uninstall it.
The feature being bundled in with GamePass makes it worth it. I used to VPN home and try and run games remotely, but it was honestly a bit of a pain. Just pressing a button and having the game launch is quite nice.
Big AI investor tells us that investing in AI is good. Oh, the surprise!
Does that invalidate this point? Yes. Because it makes no sense. The big money is not going to R&D but to build infrastructure that will be outdated in 5 years.
No, that's not the right read. He said bubbles and exuberance still produce lasting value for humanity, even when investors lose money.
Big money is going to build infrastructure which is fundamentally required for R&D. They aren't separate, they are the same thing. It sounds like you're complaining that Pfizer isn't investing in drug research, they are buying mass spectrometers and micron fidelity microscopes. Same thing!
> bubbles and exuberance still produce lasting value for humanity
Citation needed:
Bubbles destroy value, assign resources to wasteful endeavors. Just look at the financial bubble of 2008. So many people lost their jobs, so much money was moved from the average worker pocket to banks, people lost their homes, etc.
A focus on construction will have provided millions with homes, a focus on finance and money stole them of it.
The .com bubble wasted billions in bogus projects and scams. The web didn't became to improve until after the money run out and competition between companies took over.
> Big money is going to build infrastructure which is fundamentally required for R&D.
That is never true. Big money extracts value from society. It is the small investors, pension funds, etc. that move the economy. It is everyday's working class spending. It is educated people pursuing their passion that creates value.
Big money are parasites that steal from society. That pee-in-a-bottle Bezos is trying to convince people of the opposite is unsurprising.
Current AI datacenter/model development investment rate is roughly 1T/year. That's a lot. But the US economy is 33T/year. So the investment pays back (roughly) over ten years if, each year, the AI investments increase overall productivity by 0.6%, assuming the AI companies can capture half of the value of that productivity gain.
> „[AI vendors are] paying for a fixed cost with a depreciating commodity“
That's just a confusing way to say you don't think future models will be worth the development costs.
Because if future models are significantly better, why would the price of tokens to access those models deprecate?
Companies whose main core competency is writing code were already making up a big chunk of the economy before AI. Also, less wealthy companies were constrained in their use of software by the inability to afford the salaries of talented programmers (and ripoff practices from software consulting companies who in theory could help). Lowering the cost of building software systems ought to unblock a good amount of economic activity as the technology diffuses.
Those companies are certainly writing more code. But It isn’t clear that they are increasing their economic productivity. It could even conceivably have the opposite effect by fueling a race to the bottom.
e.g. an interesting possible canary in this coal mine is that there’s been a 200% increase in the rate of new apps appearing on Apple’s App Store, but it has not been accompanied by a 200% increase in the rate at which people are buying apps.
The AI pundits often seem to apply the logic that code output is directly proportional to revenue and/or profit, and as such it follows that an AI usage increase leads to more code which leads to more revenue.
I don't believe this aligns with the reality of any major company, unless your business is in the literal sense "selling code" your revenue and profit is tangential to the quantity of code you produce. Google is a good example of this: most of their revenue and profit comes from their ad network, which is disconnected from their development productivity and instead heavily reliant on network effects and time in market. If I was a new competitor with infinite AI funds to throw at whatever problem I choose, I can't simply capture their market by developing an exact copy of Google's ad platform. In the same way, Google can't substantially grow their ad network by coding "more" or "better", they still need more customers and consumers to interact with their network to see any increase in revenue.
So it doesn't directly follow that a productivity increase will inherently follow an AI usage increase.
Agreed. I think it’s more likely to expect that most of it is pure waste.
My impression is that most software development work is not profitable. Either the project is abandoned, or it fails, or it gets shipped but doesn’t generate positive ROI. But, like how venture capital works, the minority of projects that are successful make enough money to cover the rest.
Some portion of this is because demand for software projects in general is less than perfectly elastic. So more software does not automatically mean more software sales.
It also seems plausible that, in general, companies tend to fund the projects that are most likely to be profitable. They aren’t perfect at it, but I doubt they’re just rolling dice.
Which would imply that the new work companies can take on thanks to developer productivity gains will tend to be ones that are less likely to generate positive ROI.
meaning AI may only produce a net increase in waste, which only serves to erode profits.
Add to that that it’s been years now and we still don’t have an example of someone army-of-oneing a killer app or anything like that. It’s beginning to feel like another iteration of the amazing blockchain revolution that was always & forever just around the corner.
The unlocked economic activity won’t come from a Google competitor writing code faster. A lot of it will come from “boring” businesses who could benefit from custom software but haven’t had the means to create it themselves. In some cases they may not even know their problem can be solved by software, but some AI they are using for repetitive tasks will notice and offer to build an app for them.
I would go as far as to say writing more Code has almost no impact on their economic productivity. What drives those companies is infrastructure and networks
So far the place where I've seen "more code being written" having a postive effect, has been in paying down tech debt and reduction of overhead. We've rewritten services (bringing multiple microservices back under moduliths) and cut costs. But I'm talking about net-negative code. That's not the point you're making. I agree that puking out 20 new features likely wouldn't gain us more revenue.
If the quality of all apps remains high, but if there is an increase of low quality apps it may not necessarily be great for consumers as it becomes difficult to distinguish which are the good and bad quality apps, making it risky to purchase apps.
I am yet to see that ‘companies with great ideas which simply cannot afford those very expensive developers’. For the most, issue is not programmer costs. Mostly it’s inability to formulate the MVP which makes sense.
‘uber for my industry’ is not a sensible business strategy
Honestly, if you know guys whose bottleneck is pure software dev — please let me know, I have a good, experienced team in Eastern Europe, we can do wonders in product development. But coming up with sensible business ideas and executing on them in the real world is crazy hard and extremely rare.
Can you believe that the barrier to entry on a $20 Claude subscription is lower than emailing a guy to hire a team for a project you haven’t even thought of yet? Regular people who are using AI to assist with their daily work will find the chatbot offering to build software tools to automate the work for them.
You are wrong, sir. Their core competency is building out infrastructure and networks to support their software and user base. software is by far the least complicated thing they do.
what makes YouTube YouTube is not the video player it’s the servers that can handle petabytes of uploads a day and billions of views. YouTube software wise, is no different from the 100s of porn websites that are coded by small European teams
But what if it kills current ad-tech as we know it (paying to show ads on random sites without any way to verify that the site is legit), and the flow of ad money for legitimate goods turns back to journalism, magazines and other publications?
That would be half a trillion[1] redirected to regular people just from Google Ads.
The other day I watched a YouTube video on a work machine with no history and got 2 AI generated video ads for scam products before the video played.
An AI generated man talking about his product building journey to make a pressure washer hose that didn't need power (in the AI video it didn't even have a water supply connected!) that was going to be banned in a week because it was too powerful so buy now.
I've seen AI slop before and scam ads before but the combination of the two gave me some real tingly spider-sense that things are going to get worse and that some unethical people will make a lot of money from it so be in no hurry to stop it.
I mean, that says a lot about the kind of crisis out current economy is in. How much longer can the United States Be a world leader when it’s primary function is social media and advertising
Advertising is huge because it's backed by a ton of very real products that people go and buy. It matters because people don't automatically have awareness of things they could find useful.
And writing code is one of the most economically productive activities you can do. Why is it controversial that a technology is good at this?
That value for advertising goes negative on a marginal basis.
The first time (or few times) you saw an advert, you were informed of the product's existence. I now know the Hyundai Elantra exists and could potentially be suitable for my vehicle needs. Mission accomplished.
The next 10,000 times it's just fighting over share of a finite market. I am not expecting to buy another car for another few years, so reminding me that I can choose an Elantra instead of a Corolla at all times is just vapourising cash. In fact, there's a chance that you do something obnoxious in your ad and actively burn brand reputation.
You could argue it's a take on the "everyone uses a different 5% of the features" problem-- that advertisement is going to be within the first "informational" window for someone, but maybe there are more efficient ways to not blast it at uninterested audiences.
One other angle might be asking if we still need some markets to be competitive in the first place. You don't need ads if it's a "when you need X, you'll know where to find it" sort of product. If we nationalized the insurance industry alone, we'd probably eliminate a detectable percentage of ad volume.
The cost of power cost increase alone on industry gonna erase all gains from it.
You can't consider it in vacuum. AI takes limited resources. So far it winded up cost on near every consumer electronics that runs an OS, and it winded up cost of energy that is used by the entire industry and every single customer
It's not just the cost of datacenters, it's cost of infrastructure (that given current direction of US govt will just be paid from people's fucking taxes and bills..) and cost of other industries turning outright unprofitable "thanks" to demands of AI
The $1T number seems more promises than reality, which is closer to the $300B to $500B level. Still a big number, but between a third and a half of the value used in the popular media.
These are similar numbers to the dotcom bubble. With GDP growth and the percentage of productivity AI contributes staying the same in this scenario this requires regular gains in revenue or growth. If things just stumble, like with most datacenters going unbuilt the bubble will pop.
A few things, I think you’re missing the point here
- most tasks do not require the latest frontier models, even if they are a magnitude more intelligent (we don’t actually know if that will be the case). Current Gemini flash is cheap, fast, and pretty capable with good guidance for most tasks
- now that companies pay API costs instead of a subscription they will be setting restrictions on token use to not have their budget explode (like Uber in this submission), that’s a strong incentive to NOT use expensive models, and limit their thinking budget
- there is competitive pressure from China and others who can offer very decent performances at a fraction of the token price
- the price of tokens for the frontier models is likely to go up, but the price to access older models is what depreciates! The overall price per token is going down now that we are in a new world where companies understand that token maxing is one of the stupidest concept ever created by humankind.
If you have a good model router, you can route to older, cheaper models that run on older hardware, for simpler tasks. That helps labs extend the economic life of their hardware investments. They will likely fight it at first though as they see it as reducing ASP.
This is why I'm building role-model, a routing protocol and a router runtime: https://role-model.dev/
Relative to the current usage demand for tokens is effectively unlimited. If the price of tokens go down people will send more tokens to compensate. We are very very far away from a cost per token where people run out of things they want to send through an LLM.
You’re describing the issue. The problem isn’t that datacenters are under utilized, it’s that they are used to generate something that has a value going down over time, while being financed through debt with the assumption their revenue grow over time. That’s where you have a mismatch.
To simplify let’s assume a given datacenter can generate a fixed number of token per second (in reality that depends on the model tokenizer, and other model specific factors). And let’s say it is used at full capacity, split between inference and training. If people are mostly using the frontier models the datacenter might be profitable, with enough margin to cover training of future frontier model, operating costs, and repay the debt used for construction.
However if the token price goes down and the demand doesn’t change you now risk to run the datacenter at loss, with risk to not be able to repay your debt.
If price reduce and demands increase, you can reduce the part dedicated to training, but eventually run the datacenter at full capacity to generate tokens that are becoming less and less valuable over time
For better performance of ~equivalent tasks. That's what all the harness tooling people are using does: (often) increasing output quality by significantly increasing token counts.
Right. Which means tokens are actually being priced well under cost once you factor in all this datacenter/GPU capex. Also worth noting the datacenters are not purely for training. They're for inference too.
Local privacy respecting inference can be worth it. I use a local model to log everything I do all week to automate my timesheet. I also have it do a bunch of other data tasks. I won't say that larger SOTA models wouldn't do these tasks better than a local model but PII is a concern and my employer wouldn't approve of me just setting tokens on fire everyday to do what I could do myself.
Not at all! My company has 100s of clients and we track time in 6 minute increments. I feed in my browser history, terminal logs, session scripts, calendar, git commits, etc etc into it and voila it produces a highly accurate timesheet in no time flat.
Automating it has been way better for me than the alternative of breaking my flow whenever I'm switching tasks to chart my time, or logging all my hours for the week in one sitting. Different strokes for different folks I suppose.
> My company has 100s of clients and we track time in 6 minute increments.
That’s an insane level of detail and I can understand why it makes sense to use an LLM to remove that busywork, otherwise you’d be spending half of your time on this bureaucracy.
For the company I would imagine it could be more cost efficient to stop being so onerous instead of burning tokens.
Chips age and fail with age. You can check hot-carrier injection, bias-temperature instability and electromigration as they are the main aging mechanisms. All if these are a linear function of time but exponentieal of temperature. 90-100C these chips are running at are really tough, so they are likely to fail at couple of percent to 10% range in 2-3 years depending on the margins they have in the design.
The solder joints are notorious to fail at a high rate too.
Depends, the SMD caps spread across the board the tiny ones do start to fail and go out of spec over time. they are a right pain to replace and hard to spot one that has gone out of spec to cause the chip to start crashing.
Can you not just move the epxensive part (the gpu itself) to a new carrier board in that situation? Also isn't most of the cost of the GPU itself the design of the board, not actually making one, esp if you can move the heat sinks around?
You can and this is absolutely done for GPUs. It's often more feasible to jump to the next gen GPU at this point while the old part goes into the refurbished market. I believe China buys a lit of parts like these. You will never know how much lifetime is left in them though, as there's no history of the chip.
BGA Reflow rework is not rocket science, How do you think the PCBA gets assembled in the first place? Its much easier if you dont care about the boards at all and with the huge die sizes on these accelerator chips its worth it to do a board swap
There are data centers that use and rent out 10 year old server GPUs.
They can't run larger modern models. They can't run smaller models as fast as newer servers. So their remaining market is applications where customers are okay with older, smaller models and slower performance.
They have to price the service lower than competitors due to the lower performance. The older GPUs are less efficient so it costs them more to keep them running. They're paid off, but they're taking up valuable power, space, and cooling in a data center.
Eventually there is a tipping point where it's better to replace that space and power budget with something new that has more demand.
The parts are sold off on the open market. There's an equilibrium demand for the parts from other data centers keeping older servers running and from hobby people who are okay with a jet engine sounding toaster of a GPU running in their home.
except for you know the enterprise customers who won't change their code and will pay to run old inefficent hardware just to keep from dealing with upgrades?
I'd agree. but also that's too scary. and the bottleneck is the massive manual change control process since there's no automation around any of this. :)
Why take risk when you can spend money and take no risk
As long as the demand for GPUs keeps increasing, there are more data centers being built to house them.
When you have waitlists for many many months for Blackwell GPUs, keeping the old ones around as long as customers are willing to pay for them is great.
If I as a customer have a use case for a machine learning model I developed awhile ago, so an insect identification model, I had an ML researcher/eng develop it back in 2019, and it runs fine on a 2018-era T4 GPU (NVidia 2080 era), why mess with it?
What do you think are running on the T4 GPUs in AWS? A lot of the use cases I know of for them are mid-level computer vision models that don't need to be frontier level.
I can no longer edit this, but want to expand on my comment.
I've seen those vision researchers want to train on H100s at the time and being told know, wait for the T4s.
I've seen T4s running BERT models for document classification.
When there are enough Blackwells in data centers that H100s are useless for inference by your standards (I don't know if we've arrived there or not yet), there will be people who, say, want to run the Taco Bell ordering chatbot on them. There will be people who have applications that are just fine with Qwen 2.5 who will be happy renting them.
There seems to be this crazy consensus that hyperscalers are going to go into their datacenters and throw away their old GPUs. The reality is they have a ton of paying customers for them.
And there may be insect identification apps from 2019 that say "you know what? H100s have gotten cheap enough I can use a VLLM so the user can describe where they saw the insect too", or the McDonald's website support chatbot developers say "Hey, the bigger cheapers have gotten cheap enough we can upgrade our models to Qwen 2.5".
The frontier level GPUs in e.g. AWS have a huge premium. When the newer generations come out, they will be able to cut prices to a bit of a premium over the operational costs and still make a profit, and there are a ton of down-market customers who will be interested, who aren't willing to try to outbid Anthropic for Blackwells.
In addition to the physical depreciations other comments mentioned I'd also mention that old chips will settle into a low price and then actually go up on a per unit basis if you're trying to buy a significant amount of them. With a limitation on fabrication facilities continuing to pump out older cards is an opportunity cost to the manufacturers that would prefer to be producing newer cards. If you were in a place where you suddenly wanted to buy 10,000 3080s, as an example, I'm not certain if the market could actually fulfill that demand and no one with the ability to increase the available supply to meet that demand actually wants to do so.
Chips do wear out and need to be replaced (entropy do be like that and durability is not a primary concern for chip design) so you'll need to refresh your stock and, even if you don't need cutting edge models, the price of all chips at scale will go up over time. It may feel unintuitive since, when the PS3 was released PS1s were extremely cheap - but if you're struggling to understand this effect from your experiences in the consumer market you're actually looking at the price factor that starts making antiques increase in value since at a certain point they become scarce goods. The market price for an NES is higher today than it was in 2003 because the price had already bottomed out from demand from the general consumer market but the demand remaining (speedrunners and the like) is now fixed or growing while the supply is inevitably shrinking.
They do degrade physically, but the bigger thing is they stop being competitive quickly. Each year or so we see doubling of GPU speeds for the same amount of power.
If you build a 100MW data center with GPU compute and three years laster a new data center opens with the same cost for GPUs and same electricity cost you do, but can do twice as much compute, you quickly lose business unless the market is just so constrained customers can't afford to be picky. But the moment there's slack in the market you'll see major migrations off of providers that have the same cost but half, or quarter of the same performance.
So when you see someone talking about GPUs fully deprecating in value in 1-3 years this is what they're talking about. Right now it's not a big deal because there's no slack in the market. But once there is, the bottom will drop out.
Gradually, and especially when hot. Modern chips are pretty close to the physical limits of how small they can be made, and that means atomic/chemical effects like electromigration are accounted for and determine the lifetime. Every extra 10 degrees Celsius of temperature doubles the speed of chemical reactions.
When they stray too close to the line ... you get Intel's 13/14th gen chips that wear out after 1-2 years instead of 10-20 years. Intel calls it "Vmin drift" because that doesn't sound scary, but the actual point is that various wear-out mechanisms push the chip outside of its design envelope - increasing the voltage or lowering the clock speed may get it to run for a while longer, but you're living on borrowed time as the various circuits just stop working right and you get unpredictable instruction mis-execution: https://fgiesen.wordpress.com/2025/05/21/oodle-2-9-14-and-in...
sounds like planned depreciation on Intel's part, they definitely do not design server grade chips for longevity since that would harm their own revenues
It was not planned depreciation, as many chips were failing even before 2 years and this impacted not only PC Builders and Gamers, but also some server infra providers too.
This was simply poor design, it took Intel ages to really figure out what went wrong and "resolve" it.
I used to work in datacenters, during spinning disk era we had technicians from vendors basically every couple of days to replace some broken part. When the massive switch to ssd happened instead of having them every couple of days it was 3 or 4 times per month.
Despite no moving parts things broke anyway and, even if it doesn't break, the vendor can make you change the technology just by playing with maintenance cost of the older one, limiting or removing spare parts from the market.
My understanding is that a lot of AI data centers are still heavily relying on spinning HDDs, which is why seagate, western digital are selling more HDDs than ever before.
Spinning drives are still the "best" for data density and if the IO is sequential (which wouldn't surprise me with AI training workloads), the performance delta may not be that bad vs SSDs. As always, it depends on use case.
I know that a lot of cloud storage has tiered models, where the "expensive, but faster" tiers are SSDs, but then the slower cheaper tiers are HDDs, and the "cold storage" can be HDDs that are turned off all the way to tiers like AWS's S3 "deep archive glacier" tier being tape drives.
Today's data center GPUs are essentially overclocked, and so at limit of how much the chip materials can physically handle, and therefore degrade over time. For example, GH200s operate at 1W/superchip but the actual safe power is somewhere around 650W which will allow them to function for a decade or more. But that leads to around 15% slowdown and that is unacceptable in today's competition. So current GPUs are destined to be depreciating assets.
In future, we might have fixed cost GPUs but not today.
I would presume the reason they are overclocked is because they are trying to make up for the shortage. In time, the shortage of computing components will be remedied, and tokens produced at lower power pulls will be cheaper.
I assumed the issue was similar to crypto mining, where given finite amounts of space and power it makes sense to always be running the latest and most powerful GPUs instead of keeping older hardware running. There's definitely a secondary market for these GPUs as well.
Chips do deteriorate and fail naturally at datacenter scale or in timescales of decades, though not exactly like on financial reports. Leak current increases or electro-migrations occur at junctions or whatever those words mean.
And yeah, it does feel like GPUs will start losing values slower going forward with Moore's Law being dead for a while. It used to be that 3-5 years old GPUs were more useful as space heaters than GPUs, but that's much less of the case today.
Nothing is stopping them, it's just not worth it: Have a look at e.g. vast.ai's pricing (https://vast.ai/pricing).
The V100 (2017 -> 9 years old) can be rented from $0.02 to $0.37/h (right now I can find a V100 with a Xeon Gold 6140 and 48GB RAM for $0.165/h). Let's assume the guy you rent it to pins it at its 250W TDP and let's ignore the running costs of CPU/RAM/etc...
Then you draw 1/4 kwh for that compute hour. The industrial electricity prices in the US vary between 7.5 and 25 ct per kwh (depending on state, time of day, etc...), so at 100% efficiency, assuming nothing ever breaks, and the CPU consumes 0W you earn about 14ct/h.
And remember: V100s hours are sometimes sold at 1/10th the price.
If I pick average conditions you need to start thinking of whether it is worth it to rent them out: Usually it isn't unless you have them anyways and just sell idle capacity.
It's barely worth it to run them in a pure "is it profitable" sense, if we also account for the opportunity cost of taking up a slot in your datacenter it seizes to be worth it really quickly.
> There are no moving parts, I dont think memory chips or GPU chips deteriorate naturally
I believe they do, but I too would love to know more details because there are several ways this can happen. Electromigration, package failures, VRAM failures, dielectric breakdown... Hopefully there will be studies soon similar to that old Google paper on HDD failures!
Yes, even if the hardware is untouched. As technology advances, the power cost per compute cycle goes down. A gpu using old tech costs progressively more to operate compared to the newer models. So its value goes down over time = depreciation.
As for duty cycles, the chips are perfectly happy at 100% operation. Cooling and power componants fail, not the chips. But it costs manpower to repair such things and manpower is inconveniant these days. A gpu with any sort of fault just gets dumped.
When everything is said and done it'll be datacenters in American competing with ones in China that have several times lower electricity prices. Token prices will drop to a level that will be unprofitable for American data centers and they will need to close.
the hardware itself is still useful, but random failures happen every so often, so if you're trying to run a fixed sized fleet then your fleet shrinks when you can't get spares any more
When it was profitable to mine crypto with GPUs people used to sell these miner GPUs on the used market after about two years.
These were about half of the cost of an used GPU just used for gaming. By that pricr, I'd say a GPU kept busy has twice as high a chance of failure after two years of use.
So you have on one end the token revenue trending down, on the other end the training cost going up for the next frontier models, and you need to pay back your 10y debt.
0: https://youtu.be/wGZboZcSGDY?is=64GuKyqBh_4aSjTE