It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
It's one part technical, one part a product decision. The technical part is that billing is not actually instant. As a most basic example, a VM reports its billing units every X period of time it is active. If there is some network blip but it's still running, then that billing data could be delayed.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
I think there is a middle ground between deleting data and allowing 5000 VMs to be created to mine bitcoin. Obviously there are a lot of different scenarios to consider but the explosive costs seem to be constrained mostly to a couple of features which would be fairly safe to cap.
A very charitable take, in light of tech industry habits of exorbitant rent-seeking in scenarios of Platform Dominance (e.g. Google and Apple on the app store). We should remember AWS and Google companies are among the best in the world at A/B testing and extracting revenue from cloud services.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
I had a $.20/month recurring charge from AWS that I could only remove¹ by completely deleting my AWS account. That was enough to get me to give up on AWS for personal projects.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Sounds familiar. I’m being billed £0.01/month for something in GCP, I don’t know what even after digging, but I’m too fearful to complain about it or disable the account lest it somehow gets my main Gmail account blacklisted somehow.
If these cloud providers had needed a standard customer acquisition strategy to grow to their current size, hard caps and other “training wheels” features would already be in place to get people interested in and comfortable using the platform, with the hope of eventually getting a foothold into Enterprise like most SaaS startups have to do (“enjoy our product on a side project and then recommend us to your CTO!”). But AWS and GCP got to start as in-house providers for their own constellations of massive sites and back out from that to serving other hyper scale businesses first. The lack of friendly on-ramps and starter account features is a reflection of that origin more than anything.
It's definitely technically difficult. You can't easily estimate how much an operation is going to cost before you kick off that operation, which means as soon as you get close to the limit you are at risk of tripping it.
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
Stop everything is pretty damaging any real business though. Things were better in the era of VPSs. You paid for a fixed amount of compute, if you ran a stupidly expensive operation than it just maxed out your system for a certain amount of time and things slowed down. But you didn’t kill the service entirely and you didn’t have unlimited potential price
Yes and then the choice is run it and forgive it, or, stop the process midway.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
Because most enterprise users would much rather have overages in billing than outages. The opportunity costs on any serious service I deploy dwarfs usage pricing, at least at the level a generic cloud can determine.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
It’s not a binary decision though. Any sensible enterprise has many AWS accounts. Often hundreds or thousands. It’s the only clear separation of privilege.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
It is a technical reason. Basically cloud billing is much more granular and across many more services / line items than most things that basically the pipelines that figure out how much you have spent take a long time to know how much you have consumed. I believe all cloud providers with granular usage based billing have this problem.
This is one of those features that customers think they want without having thought it through:
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Counter argument, this is the sort of thing that, especially for a smaller business or individual, can be the difference between a bad night and bankruptcy.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
The alternate conversation is "the new report run had a bug and cost us $1,000,000 over the weekend" and I think that one's usually worse.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
Big companies have thousands of budgets. An email is _worthless_. In fact, it would probably cause me to lose faith in a cloud that provided that as the control.
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
These shouldn't even exist without a negotiated contract.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Monthly electricity bills are based on usage, and it works well but there’s a limit to how surprising a bill can be. The difference is the relative orders of magnitude you can be charged for these services you can go from 20$/month to 200k/month without warning.
I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
Hard caps are rare because companies find it more profitable to forgive sympathetic individuals' bills while raking in profits from corporations whose services have gone awry
Having a monthly summary or estimate of how your spending is going would be really useful, too.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
My org has a leaderboard for AI spending each month, and I have found it interesting how fast the distribution decays, just within the top 10 users. I often think “what did these people do with all those tokens?” It’s interesting to think the answer to that question is “maybe not a lot?”
Right? If spending the most is lauded, why wouldn't I use the most expensive model, automate things that don't need automating, build things I don't need to build etc just to jack the spend up?
> In an ideal world, our agents could help with this.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
Counterpoint: if you can automate API calls on the client side, why can't you automate billing caps? If you want a machine that can run 24-7 and make money for you while you sleep (which let's face it is the motivation for a lot of AI takeup), isn't the onus on you to install cicuit-breakers?
Because many cloud services have incredibly complex or opaque pricing structures that make it difficult to impossible to determine how much something is going to cost you ahead of time, especially if it's usage-based a la network egress (and the usage statistics don't update frequently enough to make such circuit breakers possible to implement client-side).
They might not be able to predict your bill but how much time do they need to add up what you already spent to minimize your overage? And TBH how much time should be acceptable to exceed your cap before it's their fault for the lag in their software.
I would just not sign up for a service without price transparency, or pre-calculate my liability based on available information before pushing the (metaphorical) Deliver Now button.
Making incredibly complex and opaque pricing structures is not necessary for the providers to charge for and make a profit on their service. And being technically difficult is a lazy excuse. Cloud platforms have to solve many, much more difficult challenges to offer their services at all, they just don’t want to invest the time in more customer friendly billing because they expect it will result in reduced revenues.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
I'm not sure where you got the impression that I'm making excuses for cloud providers. I'm just stating the way things are, not the way I think they should be.
I had an api key set to read only that somehow ran up a $400 bill, I contacted openai about it and never heard back. Not quite the same thing, but still, I find this very annoying.
We always did, the clouds convinced us that overages were the norm. You can blame credit ratings as another vector for big business to screw everyone over. Everything should have been pay in advance with an alternate billing method for overages if you want it.
The solution is to not give agents access to MCP servers.
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
It'd be nice if more than AI spend worked this way, autoscaling is almost a mixed blessing because unpredictable pricing can be worse than the cost savings...
Ubicloud does not have hard budget caps, which I only realized this morning after moving all my CI over to them over the past few months. Fortunately I didn't learn the hard way.
I understand this is snark, but if you think about it, this is already implemented in electrical infrastructure. If I use too much power, the circuit breaker trips to protect me and protect the electrical grid. OP is about a billing breaker, but the parallels should be obvious.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
The peak throughput does result in an overall monthly limit though. For a house with a 200A main breaker, that effectively limits your electric bill to $7,000/month, which is very reasonable compared to the tens of thousands of dollars in a single day that a lot of cloud billing disasters end up costing.
Why do people think new laws are needed to solve every last problem in the world?
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
Yes and no. I suspect many of the hard limits were set arbitrarily, and we'll see a relaxation of limits as people get frustrated with the limited use they get out of them. And some services will genuinely need to be re written to support higher rps or risk losing customers
I'd support this provided we have the converse as well: if the customer doesn't pay their bill on time, the service gets shut down immediately. (Disclosure: I sell SaaS services to people who don't pay their bills on time).
One of the biggest benefits of not engaging with LLMs or any of this nonsense is you dont have to care about all these "self made" problems of the LLM-gliteratti.
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
Because it’s unlikely they’ll actually be able to collect that million dollars from a lot of those customers. Rephrased: why would your vendor want to make it harder to accidentally give you a million dollars of services in exchange for debt of dubious quality?
the premise seems a bit faulty to me. why should we be giving next token predictors access to spend our money? like what great benefit do we get from this that we should allow them unfettered access, but with safeguards in the form of hard budget caps?
I don't think Simon means you should hand off the spending to agents/LLM (which would also make me uneasy) but that if you're probing one for hosting/SaaS providers they should default to recommending ones with budget caps
This is about budget caps (“I don’t want to spend more than $100, cut me off once I spend that much”) not price caps (“no one is allowed to charge more than this price per token”).
It’s incredible that in 2026, AWS and GCP are only just now introducing this. It’s possibly one of the most obviously needed features for a cloud provider.
Also, does anyone know why it’s taken this long? I suspect it’s a technical reason. While one could be cynical, I doubt it’s an intentional business/product decision. Hard spending caps are both excellent product differentiators and could possibly save these providers money as they don’t have to forgive their users when they accidentally over-use a service.
It's one part technical, one part a product decision. The technical part is that billing is not actually instant. As a most basic example, a VM reports its billing units every X period of time it is active. If there is some network blip but it's still running, then that billing data could be delayed.
The product level decision is that "shut down everything" is something the customers you want to target don't actually want. Are we including deleting RDS data? S3? Glacier storage? If so then the headline will just change from "Hobbyist got charged XXXXXX on AWS" to "Business literally had all their data deleted because a hacker took over their VM and mined bitcoin". The only people who really want this are hobbyists and it's not a market segment that's worth chasing. Easier to do the status quo of forgive afterwards then even open the can of worms of deleting all of a business's data and all their backups just because they had a 100k overrun.
I think there is a middle ground between deleting data and allowing 5000 VMs to be created to mine bitcoin. Obviously there are a lot of different scenarios to consider but the explosive costs seem to be constrained mostly to a couple of features which would be fairly safe to cap.
A very charitable take, in light of tech industry habits of exorbitant rent-seeking in scenarios of Platform Dominance (e.g. Google and Apple on the app store). We should remember AWS and Google companies are among the best in the world at A/B testing and extracting revenue from cloud services.
When you're one of only two real options out there, you can afford to demand users put up with things that on their surface seem ridiculous. Such as a billing system that (oops!) makes it difficult for customers to see where their costs are coming from, trim their largest sources of spend, notice meaningful changes in line item prices, or limit their spend. Wild how they can figure out a million different advanced services but gosh-darn-it can't figure out the hardtech of displaying line items.
Large enterprises can afford employees who are tasked full time with unwinding this capacity to mitigate the impact of these billing headaches. But I think this measure is introduced now because LLMs introduced a risk that these billing specialists could not control without caps.
I had a $.20/month recurring charge from AWS that I could only remove¹ by completely deleting my AWS account. That was enough to get me to give up on AWS for personal projects.
⸻
1. The key word was “I.” Maybe someone more skilled at navigating AWS’s menu structure than I could would have done it quickly and easily, but even though I knew what it was for, turning it off and not getting billed for it turned out to be a huge challenge. Thankfully it was only $.20, but if I were using the service for something that generated actual bills, that $.20 (and possibly more) would end up quietly siphoning money out of my pocket into Amazon’s).
Sounds familiar. I’m being billed £0.01/month for something in GCP, I don’t know what even after digging, but I’m too fearful to complain about it or disable the account lest it somehow gets my main Gmail account blacklisted somehow.
If these cloud providers had needed a standard customer acquisition strategy to grow to their current size, hard caps and other “training wheels” features would already be in place to get people interested in and comfortable using the platform, with the hope of eventually getting a foothold into Enterprise like most SaaS startups have to do (“enjoy our product on a side project and then recommend us to your CTO!”). But AWS and GCP got to start as in-house providers for their own constellations of massive sites and back out from that to serving other hyper scale businesses first. The lack of friendly on-ramps and starter account features is a reflection of that origin more than anything.
It's definitely technically difficult. You can't easily estimate how much an operation is going to cost before you kick off that operation, which means as soon as you get close to the limit you are at risk of tripping it.
Consider something like a "select * from bigtable" SQL query that might process a trillion rows. Hard to know that's going to cost $100 until after you have run it.
Yeah, we probably want some kind of traffic light system:
Green means go Orange means finish what you're doing but don't start anything new Red means stop everything
And probably a special rule to permit stable, critical spend through regardless, the same way we allow police and ambulance to run lights.
Stop everything is pretty damaging any real business though. Things were better in the era of VPSs. You paid for a fixed amount of compute, if you ran a stupidly expensive operation than it just maxed out your system for a certain amount of time and things slowed down. But you didn’t kill the service entirely and you didn’t have unlimited potential price
Yes and then the choice is run it and forgive it, or, stop the process midway.
If you stop then you have to decide whether to charge for uncompleted work.
Interesting tradeoffs.
For very small ops e.g. individual Lambda invocation you have similar concerns especially if lots are fired at once from a queue or schedule or fanout.
Advertising platforms have had this since their inception. They were just motivated because they could be left holding the bag.
Because most enterprise users would much rather have overages in billing than outages. The opportunity costs on any serious service I deploy dwarfs usage pricing, at least at the level a generic cloud can determine.
So it’s a feature the best customers don’t want, that adds risk to those customers deployments, to appease the worst customers.
At least historically. Perhaps Simon is right that the calculation has changed.
It’s not a binary decision though. Any sensible enterprise has many AWS accounts. Often hundreds or thousands. It’s the only clear separation of privilege.
There’s no reason for the majority of them to have an infinite budget cap. Prod? Sure. UAT? Why not. The sandbox environment Johnny just spun up to test some new agentic workflow? Hell no.
That's a reason to not force a hard budget cap on all of your customers, but it's not a reason to not offer one.
Features for bad customers that risk good customers are easy to say no to.
It is a technical reason. Basically cloud billing is much more granular and across many more services / line items than most things that basically the pipelines that figure out how much you have spent take a long time to know how much you have consumed. I believe all cloud providers with granular usage based billing have this problem.
This is one of those features that customers think they want without having thought it through:
“Never let me spend more than $X” also means, “Shut down my business-critical app/service/solution at 2 am on a Sunday morning because Joel in IT forgot to plan for the new report runs.”
The product design work to let customers have the first thing without risk of major pain from the second thing is non-trivial.
Counter argument, this is the sort of thing that, especially for a smaller business or individual, can be the difference between a bad night and bankruptcy.
Sure it sucks that critical services blinked out at 2am. But what sucks even more is finding out the image on my ASG had a vulnerability that allowed someone to install a bunch of bitcoin miners which kept me fully scaled from midnight to 2am.
Or more likely, that a mistake in terraform 1000xed my spending.
Most people have predictable spending and could easily say "don't spend more than 10x what I normally spend". Or 1.5x, or 2x, 3x, etc. All depending on how they want to balance a runaway cloud expense.
The alternate conversation is "the new report run had a bug and cost us $1,000,000 over the weekend" and I think that one's usually worse.
But one could just have two categories of service - the default capped plan, and a special Enterprise one where you sign a contract making it clear you understand the consequences of not having a budget limit.
Also, if your average usage is $900, set your hard limit at $2,000, not $1,000. Then when the report runs $500 over expected, you get a soft limit email and still have your report. Even a "business critical" run is probably not actually worth more than double your average spend.
Even if you set your cap at $10,000 it would be better than nothing.
The price cap should be the number you’d be willing to spend to avoid an outage vs when you’d rather kill everything and work out what happened.
I don't really understand that argument. This seems pretty obvious to me, as a customer. Is this really something that companies don't understand?
Sending an email when your budget gets low shouldn't be a big lift.
Big companies have thousands of budgets. An email is _worthless_. In fact, it would probably cause me to lose faith in a cloud that provided that as the control.
Wait, Google Cloud finally added hard caps on spending per service? I've been wanting that for so many years! They sure took their sweet time.
https://cloud.google.com/blog/topics/cost-management/new-ear...
Edit: Ugh it's fake. Literally only works for four random services, unsupported for all the rest. Completely useless for all of my projects. Also dumb that the only supported term is "monthly" considering that months are different lengths, and that they don't bother to account for credits or discounts. Maybe next decade they'll get around to implementing something useful.
The hard caps work on projects created in AI studio
What that's awesome? Yeah they definitely waited until the competitors did it first...
It was fucking on purpose, if we had a functioning government, this is one of things they would have nailed them on.
As awful a practice as it is, I'd _much_ rather have an internet where its reform is prompted by competition than by the cops.
These shouldn't even exist without a negotiated contract.
I can subscribe to your service for a specific fee on a monthly basis ($20/month say), and take the risk of losing that month's fee if your service or I make mistakes, or I can choose to drop another $20 mid-month, or anything for convenience.
Saying that the computer will "control" the billing and can run haywire tells me that I don't want to be anywhere near your pile of bad incentives.
Monthly electricity bills are based on usage, and it works well but there’s a limit to how surprising a bill can be. The difference is the relative orders of magnitude you can be charged for these services you can go from 20$/month to 200k/month without warning.
There is also a hard physical limit on how much electricity you can use before you blow out the fuse box
Also, people don't casually swing by my house and start using my electricity.
So my powerbills are predictable.
Whereas traffic spikes to websites are not.
This age of abusive AI crawlers and the non-revenue generating traffic has been a very real problem for me!
> but there’s a limit to how surprising a bill can be
I think they explicitly said that.
So, prepaid services which you top up?
I know AWS and similar sites have no such concepts, but other than those, this is an existing option of most kinds of service providers.
Realistically, most providers can't even allow you to consume $10k usage if they don't have the certainty that you can pay up. Prepaid is what gives them that certainty.
Hard caps are rare because companies find it more profitable to forgive sympathetic individuals' bills while raking in profits from corporations whose services have gone awry
Clearest explanation ever.
Plus, nobody wants to be the fired PM who said "I spent our eng. hours to achieve -20% revenue".
I wonder if BigCorp adding spending caps is due to them getting sick of customers solving it for themselves with virtual cards.
Virtual cards don't solve anything. They just get you sent an invoice instead.
It works for some non-B2Bs.
Poor old Goon Spittoon at 42 MacCroon St keeps getting my invoices.
Having a monthly summary or estimate of how your spending is going would be really useful, too.
Even if we have a negotiated yearly contract for $X spend per year, maybe we’ll hit that spend in 6 months instead of 12. Having some kind of automated telemetry saying how we’re trending would be so useful.
I’ve gotten vague warnings from customer success people saying vaguely that, but without any warning of what’s truly happening. The more we abstract away from money (tokens, credits, etc), the more we need a way to translate right back to money, to see how close we’re getting to any limits over time.
I don’t want to find out in month 5 that the contract which was expected to cover a year is now going to run out in 14 days.
Where I live, you can just not pay for something and it is cancelled.
Phone service, bank card, home internet etc.
If you don't pay your bill than they just cancel your membership and it works ok.
People in western countries are just getting shafted by companies for (mostly) no reason because an alternative balance is just inconceivable.
The west has long since let people who get off on usury hold too much sway.
My org has a leaderboard for AI spending each month, and I have found it interesting how fast the distribution decays, just within the top 10 users. I often think “what did these people do with all those tokens?” It’s interesting to think the answer to that question is “maybe not a lot?”
the answer is almost definitely "get on the leaderboard"
Right? If spending the most is lauded, why wouldn't I use the most expensive model, automate things that don't need automating, build things I don't need to build etc just to jack the spend up?
It should be illegal to not have them
Probably should also have spend controls for gambling too but seems like we're a long ways off from good legislation there.
Agreed. Or at least, customers should only be liable for expenses they incur up to the hard caps they set.
If you don’t have a mechanism for enforcing hard caps, you don’t get to send customers a bill for unlimited amounts.
Everyone in this thread should be vibe coding legislation with lean 4. We can at least have utopia for a couple months.
we should make it illegal to be unhappy too, that way we can solve depression!
Anytime someone says “there sight to be a law that…” there almost always shouldn’t be.
Haha what
> In an ideal world, our agents could help with this.
Not exactly the sort of case Simon has in mind, but I tell Claude to keep to hard daily limits on its OpenRouter spending for two long-running projects [1, 2].
A Routine for each project fires ever few hours, and Claude decides itself what to do in each session. It does tasks that require calls to other models through OpenRouter only when it is still within its daily budget for that project; after it reaches that cap, it does other tasks that don’t require extra spending.
[1] https://github.com/tkgally/je-dict-1
[2] https://github.com/tkgally/eex-dict
So weird, cause it seems a lot of services are suddenly adding them. Huh, wonder what changed?
Counterpoint: if you can automate API calls on the client side, why can't you automate billing caps? If you want a machine that can run 24-7 and make money for you while you sleep (which let's face it is the motivation for a lot of AI takeup), isn't the onus on you to install cicuit-breakers?
Because many cloud services have incredibly complex or opaque pricing structures that make it difficult to impossible to determine how much something is going to cost you ahead of time, especially if it's usage-based a la network egress (and the usage statistics don't update frequently enough to make such circuit breakers possible to implement client-side).
They might not be able to predict your bill but how much time do they need to add up what you already spent to minimize your overage? And TBH how much time should be acceptable to exceed your cap before it's their fault for the lag in their software.
I would just not sign up for a service without price transparency, or pre-calculate my liability based on available information before pushing the (metaphorical) Deliver Now button.
Sometimes the invisible hand of the market needs to be paired with a good swift legislative kick in the ass to speed things up a little.
Making incredibly complex and opaque pricing structures is not necessary for the providers to charge for and make a profit on their service. And being technically difficult is a lazy excuse. Cloud platforms have to solve many, much more difficult challenges to offer their services at all, they just don’t want to invest the time in more customer friendly billing because they expect it will result in reduced revenues.
I know your post isn’t explicitly defending the platforms, but the arguments they use feel transparently flimsy.
Stop making excuses for the cloud providers. They build these on purpose, for that purpose, working as intended.
I'm not sure where you got the impression that I'm making excuses for cloud providers. I'm just stating the way things are, not the way I think they should be.
Last time I checked it was impossible to get relevant info from API, for example for AWS.
Or if there was info that was after potentially horrible expensive operation.
That was blocking automated checking.
I had an api key set to read only that somehow ran up a $400 bill, I contacted openai about it and never heard back. Not quite the same thing, but still, I find this very annoying.
"Oh hai. I'm hooked on the drugs. Please stop me from taking more. kthnxbai."
Seriously. You all asked for this.
Can you explain your train of thought here?
It seems like you’re saying “hey, you asked for a product, so you deserve for it to have a user-hostile feature”
Like hey, you asked for trains? Well, then you have no right to complain about any aspect of a train.
Who asked for what, specifically?
I can’t think of anyone saying they would hate for AWS to support hard spending caps.
That's going to be a hard sell to service providers who rely on people basically ignoring overspend.
We always did, the clouds convinced us that overages were the norm. You can blame credit ratings as another vector for big business to screw everyone over. Everything should have been pay in advance with an alternate billing method for overages if you want it.
This could have been written in 2006
The solution is to not give agents access to MCP servers.
The entire MCP ecosystem is ludicrous. You’re paying for inference for an agent to make the same decisions over and over again, and yet the actions they’re taking can be so easily written by those same agents into a bash script you can run again and again, deterministically and for free.
MCP is the problem. Having agents “use a product on your behalf” is the problem.
The pattern you’re looking for is that agents should write scripts that use products for us, and that doesn’t need a new protocol.
It'd be nice if more than AI spend worked this way, autoscaling is almost a mixed blessing because unpredictable pricing can be worse than the cost savings...
Ubicloud does not have hard budget caps, which I only realized this morning after moving all my CI over to them over the past few months. Fortunately I didn't learn the hard way.
I'm surprised this isn't law. Should it be?
So if you use more electricity next month the CEO of the electric company should be put in jail?
I understand this is snark, but if you think about it, this is already implemented in electrical infrastructure. If I use too much power, the circuit breaker trips to protect me and protect the electrical grid. OP is about a billing breaker, but the parallels should be obvious.
If the CEO of the electric company didn't mandate circuit breakers, he should go to jail.
The breaker in this analogy is equivalent to rate quota limits, it limits peak throughput.
You can rack up an outrageous monthly electricity bill without tripping a breaker.
The peak throughput does result in an overall monthly limit though. For a house with a 200A main breaker, that effectively limits your electric bill to $7,000/month, which is very reasonable compared to the tens of thousands of dollars in a single day that a lot of cloud billing disasters end up costing.
Analogies are not the core of your cognition.
Utilities required for basic survival are expected to remain on for obvious reasons. Like not letting people die.
Why do people think new laws are needed to solve every last problem in the world?
Google implemented caps because their competitors offered them. Before that, customers could choose one of several competitors, rent the GPUs at a fixed rate, or buy the GPUs and install them on-premises.
At no point was any law needed to solve any of this.
Yes and no. I suspect many of the hard limits were set arbitrarily, and we'll see a relaxation of limits as people get frustrated with the limited use they get out of them. And some services will genuinely need to be re written to support higher rps or risk losing customers
Cloudflare also doesn't have any hard budget caps.
I'd support this provided we have the converse as well: if the customer doesn't pay their bill on time, the service gets shut down immediately. (Disclosure: I sell SaaS services to people who don't pay their bills on time).
This is how most services work...
These already exist, it's the most standard contract imaginable.
What do people think of apps switching to a lower tier of model if your $$ threshold was exceeded?
One of the biggest benefits of not engaging with LLMs or any of this nonsense is you dont have to care about all these "self made" problems of the LLM-gliteratti.
Cheers!
You mean you want to curb business Gacha?
AWS, is a loot box... Tokens are just in game currency, and that sales person is just metrics that have identified your spending as making you a whale.
Your average CTO from the last decade turned a fixed cost into variable spending that looks like a mobile game.
That feature is called a subscription. Joking aside, it is needed on the API side
surprise $10k bill is getting off easy
This guy only ended up with a ~$6k bill from his agent, not too shabby: https://news.ycombinator.com/item?id=48500012
Why would your vendor want to make it harder for you to accidentally give them a million dollars?
Ideally because I'll pick a different vendor who protects me from such mistakes.
Because it’s unlikely they’ll actually be able to collect that million dollars from a lot of those customers. Rephrased: why would your vendor want to make it harder to accidentally give you a million dollars of services in exchange for debt of dubious quality?
Because I would go to the vendor who does.
But there's no such vendor… the joys of free market.
the premise seems a bit faulty to me. why should we be giving next token predictors access to spend our money? like what great benefit do we get from this that we should allow them unfettered access, but with safeguards in the form of hard budget caps?
I don't think Simon means you should hand off the spending to agents/LLM (which would also make me uneasy) but that if you're probing one for hosting/SaaS providers they should default to recommending ones with budget caps
If you own a resource that is desired and paid for by usage then that's your income and the open market determines the money thing (price)
What on earth is Simon whittering on about?
This is about budget caps (“I don’t want to spend more than $100, cut me off once I spend that much”) not price caps (“no one is allowed to charge more than this price per token”).
A bunch of vendors provide this exact feature already, so I'm clearly not weird in wanting it.
About desire to have hard caps, to avoid runaway costs for example due to misconfiguration.
This may happen also without vibecoding.
Article seemed clear on that?
bro this is HN, the "open market" where willing customers can choose whether or not to buy a service is the root of all evil here.