points by motionlessveloc 19 hours ago

I used to work on a support team of a well known backend type service that had hard budget caps.

It was, unfortunately, a nightmare. There were tons of tickets and even threats of lawsuits from customers whose service got cut off hard at the worst possible time due to organic growth/going viral/big event/nobody knew about the limit/etc. Not only did they lose all the leads and revenue they would have gotten from that bump, but they also pissed off their own existing users who suddenly couldn't use the service either.

Generally speaking, it's much better to use alerts instead of hard limit. Even in the worst case (hackers pwn your credentials and mine Bitcoin or whatever) the rest of your business is unaffected and you can negotiate with the billing department at comparative leisure.

This is all assuming you have humans operating the service. If you're letting AI agents yolo infra in prod, you have a whole series of new problems.

KronisLV 11 hours ago

> Generally speaking, it's much better to use alerts instead of hard limit.

Probably for companies and opportunistic individuals with a high risk tolerance. As for me, just give me hard budget caps. It shouldn't be a choice between supporting one of the approaches, when you could let each client pick what they want.

Want an alert? Configure one.

Want a hard limit? Configure one.

(just make the configuration easy and visible)

  • edoceo 4 hours ago

    The thing I see is folk putting a limit when given the choice. Later the situation changes, they want mega scale. Then don't fiddle the settings. I think it's a documentation & process problem.

    1000 years ago, before CI/CD took over folk would have launch/go-live events, where teams would do a good checklist. There were pre-launch meetings to review, which included load checking.

walrus01 18 hours ago

That doesn't mean you couldn't have a service which by default has no hard cap, and people have to opt into it. You could even put a user interface thing where people have to type a whole sentence perfectly matching and hit OK, like "I understand that enabling a hard billing cap will shut off services if it exceeds my monthly quota". Wrap it in as much service agreement contract, TOS language as is necessary.

Heck, have it do the equivalent of send people a DocuSign equivalent PDF to sign acknowledging the risk before enabling it. Would it still stop pissed off people? Probably not. Would it help with the risk of lawsuits, very possibly.

reticulates 18 hours ago

Yes, this is an important point although it has changed with A.I. Software is traditionally very high margin and so a $10k bill can be written off by the provider without any meaningful loss.

As a customer, the big number is scary and causes panic but for the provider… customers constantly fail to pay bills, providers are constantly writing off bills because it just isn’t worth the cost to chase, if a customer says “hey that usage was a mistake” it’s usually worth it to write it off to save the relationship. If you write off a big bill that wouldn’t have been paid anyway, the customer will perceive you as wonderful and benevolent and be loyal for life when they are ready to spend their money.

With tokens though the actual cost being incurred is much, much higher. If your service is just a wrapper around tokens, and a customer incurs $10k of usage that you paid OpenAI $5k for, it becomes much more difficult to write off.

Google Cloud is one of the few services that actually pursues unpaid bills even on their high margin services.

  • preommr 17 hours ago

    Google Cloud is also the scariest because of how much damage it can do and how bad their payment system can be.

    I recently loaded up on prepaid api credits for gemini and it somehow triggered some billing shenanigans in my linked accounts where it said I had a negative balance (from the credits), and they were going to discontinue my services. I had to reset some settings to sort it out, mainly using their chat ai and mine (because theirs gave me right status info, but wrong conclusions).

    It's pretty messy across like aistudio.google.com, and their typical console, and google workspace business account. I'd be so fucked if they froze my account, I'd rather just pay openrouter to access credits in the future.

    • mschuster91 13 hours ago

      > Google Cloud is also the scariest because of how much damage it can do

      Google is the only cloud platform I'll not just never use, but personally discourage anyone from using it.

      Simply because there have been way too many horror stories on here about people who had gotten their personal gmail accounts frozen for whatever BS reason - and absolutely zero recourse. With any other large service you can always get ahold of a human, with anything tied to Google it's impossible and even raising a major stink on HN or "legacy media" often does not help.

  • sandworm101 11 hours ago

    That will change as data centers become more "compute hubs" than service providers. When a good chunk of your bill is the cost of the electricity the compute center will have little margin to negotiate. Power companies do not care about customer mistakes.

mybrowsercache 9 hours ago

My mobile phone contract reduces my speed a lot when the hard limit it reached. I still get some connectivity. What Never happens is that I have to pay extra. If a company can’t give me hard limits on the cost I have to pay, I will not do business with them. Notifications when close to the limit are fine, but when I don’t react to them I want things to crash rather than having to pay indefinite amounts.

  • collingreen 7 hours ago

    Isn't that a soft limit? I guess it's a hard limit from the "zero additional cost" perspective but sift from the "stuff still works but in a degraded state" perspective.

    • oblio 7 hours ago

      AWS and co have basically 0 built in limits, either hard or soft.

    • bluGill 4 hours ago

      It's indirectly a hard limit. You can calculate at the reduced speed that you're given exactly how much data you can consume after you reach that limit. And there's your true hard limit.

seviu 4 hours ago

As somebody developing hobby apps, there are things i just can’t use. Take an example cloudflare workers. I recently worked on a poc for an app. I ended up hosting it at home due to the uncertainty.

The caps were very generous but I wasn’t going to spend a minute worrying about what if a malicious actor took over

hypfer 11 hours ago

I'm of course lacking specifics, but it sounds like this take-away might not be the best solution?

A better approach to that scenario might be to split into base load and peak load infra, similar to how we do it with the power grid.

Base load could be an actual metal server that you pay x amount of money for and is fully yours, with peak load being handled by autoscaling dynamic stuff.

Both things being on a fixed budget, if that budget would run out, the service would not degrade as in "grind to a halt" but as in "slows down", which is probably not a downtime as part of a communicated SLA.

My point being that cloud and on-demand is useful, but a hybrid approach might in many cases make more sense. Though of course YMMV. I do not know what the requirements of your product are.

  • kccqzy 9 hours ago

    Depending on the size of the peak, overloading your base load could very well grind it to a halt. At times of overload, you need loadshedding to recover, which is exactly the opposite of what you are proposing here.

    • hypfer 9 hours ago

      I guess that depends a lot on the specific scenario and use-case.

      • kccqzy 6 hours ago

        What use case do you have in mind? In my experience every use case I can think of won’t work with your model. The diurnal variation between nighttime demand and daytime demand (in a given timezone) is too great.

        • hypfer 6 hours ago

          I'm kinda not interested in this bad faith debate, but I'd like to call out that I perceive it as such.

mcintyre1994 10 hours ago

I can understand how that causes issues, but surely it also shows that alerts don't work either? These companies were presumably missing/ignoring alerts with scary messaging that they're about to get shut off, and would also miss/ignore alerts with scary messaging that their bill is getting too high.

  • pixl97 9 hours ago

    Ya, especially on the open internet where the time between the first alert and bankruptcy could be minutes.

layer8 6 hours ago

Why not both? Have a soft cap with alerts, and a higher hard cap as a failsafe.

aenis 18 hours ago

I think it's not generally 'much better' to use alerts. People - end users - are by now quite used to seeing things go down for a while. No biggie. But a infra oopsie can kill a company in ways a short outage won't.

And its of course not just people yolo'ing with AI. People were quite capable of causing such outages themselves just fine. Distributed, serverless systems are hard.

  • traceroute66 12 hours ago

    > I think it's not generally 'much better' to use alerts.

    Anyone who says its 'much better' to use alerts instead of hard caps needs to Google 'alert fatigue'.

    Alerts are soft. You ignore them or miss them, nothing happens except you spending $$$$$$$$$$ more.

    Great if you're the cloud provider raking in the cash, but a poor way to run your infrastructure.

    Hard caps force you to implement correctly (to control costs in the first place) and have correct monitoring in place (to keep a healthy cap buffer).

    So it means you can't just vibecode some slop and blindly devops it via Github CI/CD. You actually need to think and reason about your infrastructure.

  • bluGill 4 hours ago

    Alerts are only good if somebody responds to them each and every single time. If you don't have someone with authority who can respond to every alert and will do that as part of their job, then they're worthless. Part of responding of course is understanding what it means.

Shorel 10 hours ago

While you have a valid argument for why a customer would prefer the alerts...

It should be the choice of the customer.

Imposing one option: only hard budget cap or only alerts is the wrong design decision.

  • pixl97 9 hours ago

    Really it depends on the average loss for the service provider by customers that can't pay. If overages don't get paid by customers than all service providers will move to hard caps to prevent loss. If all providers do it, customers won't have much of a choice, except for higher caps with credit checks first.

kurthr 6 hours ago

It may be a dumb observation, but until I actually read the phrase "AI agents yolo infra in prod", I hadn't internalized that each agent is effectively always YOLOing! Because, of course they are. It is their purpose.

Kim_Bruning 7 hours ago

OH! Your actual budget is not a scalar; it's a function, a function of downstream!

figassis 11 hours ago

Why is this mutually exclusive. Is this even an argument? Some people complain becuase the hard cap kicked in. Others want hard caps. There is absolutely nothing stopping both from being implemented and offered as a choice. But we have spent decades being told "we know what's better for you". As an engineer, I think we need to come back down from the stratosphere.

spoonyvoid7 18 hours ago

> you can negotiate with the billing department at comparative leisure.

I'm curious. How likely is the billing department to waive off a huge bill as bad debt because an inexperienced builder misconfigured their infra or was hacked?

  • reticulates 18 hours ago

    Not the OP but run a high margin service that has customers run up accidental bills often. Customers running up bills intentionally and then not paying is even more common. There is almost no situation where trying to force a customer to pay makes sense, we write off any amount without question. The goodwill is worth it every time. Most SaaS companies don’t even have the processes in place for debt collection anyway.

hashstring 1 hour ago

yolo in dev can still lead to high bills

hn_throwaway_99 3 hours ago

I think you may be confusing the issues here, especially if the "well known backend type service" you are referring to is AWS.

I got bit by some undocumented (or at least very poorly documented) hard limits in AWS when our app went viral. It was extremely difficult to just find out what these hard limits actually were. And most importantly, we never explicitly set them, they were just hidden defaults in AWS.

That's very different from having an easy to use, visual dashboard of where and what all your hard limits actually are, and make it it extremely easy to turn them off or on at a moments notice.

ndsipa_pomu 11 hours ago

I don't see how it can be a problem if a service has mandatory settings for soft and hard limits. Those who don't want to disrupt services just unset the hard limit. Those who complain after the fact can be pointed at their explicit choices and have them explained to them - the service provider can hardly be held liable for a customer choosing the wrong value when it's easy to set the right value.

This is not rocket salad to implement correctly.

  • pixl97 9 hours ago

    If AI bills are that much higher, and margins are that much lower then I'd expect upsetting hard caps to start requiring credit checks.

dvfjsdhgfv 13 hours ago

> Generally speaking, it's much better to use alerts instead of hard limit.

You can't extrapolate your own experience and generalize like that. There is a huge number of people who explicitly want hard caps, period. AWS giving in and finally offering this option after 20 years of people begging them to do that is a sign of that.

clickety_clack 19 hours ago

I mean, I understand why _you_ want it that way, but that doesn’t mean that there can’t be hard budget caps for other people. You could even have both, with alerts at a level lower than the cap. I just want some kind of control over it.