The AI Hosting Inflection Point

I am a heavy API user. The reason is arithmetic, not ideology.

Most of my current AI work is public, bursty, and sensitive to how quickly the model landscape is improving. I use APIs because they let me access the capability I need without buying a large hardware system that may be overtaken by a cheaper, more efficient option within 12 months.

I have run local models. I understand the appeal of owning your own intelligence, keeping the data on the machine, and having a system that does not depend on somebody else’s pricing or access rules. I love local AI. For my current workload, though, the economics of renting capability are better.

Privacy is not a footnote here. I am very privacy-conscious, and I would not send sensitive client information, private documents, personal records, or regulated data to an external API just because it was convenient. Local AI is the right answer when the data needs to stay local.

I have run local models. I know how to set them up, what the hardware requires, and where the trade-offs are. I am more than capable of running them. I have chosen not to prioritise owning the hardware right now because the economics do not justify it.

That does not make local AI a bad idea. I still think having local capability is worthwhile. It gives you privacy, control, resilience, and an option that does not depend on somebody else’s pricing or access rules. I just do not think it is automatically the cheapest way to access the best current models during this economic inflection period.

This is a narrower argument about general use: public web content, SEO research, open-source code, public documents, writing, and short bursts of agent activity. For that kind of work, the economics did not justify a five grand box.

So this is a narrower argument about infrastructure economics. Every workload does not need the same setup, and I am not prepared to accept the current rorting of technology prices as normal.

The break-even point has moved

An AU$5,000 local AI box sounds like a one-off purchase. It is not. I am using AU$5,000 as a rough all-in estimate for a 3090-based system, not claiming every build costs exactly that amount.

Spread across four years, the hardware costs:

AU$5,000 / 48 months = AU$104.17 per month

That is the clean version. It ignores electricity, cooling, storage, repairs, the value of the space it occupies, and the fact that the machine will age while API models keep changing underneath you.

Now compare that with your API bill.

If your current usage costs AU$50 a month, four years of API spend is AU$2,400. You are still AU$2,600 below the hardware purchase before the local box has paid for itself.

At AU$100 a month, four years of API spend is AU$4,800. That is close to the hardware figure, but still below it before electricity and maintenance.

At AU$300 a month, the conversation changes. You are spending AU$14,400 across four years. A local machine starts to look much more reasonable, especially if the workload is steady and the hardware can keep up.

Here’s the problem: most people compare a hardware purchase with one API invoice. That is not the comparison. The comparison is total cost over the useful life of the hardware.

Resale changes the calculation, but it does not erase it. If the AU$5,000 machine is worth a third of its purchase price after four years, the net hardware cost is still roughly AU$3,333, or AU$69.44 per month before electricity and maintenance.

APIs are dropping faster than hardware costs

The price of the machine is mostly fixed. A GPU bought today still costs what it costs today. It may become cheaper later, but you have already paid for it.

API pricing is different. Models improve, competition increases, providers optimise their infrastructure, and the cost per token often comes down over time. That does not make every API cheap. The result depends on the model, provider, workload, and how often you use it.

That creates an awkward situation for anyone making a hardware purchase based on today’s workload. You are locking in a large cost while the alternative is becoming cheaper and more capable.

The box does not get a software update that halves its purchase price.

An API can become cheaper without you replacing a GPU.

Hardware prices, meanwhile, are doing what hardware prices do. Exhibit A:

A used RTX 3090 still costs AU$1,500

The hardware side of the calculation is not theoretical either.

At the time of writing, I found a Melbourne listing for a used GALAX Nvidia GeForce RTX 3090 SG. It was listed at AU$1,500, 17 hours old, with the seller describing it as used in good condition. Bought new in 2021. A little dusty around the fans. Original box and accessories included.

That is a graphics card bought in 2021, being sold second-hand, for AU$1,500.

The 3090 is still a useful card for local AI. That is exactly why the price is interesting. It is not obsolete junk. It is capable hardware with a large amount of VRAM, and the market still knows what that is worth.

But once you add one to the rest of a machine, the five grand estimate arrives quickly. GPU, CPU, motherboard, memory, power supply, storage, case, cooling, and the electricity to keep the thing running.

The API comparison is no longer competing with a cheap spare computer. It is competing with a serious capital purchase.

The GPU is not the computer

This is the part that gets lost in local AI pricing discussions.

The AU$1,500 RTX 3090 is only the graphics card. You still need a system powerful enough to run it: CPU, motherboard, memory, power supply, storage, case, cooling, and an operating system. Then you pay for the electricity every time you use it, whether the workload is earning money or you are testing a new model because you went too far down the rabbit hole.

A GPU price is not a system price.

And the 3090 gives you 24 GB of VRAM. That is useful, but it is still 24 GB.

Compare that with the GMKtec EVO-X2 AI Mini PC AMD Ryzen AI Max+ 395. The product page currently shows the 128 GB RAM + 2 TB SSD configuration at US$3,649.99, with an AU power plug selected. That is an entire computer, not a GPU waiting for the rest of the computer to be built.

The comparison is not perfectly like-for-like. The mini PC’s 128 GB is system memory in an integrated design, not 128 GB of dedicated GPU VRAM. How much can be allocated to graphics or AI depends on the system and software.

But that distinction does not make the comparison less useful. It makes it more honest. The GMKtec is not a cheap alternative to the 3090. It is a serious all-in-one machine, and at US$3,649.99 before any Australian conversion, GST, shipping, or import costs, it is not a small purchase either.

It still does not change my conclusion. For my public and bursty workload, even a sensible current machine is a capital purchase that sits idle between bursts. The question is not whether the mini PC is good. The question is whether I need to own that capability often enough for it to earn its keep.

The local box has a capital cost before it has processed a single token. It has an electricity cost every time it runs. It has a physical footprint. It has a replacement problem when the next generation arrives.

This is the inflection point: the break-even calculation has moved. A workload that used to justify owning hardware may now be cheaper to rent, particularly when it is intermittent rather than constant.

That hardware bill sits inside a bigger infrastructure problem. If the industry is spending extraordinary amounts of money building AI capacity, then passing those costs through as if every user needs a $5,000 local machine is not a neutral technical decision. It is how hype turns into household hardware bills. I wrote more about that in Could All These AI Data Centers Be Worthless? Let’s explore..

This is not local AI versus APIs

The argument gets worse when it turns into a culture war.

Local AI is not pointless. APIs are not automatically better. They solve different problems.

Local makes sense when you need:

  • Sensitive data to stay on your own machines
  • A controlled or regulated environment
  • Offline access
  • Predictable latency without an external dependency
  • Heavy inference running around the clock
  • A workload large enough to keep expensive hardware busy

APIs make sense when you need:

  • Public-facing content and research
  • Short bursts of inference
  • Flexible access to the newest models
  • No hardware maintenance
  • A low upfront cost
  • The ability to scale usage up or down without buying another machine

The useful question is not whether you are ideologically pro-local or pro-API.

Both camps have very strong opinions. People get aggressive about buying a GPU and running your own intelligence, and I understand why. There is something valuable about owning the capability, keeping it under your control, and not asking a company or government for permission every time you want to use it.

But for the general user, running a local AI system is still a lot of infrastructure for a problem they may not actually have. If the work is public, occasional, and cheaper to run through an API, a five grand box is difficult to justify. That is not a failure of local AI. It is just the wrong tool for that workload.

The caveat is government access. APIs can be restricted, providers can be pressured, and a government could limit access to particular models almost overnight. That makes local capability a completely valid form of resilience. I would rather have the option than discover I had built my entire workflow around access that somebody else can switch off.

That is why I see local capability as insurance, not as a reason to buy a box tomorrow. Insurance costs money. For my current workload, I have decided the premium is not worth paying yet. If access restrictions become real, that calculation changes.

And if you are technical enough to build and maintain your own local GPU setup, you are probably technical enough to find another model, provider, or lawful route if your local or national government starts trying to block access to models developed elsewhere. The argument for owning local intelligence is real. I just do not think one local restriction automatically traps a technically capable person forever, particularly while competition from Chinese labs continues to add capable alternatives to the market.

So the question remains simple: what does this workload actually require?

Bursty workloads are expensive to own

My own workload is not a data centre.

I might spend an intense afternoon using AI for code, content, research, or planning. Then I might use very little for two days. Then a project arrives and usage spikes again.

That is a bursty workload. I pay for capacity when I need it and stop paying when I do not.

A local box charges me every day, whether it is working or sitting under the desk looking expensive.

This is the same reason most businesses do not build their own electricity grid. Owning infrastructure can make sense at sufficient scale. Below that point, the fixed cost is just part of the scenery.

Very impressive scenery, admittedly.

The maths you should actually run

Start with the hardware cost, then include the costs people conveniently leave out:

  1. Purchase price
  2. Electricity over the expected lifespan
  3. Cooling and storage
  4. Maintenance and replacement parts
  5. The cost of your time managing it
  6. The likely resale value at the end
  7. The cost of buying newer hardware when the models outgrow it

Then calculate your API spend over the same period.

Use this simple comparison:

Local monthly cost = total ownership cost / useful months

API monthly cost = average monthly API spend

If local monthly cost is higher, rent the capability.

If API monthly cost is higher and the workload is consistent, buy the capability.

If the answer changes depending on the month, use a hybrid setup. Keep sensitive or high-volume work local, and send public or occasional work to an API.

That is the sequence I followed. What data am I handling? Mostly public data. How often am I using the system? In bursts, not continuously. What do I actually need to run? The best available capability for the task, not one fixed model sitting on one machine. Once I answered those questions honestly, the decision became much less ideological.

These figures are not universal constants. They are a sensitivity test. Change the API spend, electricity rate, utilisation, hardware lifespan, or resale value and the break-even point moves with it. That is the point of doing the maths, not pretending one number applies to everyone.

The answer is allowed to be boring. Boring is usually where the margin lives.

When local wins

The five grand box becomes easier to justify when the workload is heavy, predictable, and difficult to send elsewhere.

A company processing private documents every hour has a different calculation from someone summarising public webpages twice a week.

A regulated organisation has a different calculation from an independent builder writing a blog post.

A developer running models continuously has a different calculation from someone who wants a local model because owning computers feels good.

That last reason is valid. It is just not an economic argument.

Privacy, control, reliability, and sovereignty all have real value. Put a price on them instead of pretending they are free, then include that value in the calculation. Privacy has a price for me. It is one reason local capability remains on the list, even when it is not the cheapest option today.

When APIs win

APIs currently win for public-facing or bursty workloads where the main requirements are capability, flexibility, and low fixed cost.

They also win when the model landscape is moving faster than your hardware budget. You can try a new model without replacing a GPU. You can scale up for a project and scale back when the project ends. You are paying for use rather than potential use.

That matters more than people admit.

Most computers are bought for the workload we imagine having. The bill arrives for the workload we actually have.

I will probably own a local rig eventually

I will probably own local AI capability eventually. I like the idea too much not to. But the reason I am not buying a box today is not simply that my workload is too light.

It is that the box may age faster than the workload grows.

If I spend several thousand dollars on a 128 GB local AI system now, I could be buying into a hardware target that looks very different within 12 months. A model that currently needs a serious desktop may become smaller, faster, and cheap enough to run on a much less expensive machine. It may even reach the point where something with DeepSeek V4 Flash-class capability runs locally on a phone. I do not know whether that exact timeline lands. The direction is the part I cannot ignore.

That is the argument I explored in Could All These AI Data Centers Be Worthless? Let’s explore.: abstract capability can improve much faster than concrete infrastructure can depreciate. The data centre, GPU, or mini PC does not become more efficient just because the model running elsewhere does.

So I am not waiting for my workload to become large enough to justify the current box. I am waiting to see whether the next capability curve makes the current box a bad purchase.

For now, the right answer is a boring hybrid: use APIs for public content, research, code, and bursty work. Use local models when privacy, control, offline access, or sustained utilisation make the fixed cost worthwhile.

The best AI setup is not the one that wins an argument online. It is the one that does the work at the lowest sensible total cost, without locking you into hardware that ages faster than your use case.

Calculate the four-year API spend. Add the full ownership cost. Include the value of privacy. Then decide whether you are buying capability, or simply buying today’s version of it.

Full stop.

Want us to do this for you?

Get a free audit showing exactly what's costing you rankings.

Get The Teardown

Get your free site teardown.