Banks Are Giving AI Agents Names, Logins, and Managers. The Label Is the Easy Part.
A July 2026 report revealed banks like BNY are treating AI agents as named digital employees — with...
New VentureBeat Pulse data shows enterprises are buying AI infrastructure faster than they can measure what it costs — 83% run their GPUs cold, and fewer than half can track their compute spend. The same trap is quietly forming under small businesses. Here's how to stay out of it.
On July 16, 2026, VentureBeat published a wave of Pulse Research that put a number on something a lot of operators already felt in their gut: most companies cannot see what their AI actually costs.
Across 107 enterprises, spending on AI infrastructure is accelerating well ahead of the ability to measure or steer it. Fewer than half — 44% — can rigorously track what their compute costs. Meanwhile 83% report GPU utilization at or below 50%, which is a polite way of saying most of them are paying full price for hardware that sits idle half the time.
VentureBeat named it directly: the compute gap. The distance between how aggressively organizations are investing in AI and how little of its economics they can actually see.
That is an enterprise dataset. But the pattern underneath it is not an enterprise problem. It is forming right now under every small business that spun up an AI tool, connected it to a few systems, and started running real work through it without ever asking the boring question: what is this costing us per unit of work, and would we know if that number tripled?
Most can't answer. That's the gap. And it's cheaper to close before it's expensive.
The compute gap is not a discipline problem. It's a design problem. It shows up because the tools people adopt first are built to hide the meter, not expose it.
You sign up for a per-seat plan. You get a flat monthly number. That number feels like the cost. It isn't. The real cost is the model usage underneath it — tokens consumed per task, retries when a run fails, the long-context calls that quietly cost ten times a short one. On a flat plan, that variability is invisible to you because the vendor is absorbing it and pricing it back in with a margin.
That works until it doesn't. The moment your usage grows past what the flat rate assumed, one of two things happens. The vendor raises the price, or the vendor throttles you. Either way, the first time you learn what your AI truly costs is the moment you've already lost the standing to do anything about it.
The VentureBeat data shows how deep this runs even at the top of the market. When enterprises choose infrastructure, only 8% decide on headline token price. They choose on integration with their existing stack (41%) and total cost of ownership (35%). They have already learned that the sticker price is not the cost. The cost is everything around it — and most of them still can't measure that cost cleanly.
If the largest, best-resourced buyers in the market are flying this blind, the small business running three AI tools on three separate flat plans is not in a better position. It's in a worse one.
Here is the finding from the same report that should stop any operator cold. A clear majority of enterprises — 64% — plan to switch or add an infrastructure provider within twelve months. And 38% plan to do it within a single quarter.
Read that again. Two-thirds of organizations intend to change the foundation their AI runs on inside a year. This is a category with churn intent you'd normally see in a fickle consumer app, not in the layer businesses build their operations on.
They want to switch because the market is moving underneath them. New models ship every quarter. Specialized compute providers that barely registered a year ago are now the top thing 45% of them plan to evaluate. Prices swing. Capabilities leapfrog. The provider that's right today is a coin flip to be right in six months.
Now combine the two findings. Most businesses can't measure what their current setup costs, and most of them plan to move off it soon anyway. That is the worst possible position to negotiate from: you don't know your own numbers, and you're about to make a foundational change based on not knowing them.
We wrote recently about why model lock-in stopped being abstract after a frontier model got pulled offline worldwide with no warning. The compute gap is the financial version of the same lesson. Lock-in isn't only a risk when a vendor disappears. It's a daily tax you pay when you can't see your costs and can't move without a rebuild.
The fix is not a spreadsheet. It's an architecture where cost is a first-class, visible property of the system — and where switching providers is a configuration change, not a migration project.
A well-run AI setup answers three questions on demand, without a data project:
When cost is visible like this, the 64%-plan-to-switch problem stops being a threat. Switching is just changing a setting, because your workflows were never welded to one provider's pricing in the first place.
The bad version is the default version, and it's seductive because it's easy at the start.
You run each AI function on a different flat-rate tool. Sales on one, support on another, content on a third. Each one gives you a tidy monthly invoice and hides the usage underneath. Six months in, the invoices have crept up, you can't attribute any of it to specific work, and when you try to move a workflow to a better or cheaper model you discover it's trapped inside a product that only runs one vendor's models.
You didn't choose that trap. You defaulted into it, one convenient signup at a time. That's how every compute gap gets built — not through a bad decision, but through the absence of a decision about where cost and control should live.
The single most striking number in the VentureBeat report is that 83% of enterprises run their GPUs at 50% utilization or less. They bought dedicated capacity, and most of it sits cold.
A small business isn't buying GPU clusters. But the same waste shows up in a different costume: paying for standing capacity you don't use, and paying a premium for burst capacity you can't predict.
This is where the shape of your infrastructure matters more than the price tag. Ephemeral, session-based tools spin up fresh for every task and tear down after — which sounds efficient until you count the cost of re-establishing context, re-loading data, and re-running setup work on every single invocation. You pay that tax constantly and it never shows up as a line item you can see.
Persistent infrastructure runs the other way. An always-on agent server holds context, memory, and connections between tasks, so work doesn't restart from zero every time. The cost is steady and legible instead of spiky and hidden. We've made the full case for why persistent servers beat ephemeral containers for real business work — the cost argument is one piece of it, and it's the piece the compute gap makes urgent.
The lesson from the idle-GPU number isn't "buy less." It's "buy infrastructure whose cost you can actually see and match to real work." Utilization you can't measure is money you can't manage.
You don't need a finance team or an infrastructure audit to get out ahead of this. You need a handful of concrete moves. Here's the order we'd run them in.
Inventory every place AI is costing you money. List every AI tool, plan, and API your business pays for. Most operators are surprised by the length of this list — that surprise is the compute gap in miniature.
Find the ones that hide the meter. For each item, ask: can I see what this costs per unit of work? If the answer is "no, it's a flat plan," flag it. Flat plans aren't automatically bad, but a flat plan you can't see behind is a cost you don't control.
Separate cost from provider. Identify which of your AI functions are welded to a single model vendor. Those are the ones that'll hurt when — not if — you want to switch. The report says 64% of organizations plan to move within a year. Assume you'll be one of them.
Consolidate onto infrastructure that bills at list rate. Instead of a patchwork of flat-rate tools each taking a margin on hidden usage, run your AI coworkers on infrastructure where you pay the model provider's actual rate and see the meter. Transparent billing isn't a nicety — it's the only way the other three steps stay true over time.
Route by stakes, not by habit. Once cost is visible, match the model to the job. Low-stakes, high-volume work goes to cheaper models. High-stakes work goes to the strongest model available. This one discipline routinely cuts real spend without cutting quality, and it's only possible when you're not locked to one provider.
Do these five things and you've closed the gap the VentureBeat data is describing — not by spending less on AI, but by finally being able to see, attribute, and steer what you spend. That's the difference between AI as a mystery line item and AI as a managed part of your operation.
Q: What is the "AI compute gap"? A: It's the distance between how much a business is spending on AI infrastructure and how much of that spend it can actually measure and control. VentureBeat's July 2026 Pulse Research found fewer than half of enterprises (44%) can rigorously track their AI compute costs, even as spending accelerates. The gap forms when investment outruns visibility.
Q: Why can't most businesses measure what their AI costs? A: Because the tools they adopt first are designed to hide the meter. Flat per-seat plans bundle variable model usage into a fixed price with a margin on top. You see a tidy monthly number, not the real per-task cost underneath it. The first time many businesses learn their true cost is when the vendor raises the price or throttles usage.
Q: Is the compute gap only an enterprise problem? A: No. The VentureBeat data is from enterprises, but the pattern is worse for small businesses, which tend to run several flat-rate AI tools in parallel with no way to attribute cost to specific work. A small business running three AI tools on three plans has three compute gaps, not zero.
Q: How does model-agnostic infrastructure help with cost? A: When your workflows aren't welded to one provider's pricing, switching models is a configuration change instead of a rebuild. You can route low-stakes work to cheaper models and high-stakes work to stronger ones, and see the cost difference immediately. Given that 64% of organizations plan to switch providers within a year, that flexibility is a direct hedge against both price hikes and lock-in.
Q: What does "billing at list rate" mean and why does it matter? A: It means you pay the model provider's actual published price for usage, rather than a marked-up rate bundled into a flat fee. List-rate billing keeps the meter you see identical to the meter that exists — which is the foundation of every other cost-control move. If the number is marked up and hidden, you can't manage it.
Q: What's the first step to closing my own compute gap? A: Inventory every place AI is costing you money — every tool, plan, and API. Most operators are surprised by how long that list is, and that surprise is the gap itself. From there, flag anything you can't see behind, and separate the functions that are locked to a single vendor from the ones that aren't.
The businesses that stay out of the compute gap aren't the ones spending the least. They're the ones who can see their spend, attribute it to real work, and change providers without a rebuild. That's an architecture decision, and it's easiest to make before the invoices pile up. If you're ready to stop using AI tools and start running a real team of AI coworkers — on infrastructure that bills at list rate and never locks you to one model — Associates AI Teammates gives you a 14-day free trial with no credit card required. Start your free trial at associatesai.team.
Written by
Founder, Associates AI
Mike is a self-taught technologist who has spent his career proving that unconventional thinking produces the most powerful solutions. He built Associates AI on the belief that every business — regardless of size — deserves AI that actually works for them: custom-built, fully managed, and getting smarter over time. When he's not building agent systems, he's finding the outside-of-the-box answer to problems that have existed for generations.
More from the blog
A July 2026 report revealed banks like BNY are treating AI agents as named digital employees — with...
A July 2026 national survey found 74% of small businesses use AI while 78% still don't trust it to h...
New July 2026 research traced 2,422 AI-generated sentences back to their sources and found that 76%...
Want to go deeper?
Get started today. Hire your first Teammate in minutes and put it to work on what you're reading about.
Get Started