The AI cost crisis is a GTM ops problem. Here’s how to fix it.

How B2B SaaS companies are bringing AI spend under control without killing the workflows that actually work.

A founder we work with recently told us their internal AI bill had crossed $80,000 a month. Not customer-facing product. Internal. Marketing, sales, RevOps. Mostly Claude.


They had no visibility into who was spending what. No model defaults. No token budgets. Just a company-wide mandate to “use AI” and a credit card getting hammered at the API layer.


This is not an unusual story right now.


The AI cost crisis spreading across SaaS isn’t a technology problem. It’s an operations problem. And it was entirely predictable.

How we got here

Starting in late 2024, investors pushed founders to show AI adoption. Founders pushed their teams. Teams adopted AI tools without guardrails, governance, or any sense of cost-per-workflow. The FOMO was real and the budgets were effectively unlimited.


Now the bills are arriving. Kyle Poyar at Growth Unhinged reported in July 2026 that Anthropic’s enterprise pricing restructure removed token subsidies above 150 seats, and that Fable 5 runs at roughly 3-5x the per-task cost of Opus 4.8. Claude Code, Cowork, and third-party AI tooling compound on top of base model costs.


Poyar profiled Pylon, a B2B support platform, where engineering was spending $5,000 per person per month across AI coding tools, and the support team was triggering a Claude skill on every ticket at $2 per ticket in base inference costs alone. No one had measured any of it.


The problem wasn’t AI. The problem was that nobody built the ops layer around it.

What we’re seeing with GTM teams specifically

Sales and marketing tend to be the hardest places to justify AI spend, and also the places where spend is most undisciplined.


Sales teams use AI for meeting prep, email drafts, competitor comparisons, and list building. These are low-leverage applications of expensive models. An AE running Opus on a routine follow-up email is the equivalent of calling in a specialist to fill out a form.


Marketing spend is worse. AI-written copy is still producing mediocre output at scale. The use cases that actually work, creative brainstorming, positioning refinement, pattern analysis across large content sets, are high-value but low-volume. They don’t explain a $700-per-person monthly AI bill.


The hard question every GTM leader needs to ask: is the spend producing measurable output? If you can’t point to pipeline, conversion, or efficiency metrics that have moved because of AI, the spend is not justified yet.

The architecture problem nobody’s talking about

Here’s where most companies made their biggest mistake, and it’s the one hardest to fix retroactively.


They adopted AI reactively. They gave people access to frontier models, told them to figure it out, and called it transformation. What they built was a collection of ad hoc prompts, disconnected workflows, and model dependencies with no governance layer underneath.


This matters for costs because non-deterministic AI, models called directly, without structure, is expensive to run and unreliable at scale. Every task requires a large model. Every output requires human review. Every workflow breaks when the model updates.


The companies managing AI costs well have something in common: they built a harness around the model. A router that sends tasks to the right model at the right cost. Tools and MCP servers that pull structured context so the model isn’t doing expensive reasoning work to find basic information. A validator that checks output before it ships.


This is the architectural frame: a router, tools, MCP servers, and a validator. The model sits in the middle and does the reasoning. Everything else is engineered to be deterministic.


This is not a small build. But the economics are significant. When you route a content brief to Sonnet instead of Opus, you’re spending 10x less per task. When you pull CRM data via an MCP server instead of asking the model to figure out what it needs, you eliminate thousands of tokens per call. When you batch non-urgent tasks through the Anthropic Batch API, you cut costs 50% automatically.


The companies that built on top of frontier models without architecture are now paying the full cost of that decision.

A practical framework for getting costs under control

If you’re starting from scratch, here’s the order of operations. It maps directly to the framework we’ve applied across RevOps engagements for over a decade.


Step 1: Get visibility. You can’t manage what you can’t see. Map AI spend by team, by tool, and by use case. This alone will surface the obvious waste. Poyar reported that Pylon’s CEO had spent $4,000 in three days on a single Claude Code analysis before he knew the meter was running.


Step 2: Set defaults, not bans. Model defaults are the highest-leverage lever most companies haven’t pulled. Setting Sonnet 4.6 as the organizational default instead of Opus, and requiring approval to use heavier models, reduces costs dramatically with almost no quality loss for everyday tasks. Poyar noted that Vanta blocked Fable entirely and set Sonnet as default, with no meaningful quality difference reported.


Step 3: Assign token budgets by team. The goal is visibility and a small amount of friction before spend escalates. Poyar cited Tesla capping spend at $200 per week and Uber at $1,500 per employee per month. The right number depends on the use case. The point is that a number exists.


Step 4: Audit your workflows against the cost. For each major AI workflow, ask: what does this cost per run, what output does it produce, and is there a cheaper way to get the same result? High-frequency, low-complexity workflows, ticket summarization, email drafts, data formatting, should be running on cheaper models or automation, not frontier AI.


Step 5: Build the architecture layer. If you have AI workflows generating real value, invest in making them deterministic: a routing layer that matches task complexity to model cost, structured context retrieval so the model doesn’t burn tokens on search, and a validation step so output quality doesn’t regress when models change. This is where twelve years of building RevOps systems has taught us the most, reliable outputs come from reliable architecture, not from hoping the model gets it right.

What comes next

AI pricing is not going to get simpler. The frontier model providers are under margin pressure and are moving away from flat subscription models toward consumption-based pricing. Fable 5 is the clearest signal yet: the best models are going to cost significantly more, and the pricing advantage for light usage will erode.


The companies that will manage this well treat AI spend like any other strategic investment. Not special budget. Not FOMO spend. A line item that needs to justify itself with measurable output.


The ops layer that most companies skipped on the way in is the thing that will determine who comes out of this with AI that works and a P&L that holds.
That work is what we do. Domestique is a HubSpot Platinum Partner and fractional RevOps firm serving Series A-D B2B SaaS companies. We’ve spent twelve years building the operational frameworks that make GTM systems reliable, AI-augmented or otherwise.


If you want to stay current on how GTM ops is evolving, subscribe to our newsletter below.

Sources: Kyle Poyar, “The AI cost crisis is entirely self-inflicted (and Fable 5 just made it worse),” Growth Unhinged, July 8, 2026.

Previous
Previous

HubSpot Optimization: A Practical Playbook

Next
Next

Upcoming RevOps Masterclass: The GTM Stack Hyper-Scalers like OPenAI & Anthropic Use