HomeBlog › Article
Insight · August 9, 2026

The Token Tax: AI Security Is Rebuilding Splunk's Pricing Problem

Security teams already learned what volume-based pricing does to coverage. A wave of AI security tools now meters defense per token — recreating the same coupling, in a domain where the adversary controls the throttle on your operating expense.

Every security buyer over the age of thirty has a Splunk story.

It usually goes like this. The platform worked. Analysts loved it. Then the business grew, telemetry grew with it, and the invoice grew faster than either. Somewhere around year three, a meeting happened that nobody wants to admit to: a room full of security people deciding which log sources to stop collecting. Not because those sources lacked value. Because ingest cost money, and the budget was the budget.

That meeting is the real legacy of volume-based pricing. Not the invoice — the behavior it forced. Security posture became a function of the finance spreadsheet. Teams learned to shrink the aperture and hope the thing they stopped watching wasn’t the thing that mattered.

The industry swore it had learned. Vendors built entire go-to-market strategies around not being Splunk. Flat-rate pricing. Ingest-free tiers. Pipeline tools whose sole job was reducing what hit the meter. The whole world of “Data Optimization” sprouted several companies to filter, route and store data efficiently to lower cost.

And then the same industry wired its next generation of products directly into a per-token meter.

The new meter looks familiar

A large share of the AI security tools now reaching market — AI SOC analysts, agentic triage, LLM-powered detection engineering, “reasoning” layers bolted onto existing platforms — work by shipping your telemetry to a third-party frontier model and paying per token for the response.

Look at what that means structurally, not functionally. Cost scales with the volume of events analyzed. That’s the same axis Splunk billed on. The unit changed from gigabytes to tokens; the shape of the curve did not.

Except this version is harder on the vendor, because tokens aren’t a pricing decision. They’re cost of goods sold. Classic security SaaS runs at 75–85% gross margin because serving the ten-thousandth customer costs roughly nothing. A product that calls a frontier model on every alert carries real, recurring, per-transaction cost that never amortizes away. Every event analyzed is money out the door — at 2 a.m., during an incident, forever.

Run the vendor’s math

Suppose a mid-sized enterprise generates a few hundred thousand security-relevant events a day worth reasoning over. Suppose each one costs a couple of cents in inference once you account for context, retrieved evidence, chain-of-thought, and the follow-up calls that agentic workflows make on their own. Take the arithmetic out to a year and you land somewhere in the seven figures — for one customer, before the vendor has paid an engineer.

Adjust the assumptions however you like. The exact number isn’t the point. The point is that the number is large, it’s variable, and it lands on the vendor’s income statement.

Which leaves a venture-backed startup exactly three exits, all of them bad for you:

  • Eat it. Sell at a healthy-looking price and absorb the inference cost. This works beautifully in the pilot phase, which is precisely why pilots feel so good. It stops working the moment usage compounds. Somewhere between Series B and Series C, a board deck contains a slide about gross margin, and the strategy changes.

  • Pass it through. Move to consumption pricing and hand the customer the meter. Now you own the variance. Your security spend becomes a forecast rather than a line item, your renewal comes with a true-up conversation, and finance starts asking why the security bill moved 40% in a quarter when nothing about the company changed.

  • Degrade quietly. This is the dangerous one. Route to smaller models. Sample events instead of analyzing all of them. Reserve the expensive reasoning for what a pre-filter deems “interesting.” Shorten context windows. None of it appears in a release note. The product still returns confident-sounding answers. Your coverage has simply become a function of somebody else’s margin target — and you have no instrument that tells you it happened.

Most vendors will do all three, in that order.

Then the attacker gets a vote

Here’s the part that has no Splunk analogue, and it deserves more attention than it’s getting.

When your defenses cost money per event analyzed, generating events becomes an attack. Not a distraction — an attack with a direct financial payload. Spray enough traffic, fire enough prompt-injection attempts, trigger enough agent loops, and you’re not just creating noise for analysts to wade through. You’re spending the defender’s budget on their behalf.

Denial-of-wallet is a recognized problem in serverless architecture. Nobody expected it to become a property of the security stack itself. But that’s the natural consequence of putting a metered third-party model in the enforcement path: the adversary controls the throttle on your operating expense, and the only defenses available are throttling, sampling, or capping — each of which is a decision to stop looking at things during exactly the period when things are happening.

A security control whose cost curve is set by the attacker is not a control. It’s an exposure with a dashboard.

What actually went wrong with Splunk

It wasn’t the price. Splunk was, for a long stretch, worth what it cost. What went wrong was the coupling — the fact that the cost of defending scaled linearly with the amount of activity worth defending. That coupling meant growth was punished, incidents were expensive to investigate at the exact moment investigation mattered, and the rational move was always to see less.

Token-metered security recreates that coupling in a domain where AI-driven adversaries are about to make volume cheap for the attacker and expensive for everyone else. That asymmetry is the whole problem. The economics run backwards.

Questions worth asking before you sign

If you’re evaluating anything marketed as AI-powered security, the pricing page will not tell you what you need to know. The architecture will. Ask:

  • What is your inference cost per event, and who absorbs it when my volume goes up 10x? If the answer is vague, the answer is “you, eventually.”
  • Which third-party model providers sit in the enforcement path? You are inheriting their pricing changes, their rate limits, their deprecation schedules, and their outages.
  • Does analysis quality vary with load or with my tier? Ask directly whether the system samples, and what it does when it hits a budget ceiling.
  • Can an attacker increase my bill? Watch the reaction to this one. It’s frequently the first time anyone has asked.
  • What’s p99 latency when the model provider is degraded — and does the product fail open or fail closed?
  • Where does my telemetry go to be analyzed? Metered inference usually means your data leaves your boundary. That’s a compliance conversation, not just a cost one.
  • What does renewal look like at 3x today’s usage? Model the bad case now, while you still have leverage.

The architectural answer

The way out isn’t to reject AI in security. It’s to stop renting intelligence by the call in the one place where volume, urgency, and adversarial pressure all peak at the same moment.

This is why Salience Cyber built the CognitionAI Engine the way we did. Purpose-built inspection at the point of interaction, in the browser and on the system. No kernel agent on the host. No third-party generative model in the path — which means no per-token meter running underneath your security posture, no data leaving your environment to be reasoned about, and no scenario in which an adversary spends your budget by generating traffic.

The prevention-first posture compounds the same advantage. Detect-and-respond is expensive by construction: every alert becomes an investigation, and in a token-metered architecture every investigation becomes an invoice. Neutralizing a threat at runtime ends the cost, the workflow, and the exposure in the same motion.

Your cost of defending should be predictable. It should not scale with how hard you’re being attacked. And nobody on your team should ever again sit in a room deciding what to stop watching because the meter is running.

We already ran that experiment. We know how it ends.


Get in touch to talk about what predictable AI-native defense looks like in your environment.