Cisco on Sept. 15 expanded Splunk’s artificial intelligence stack with tools that track agent behavior, token consumption and the infrastructure supporting AI workloads. The announcements at Splunk’s .conf26 event in Denver turn a growing executive concern—unpredictable AI operating costs—into something technology and finance teams can monitor alongside applications, networks and security.
The central addition is Tokenomics, a capability inside Splunk Agent Observability that attributes token spending across AI agents and employee use of coding tools such as Claude Code, Codex and Cursor. Cisco said the system can also forecast consumption before a billing cycle closes. Separately, Cisco AI POD for Splunk brings Splunk AI to self-managed, private-cloud and air-gapped environments using Cisco infrastructure, Nvidia accelerated computing and a Kubernetes-based architecture.
Network World independently reported that the products are aimed at companies moving agents into production, where teams need to understand behavior, cost and infrastructure impact. SDxCentral also described the release as an effort to make AI spending observable rather than leaving it buried in provider invoices.
The new control layer is about cost and conduct
Enterprise AI budgets are becoming harder to read as usage spreads beyond centralized data science teams. A developer may use a coding assistant, a customer-service group may deploy an agent and an operations team may run automated investigations, each drawing on different models and pricing structures. Traditional cloud-cost tools can show a bill, but not necessarily which agent, workflow or user created the expense—or whether the output justified it.
According to Cisco’s announcement, Tokenomics is designed to connect that usage to teams and business outcomes. Splunk Agent Observability also evaluates model and agent behavior across the AI stack and can apply runtime guardrails intended to stop inaccurate, unsafe or sensitive-data-leaking actions.
That combines two questions that companies often manage separately: What is this agent allowed to do, and what does it cost when it does it? The pairing matters because an agent that completes a task correctly may still be economically inefficient, while a cheaper workflow can create larger downstream costs if it produces errors or requires extensive human review.
On-premises AI becomes part of observability
Cisco AI POD for Splunk addresses a different barrier. Financial institutions, governments, health systems and other regulated organizations may need machine data to remain inside controlled environments. The new configuration lets those customers run Splunk AI on premises, in private clouds or in air-gapped systems instead of sending sensitive operational data to a public cloud service.
Cisco said the AI POD is available now. Splunk AI Assistant is also available, while Agent Launchpad, a tool for building custom agents, is expected later in 2026. Supported self-hosted models include Cisco’s Deep Time Series Model, Google Gemma 4 and OpenAI’s GPT-OSS 20B, with Nvidia Nemotron models planned for the coming months.
The company also introduced Observability Studio, which brings instrumentation earlier into application development, and a Network Intelligence App that connects Cisco network topology, device health and events with Splunk data. Together, the products position observability as the operating layer for AI systems rather than a troubleshooting tool applied after deployment.
Analysis: Tokens are becoming a unit of operations
The strategic signal is larger than any individual Splunk feature. Tokens are becoming an operational unit similar to cloud compute, storage or API calls. Once AI moves into recurring workflows, leaders need allocation rules, forecasts, exception alerts and a way to compare spend with useful work. Without those controls, experimentation can quietly become a permanent cost center.
That does not mean a lower token count is automatically better. More capable models may use additional tokens but reduce review time, improve accuracy or resolve a task that cheaper models cannot. The useful measurement is therefore cost per acceptable outcome, not cost per token in isolation. Splunk’s approach is notable because it places model behavior, infrastructure telemetry and spending in the same operating view.
There is also an organizational consequence. Finance teams will increasingly need data from AI observability systems, while technology teams will need policies that reflect budgets and business priorities. The resulting discipline looks less like buying software seats and more like FinOps: continuous measurement, attribution and optimization across a changing portfolio of providers and workloads.
What leaders should measure next
Companies do not need to wait for a large agent fleet to establish controls. A useful baseline should identify every production agent, its owner, approved models, data access, expected task and escalation path. Cost reporting should distinguish experimentation from recurring operations and should attribute usage at least to the application and team level.
Performance metrics should include completion quality, human-review time, failure rates and the cost of correcting errors. Security teams should monitor prompt injection, sensitive-data exposure and unexpected tool use. Procurement and finance should also test how quickly workloads can move between models when prices, performance or risk conditions change.
Cisco’s release does not resolve the broader question of how much value enterprise agents will ultimately create. It does clarify the management problem. As agents become infrastructure, their costs and conduct can no longer sit outside the systems companies use to run the rest of the business.
