When AI Agents Execute the Work, Token Cost Is Only Part of the Economics
When AI Agents Execute the Work, Token Cost Is Only Part of the Economics
Alibaba Cloud's AgentCore offers a concrete example of how enterprise AI economics can change once agents move beyond generating responses and begin executing end-to-end workflows. Token and infrastructure costs still matter. But when an agent also uses tools, retries failed steps, handles exceptions, accesses memory, and requires human review, enterprises may need a higher-level metric: the total cost and reliability of completing an acceptable unit of work.
Alibaba Cloud's AgentCore offers a concrete example of how enterprise AI economics can change once agents move beyond generating responses and begin executing end-to-end workflows. Token and infrastructure costs still matter. But when an agent also uses tools, retries failed steps, handles exceptions, accesses memory, and requires human review, enterprises may need a higher-level metric: the total cost and reliability of completing an acceptable unit of work.
Enterprises comparing AI platforms often begin with model pricing and token costs.
That makes sense when AI is primarily used for question answering, summarization, content generation, or a single API call.
Agentic workflows are different.
Once an AI system is expected to complete multi-step tasks, use tools, access memory, manage permissions, interact with other systems, and recover from failures, the economics extend beyond model usage.
Alibaba Cloud's AgentCore makes that distinction more concrete. The platform combines agent execution with model connections, tools, memory, governance, and monitoring. Its cost structure also includes compute, gateway usage, storage, data transfer, and other infrastructure required to run agent workloads.
That creates an important enterprise distinction:
The vendor's billing unit and the enterprise’s business-value unit may not be the same.
A vendor may charge for tokens and infrastructure. An enterprise deciding whether to deploy or scale an agent may need to measure something different: the total cost and reliability of completing an acceptable unit of work.
1|Agent Costs Extend Beyond Tokens
Alibaba Cloud's AgentCore provides a platform for building and governing AI agents.
Agents can connect to models, MCP tools, skills, long-term memory, and other execution resources. The platform also provides controls for permissions, governance, and monitoring.
AgentCore is now a separately billed service. Its charges include resources such as:
compute
gateway usage
storage
data transfer
other infrastructure required to run agent workloads
This does not make token economics irrelevant.
Models still consume tokens, and model usage remains part of the overall cost structure.
What changes is the number of cost layers an enterprise may need to consider once an agent begins executing a full workflow.
2|Billing Units and Business-Value Units Are Different
Consider two agent systems.
The first uses a cheaper model and consumes fewer tokens.
But it also:
fails more often
requires more retries
needs more human intervention
takes longer to complete a task
The second uses a more expensive model but:
completes more tasks on the first attempt
requires fewer retries
needs less human intervention
finishes work faster
If an enterprise compares only token prices, the first system may appear cheaper.
The conclusion may change once retry costs, human review, exception handling, failure recovery, and elapsed time are included.
That suggests a different question for enterprise buyers.
Instead of asking only:
How much does this model cost per million tokens?
How much does this model cost per million tokens?
they may also need to ask:
What does it cost to complete one acceptable unit of work?
What does it cost to complete one acceptable unit of work?
Those questions measure different things.
The first captures an underlying usage cost.
The second is closer to the economics of the business process the agent is supposed to perform.
3|Agent Pilots May Need Different KPIs
Many generative AI pilots track metrics such as:
number of users
number of prompts
token consumption
response quality
user satisfaction
Those measures remain useful when the system functions primarily as an AI assistant.
If the goal is for an agent to execute a workflow, enterprises may need another set of metrics.
Task completion rate
How many assigned tasks does the agent actually complete?
Human intervention rate
How often does a person have to take over?
Retry rate
How many attempts are required to finish an average task?
Exception rate
How often does the workflow stop because of data, permission, tool, or process failures?
Time to completion
How long does it take to produce an acceptable result?
Failure-recovery cost
What people and infrastructure are required when the agent fails?
These metrics reveal costs that token consumption alone cannot capture.
Two agent systems could each generate $1,000 in model charges. If one produces 10,000 acceptable completed tasks and the other produces 5,000, their business economics are materially different.
Token cost can remain part of the numerator without being the most useful denominator.
Token cost can remain part of the numerator without being the most useful denominator.
4|Model Benchmarks Still Matter, but They Cannot Answer the Whole Agent Question
Enterprises comparing models will continue to care about:
reasoning performance
coding capability
latency
benchmark results
context windows
price
Those factors remain relevant
But an agent's ability to complete enterprise work also depends on other components:
tool-use reliability
memory accuracy
permission handling
multi-step orchestration
failure recovery
human oversight
For an agentic workflow, a model benchmark can help answer:
How capable is the model?
How capable is the model?
It cannot, by itself, answer:
How reliably can the overall system complete the business process?
How reliably can the overall system complete the business process?
Enterprises may need both answers.
5|This Is Not Yet a Shift to Outcome-Based Pricing
There is an important evidence boundary.
Alibaba Cloud still bills AgentCore according to underlying resource usage, including compute, storage, gateway usage, and data transfer. Model usage also continues to generate token costs.
There is not enough evidence to conclude that enterprise AI has broadly shifted from token-based pricing to outcome-based pricing.
Nor is Alibaba charging customers according to completed business tasks.
The more defensible conclusion is narrower:
Even when vendors continue to bill by underlying usage, enterprises can evaluate value at a higher level than the vendor's billing model.
Even when vendors continue to bill by underlying usage, enterprises can evaluate value at a higher level than the vendor's billing model.
The vendor's invoice may be denominated in tokens, compute, and storage.
The enterprise decision can instead ask:
What is the total cost and reliability of completing one acceptable unit of work?
What is the total cost and reliability of completing one acceptable unit of work?
Today's Decision Insight
The most important change introduced by AI agents may not be how many additional tasks AI can perform.
It may be how enterprises need to measure AI economics.
When AI primarily generates an answer, token pricing, model quality, and response performance can provide much of the information needed to compare alternatives.
Once an agent begins executing an end-to-end workflow, the cost structure becomes more complex.
A single task may require model inference, tool calls, data retrieval, memory operations, system actions, retries, human review, and exception handling.
A lower-cost model therefore does not necessarily produce a lower-cost workflow.
For CIOs, CTOs, enterprise AI leaders, and operations executives, that has a direct implication for pilot design.
If the objective is for an AI agent to take responsibility for actual work, the evaluation cannot stop at:
Is the model good enough?
Is the model good enough?
or:
Are the tokens cheap enough?
Are the tokens cheap enough?
It may also need to ask:
Can the agent reliably complete the work? What does one acceptable completion actually cost? And how much human intervention is required along the way?
Can the agent reliably complete the work? What does one acceptable completion actually cost? And how much human intervention is required along the way?
Alibaba Cloud's AgentCore does not prove that agents have replaced models as the primary unit of enterprise AI purchasing. Nor does it show that outcome-based pricing has become a market standard.
What it does provide is a concrete example of why enterprise AI economics may need another layer of measurement.
As AI moves from generating outputs to completing work, the denominator used to evaluate its economics may need to change as well.
As AI moves from generating outputs to completing work, the denominator used to evaluate its economics may need to change as well.
Decision Question
When an AI agent executes an end-to-end workflow, what should an enterprise measure beyond token and infrastructure costs to determine the true cost and reliability of completed work?
When an AI agent executes an end-to-end workflow, what should an enterprise measure beyond token and infrastructure costs to determine the true cost and reliability of completed work?
Turning Global AI Signals into Business Decisions

