Enterprise AI Is Still Day One: Token Costs, Deployment Chaos, and a 10-Year Diffusion Curve
Despite three to four years of AI momentum, enterprise deployment is still remarkably nascent. Levie, drawing on conversations with hundreds of Fortune 500 CIOs, notes that most organizations only just finalized rollout plans for basic AI chat. The deeper irony is that the pace of breakthroughs is itself a drag on adoption — each new capability wave tends to make the prior architecture obsolete before it is fully implemented, creating a state of unprecedented IT inconsistency. The realistic timeline for full agentic diffusion across knowledge work is ten years, not two or three.
Token cost and budget governance has quietly become the number-one enterprise AI pain point. Unlike flat SaaS fees, a single agentic task can consume hundreds of dollars of compute, and line-of-business teams have no FinOps muscle to manage it. IT budgets (3–7% of revenue) cannot absorb productivity gains that should flow to the full business. The result: AI spend is escaping to marketing, sales, and legal budgets — departments with no tooling to measure ROI on tokens. Levie sees a potential $5B opportunity in AI compute ERP — governance, allocation, and ROI measurement for inference spend.
The coding environment is the exception that proves the rule for agentic AI. It has uniquely favorable properties — technically fluent users, verifiable outputs, self-contained codebase context, and no entitlement complexity — that simply do not exist in general knowledge work. Agents in broader enterprise contexts immediately collide with data access, permissions, semantic inconsistency, and unstructured context spread across dozens of systems. Chat worked because it only touched search and the LLM; neither had a permission problem. Agents touch everything.
On business model evolution, Levie predicts every enterprise software company that survives the AI transition will operate dual pricing: seat-based for human users plus consumption-based for purely agentic workloads. Some Box customers are already asking whether agents should hold their own seats for governance and retention purposes. Labs, meanwhile, face a structural tension — they want to move up into applied use cases to control the end-user account and avoid being commoditized, but maximizing inference volume may ultimately favor maintaining a healthy ecosystem of vertical players above them, echoing the AWS marketplace dynamic.
On jobs, Levie pushes back on doomerism: Jevons paradox means more capable engineers unlock more projects at large industrials, small businesses will staff agent-adjacent roles they never had before, and the bridge layer between raw AI capability and working enterprise workflows — covering data integration, change management, domain expertise, and ongoing maintenance — is itself a large and durable employment category. His practical advice: employees should spend 5–10% of their time and $50–100/month developing hands-on fluency with agents now, before the transformation fully arrives.
Aaron Levie (Box CEO) argues enterprise AI adoption is far earlier than the hype suggests, with rapid capability breakthroughs paradoxically slowing rollouts by obsoleting architectures mid-implementation. Token cost management has emerged as the hottest operational issue, and the full diffusion of agentic AI across knowledge work is realistically a decade-long transformation.
Enterprise AI deployment is still incredibly early even though the AI wave has been going for three to four years.
Some use cases companies will eventually realize don't need an agent at all — software running on a CPU is actually cheaper than having an agent re-render a UI every single morning at $30 per day.
It's it's just I'm smiling at the you know old problems are a new again and effectively we're talking about a semantic layer which you know I guess it is getting rebranded as an anthology and that's the new new thing when in reality it's been the same problem for 20 years plus.
it's actually why the doomers are also wrong about jobs because this is actually going to be a very real sustaining job that is not like a one-time you implement the the agent and you upgrade the system and then it kind of works forever
People who have an eye for design will become both designers and managers of agents doing design — AI does not convert copywriters into world-class designers
There are multiple tasks where hiring a person to use agents would be economically justified overnight, but executives won't personally do all the wiring and prompting themselves.
Compressing AI capability into vertical applied use cases — handling implementation, workflow integration, and abstracting complexity from knowledge workers who just need to move on with their day — is where vertical players will find their competitive opportunity.
You could probably do this diffusion in like 2 years or 3 years and we could probably all do the change management like collectively as an ecosystem.
The technology is getting so advanced that it makes obsolete the prior thing that you implemented which actually means that the rollout takes longer.
The bigger question is Silicon Valley engineering versus sort of non-engineering knowledge work.
I have a decent kind of sample size, probably a couple hundred CIOs just this year across the kind of Fortune 500 Global 2000 type of cohorts.
Everybody just started figuring out like their final rollout plans for like the chat system in the enterprise.
We still have five years of growth of like the chat system for knowledge work.
The productivity gain from chat is sort of rate limited by the human's ability to like have a conversation.
CIOs know that their engineering teams are using Cloud Code and CodeX and Cursor, and they're seeing the productivity gains come out of those teams.
The tone from enterprise CIOs is actually remarkably optimistic and excited and positive as opposed to a typical trough of disillusionment.
Microsoft canceled their internal cloud code licenses this week after token-based billing made the cost untenable.
One agent could be consuming a thousand dollars of compute on a single task, so you can't lump that all into a twenty dollar per user per month fee.
The period where everything has been flipped on its head on cost modeling and token budgeting basically all correlates to the Anthropic revenue curve.
The per token cost for frontier tokens has increased, which is completely opposite from the narrative that the cost of tokens was always going down.
The data center providers and the labs have pricing power and don't need to lower prices on anything right now.
We've compressed what should normally happen in like 10 years of rollout into like 18 months.
Enterprises are having an uncomfortable acceptance surprise about token bills as opposed to an I'm not doing this anymore surprise, because they're empirically getting the productivity gains.
IT spend is basically somewhere between like three to seven percent of corporate revenue in a company.
Employees don't actually really know what the cost of compute is, so they're going to go about using these systems as freely as possible.
One little task given to an agent could cost $200 because you just happen to structure the query wrong and now it's going to go fan out across a bunch of systems.
OpenAI has a dedicated capacity program where if you know your workload, you're able to lock in certain pricing that helps support that.
The average enterprise will certainly be using half a dozen models in their organization.
For customer service interactions, you can now cap that at 50 cents per million tokens and it will never go higher than that because you might swap it out with an OSS model.
in coding you have a highly technical user. You have models that are hyper trained on coding. You have you know effectively verified uh you know verifiable work because like the code either like runs and you can QA and you can have tests on it
That technical user, like the moment the agent does something stupid or runs into a problem, the user themselves know how to go fix it and get it back on track.
the code base has so much of the context in coding, whereas in the rest of knowledge work, the context lives across like 20 different things, some digital and some very not digital
chat was like basically it could do two things. It could it could access search and it could access the LLM. And that and that was amazing. And but guess what? Neither of those things has a permission problem. Neither of those things required wiring up some other system where you could have massive data leakage.
diffusion is going to take time
we should totally be thinking on the order of 10 years as like a as a rough type scale for like whatever it and this might be
the breakthroughs keep happening faster than the customer can implement any kind of standard architecture. And and those breakthroughs often times basically undo or make obsolete the last thing you implemented.
if you went to an enterprise right now, it's actually a period of maybe the least amount of of consistency I've ever seen in in IT
I could probably lay out up to 10 to 15 reference architectures to all solve that problem
for 20 to 30 years in IT, it was sort of okay to to sort of have all these systems, some redundant, some not well managed. You could kind of throw humans at the problem
most enterprises have five different places where their contracts are being stored, you know their road maps are across you know 30 different locations uh inside of their their data environment
do you have technical people in your organization that you can say, I'm going to have you go sit next to the business or within the business, and your job is to understand the patterns of how these people work, and make sure that they have the ability to use agents to go and and do that work
once the model changes, there's another set of work to be done. Do you have to make sure like did you get the gains of that model improvement? Or did you have to leave behind some scaffolding that you had to build for the prior model?
we built this insane technology that's like incredibly using computers, incredibly at using software, incredibly at writing code, incredibly incredibly at writing tool using tools. And but guess what? It like has, you know, a fixed amount of memory. It has a fixed amount of context it can work with.
Complex queries involving Box data, Salesforce data, and Workday will be done fully headlessly inside of coworker Codex or something similar
Building a data room and sharing contracts correctly is slower via text than via a graphical user interface
By volume, database queries headless will just be a hundred times larger than the interface-driven way of doing work
Agents are going to be banging on enterprise systems far more than humans ever did
Any enterprise software company in three years from now that gets through the AI transformation period will have both a seat business model and a consumption business model
Some Box customers are already exploring whether agents should have a Box seat because they need to store data that gets retained and governed over a long period of time
Agent seats would probably need to be cheaper than regular end user seats
Box launched a cloud for legal solutions announcement enabling an agent to read through 100 contracts or review a data room for risks
Box has had an API available almost on day one of the business
Coding has a unique property different from most knowledge work: sloppy agent-written code largely doesn't matter as long as the software runs, short of security risks or memory inefficiency
In legal work, unlike coding, you cannot have an agent write a contract whole cloth because there is no way to fully verify it — a lawyer still has to attest 100% validity
The Financial Times published an article approximately 3 weeks prior about lawyers being inundated with contracts and legal questions their clients generated via ChatGPT that lawyers now have to adjudicate
Small businesses will for the first time be able to augment functions they wouldn't have had internally before with agents, and each of those functions will often need a human to make the agent effective
Startups are hiring as fast as possible because their productivity gains from AI are causing them to need to hire for additional job functions
Box continues to hire across marketing, engineering, IT to build agents, and sales reps, with the contours of job titles not shifting as much as expected
In 10 years, going into marketing will require being a CS minor equivalent for agents — deploying a full end-to-end marketing campaign as one of the standard tasks of the role
Companies that already reached saturation of demand for a function with humans represent roughly 10% of the economy; the remaining 90% have unmet demand that agents will unlock
Society works really well when people want to work at companies and can feed their families — destroying that for one extra point of operating margin is a net negative outcome
Perplexity Computer does a better job than any other computer-based agent for being a workhorse — going through websites, doing search-related tasks that require clicking and reading pages
As an executive who starts using AI agents, you will see lots of areas where you should actually hire more people because the output creates value that needs humans to run with it.
Unless the labs build out the equivalent of hundreds or thousands of people for every single vertical and every single line of business, there is a lot of opportunity in that bridge area of work.
Labs are moving up into applied use cases and doing some well and some not well; for some announcements Levie now uses the lab directly instead of the vertical application, and for others the lab's version is still the poor man's version.
A lab does not want a vendor above it that can swap it out at any moment based on token cost, creating strategic incentive for labs to control the end-user account.
For AI labs, the ultimate biggest prize is the maximum amount of inference, which creates incentive to maintain a balanced ecosystem rather than competing in every vertical.
If some companies lean too heavily into a non-ecosystem approach, competitors emerge and balance the market — capitalism self-corrects.
The breakthroughs keep happening faster than the customer can implement any kind of standard architecture and those breakthroughs often times make obsolete the last thing you implemented.
The token cost and budgeting and budget planning probably is at least 1/3 of the hottest button issues that relate to AI, and it might even be tied for number one like half the time.
The line of business doesn't necessarily know how to budget for compute. They don't have finops for the marketing team or for the sales team.
One of the things that we don't have tooling for is how do you measure the ROI on the tokens.
the agent equivalent that would have been doing coding that just can consume all of the code base that it needs and generate whatever it needs. That agent in knowledge work is going to either bounce up against an entitlement issue like immediately and it's not going to have access to a resource, or it'll have access to too much in terms of resources and then start to answer questions with data that it shouldn't have
one of the memes is nobody's signing up for more than like one-year deals with the labs. And and part of that is because of the the pace of innovation that's happening.
everybody's getting a different definition to to their query because actually the way that company calculated things is like it's like no they do an FX adjusted number or they do a or they they they measure their you know net retention rate differently than what the model was trained on
The next generation leaving college faces potential complete and utter fear about whether jobs exist on the other end of AI transformation
Some argue that vertical apps or function-specific apps will be rendered not useful by a future training run that subsumes their capabilities — the 'wrappers on the model' critique.
Box's singular concept has always been to take super advanced technology breakthroughs and bridge them to the real world.
AI costs will escape the IT budget and move to line of business budgets because AI productivity gains extend beyond the 3 to 7 percent of revenue that IT spend represents.
You are going to have to centralize the management of IT systems and what you procure, but also decentralize the decision-making of how to use these things because the CMO should decide whether to spend a million dollars of compute or on marketing events.
Frontier AI model capabilities get applied to high-value tasks like coding and contract processes, but once a task can be performed reliably, you peel it off to a lower cost model and run it on an ongoing basis.
you've got kind of kind of five or six reasons that that AI coding looks very different from the rest of knowledge work
most agentic challenges I I think are kind of inversions of of of uh basically like you have a data challenge. Like the agent can't get access to the right information to to do the work. Uh maybe they have access to too much information...Or they don't have enough context to be able to execute the task
New technology mediums don't fully eradicate prior mediums — people end up with multiple devices (iPad, MacBook, iPhone) each doing something different, suggesting AI interfaces will coexist rather than one replacing the other
Enterprise software pricing will evolve to a dual model: a seat-based model for human end users (who have a right to use software and data via agents up to a certain allocation), plus a consumption model for purely agentic workloads beyond that threshold
AI can accelerate contract review or generation massively but there is still a lawyer on either end doing real work — removing one bottleneck still leaves another bottleneck, so jobs don't get eliminated as the first order effect
Jevons paradox applies to AI: large industrial companies like Caterpillar, Eli Lilly, John Deere will light up far more technical projects because one engineer now has the capacity of three to ten, increasing their demand for engineering capacity rather than reducing it
Future org charts will embed an AI IT capacity in most functions — an AI person or team in sales, marketing, and engineering — whose job is to identify daily workflows and bring automation to multiply output (e.g., testing five times the number of campaign ideas)
A useful mental model for understanding agentic AI value: ask 'what would I do if I had an unlimited chief of staff I could throw any task to?' — this opens up thinking about how to rewire organizational workflows
A 'bridge layer' between AI capability and end-user workflow is needed, and this layer is not merely a model wrapper — it encompasses data integration, bespoke workflows, change management, implementation, ongoing support, and domain expertise.
The hyperscaler precedent — where AWS marketplace pulls through products it could otherwise compete with because the bigger prize is infrastructure consumption — may be a model for how AI labs balance ecosystem partnership versus direct competition in the applied layer.
There's probably a $5 billion startup waiting to happen just in like ERP for your AI compute — how do I decide that all of this stuff is being used in the right way, how do I measure the value being produced, how do I make sure it rolls out to the right teams.
if you're going to have a world of agents and you want to have some flexibility on what agentic platform you deploy and what what type...then you need to get your data into a format that is going to work within that that kind of agentic ecosystem
Employees should spend 5–10% of their time getting really good at AI tools — use Codex, Co-work, Perplexity Computer, Cursor, connect them to a couple systems, try them on a personal workflow, and develop fluency
Employees should spend 50–100 dollars per month on AI tools — turn off a cable subscription if needed — to gain hands-on fluency with agents
“I just know enterprises.”
“Adam Smith, you know, figured this out a long time ago. Like division of labor is like a really powerful thing. Agents haven't fundamentally changed the concept of division of labor.”
“the difference between us doing a prompt with AI, seeing this incredible outcome, and we're like, "Oh my god, like obviously that thing could completely destroy this one application." To then the ongoing daily sort of mechanics of that product, the implementation of it in an in a workflow, the knowledge worker that doesn't have time for any of this stuff.”
Shared with Earmark · earmark-ai.com