Aaron Levie: Enterprise AI Is an IT Architecture Problem Wearing an AI Costume
Aaron Levie draws a sharp distinction between two AI deployment realities. Agentic coding has taken off because engineers already have access to the right data, can verify outputs, and the models were trained heavily on code. Knowledge work agents face the opposite conditions: fragmented data sources, non-technical users who can't steer agents effectively, and outputs that are hard to verify. He estimates agentic coding could compress a 12-month engineering project into two months, and that Box may do 3–10x more work as a result.
The enterprise deployment challenge is fundamentally an IT and data architecture problem. Levie's central argument — "if you show me your IT stack, I'll show you what you're going to be able to get from agents" — holds that scattered, poorly maintained sources of truth, missing permissions infrastructure, and legacy workflows prevent agents from functioning. Companies often need to re-engineer processes, migrate data, and set up skills files, MCP servers, and proper access controls before agents deliver real value. What looks like an AI conversation is frequently masking an IT architecture conversation.
Levie predicts a new internal role will emerge: an IT/business AI automation engineer embedded within lines of business, acting as a forward-deployed engineer for agent deployments. He also expects model routing to matter more as frontier model costs stratify — using the most capable models for planning and review while routing high-volume token work to cheaper alternatives. On budget, he suggests AI spend may need to reach ~5% of revenue, which likely requires line-of-business ownership rather than central IT budgets alone.
The bottom line is that the best AI implementations emerge from the employee level up, not the C-suite down. The strongest practitioners combine deep domain expertise with hands-on experimentation. Agents are simultaneously the most technical solution ever deployed to non-technical people, making CIO judgment and architecture more consequential than at any prior point in enterprise IT history.
Box CEO Aaron Levie argues that agentic AI has clearly worked for engineering teams but remains genuinely hard to deploy for knowledge workers. The real blockers are data architecture, permissions, and process re-engineering — not the models themselves. ROI requires high-judgment people close to the business, constant iteration, and a willingness to push past the initial hype.
CEOs should use AI technology so much that they get to the other end of the psychosis and can see in a practical and pragmatic way all the places where humans are still necessary to get ultimate gains.
I would just be playing with this technology constantly. I'd be I'd be pushing the limits. I'd be breaking things
I'm often prototyping in parallel just to see are there things that that I'm missing and can this expand the kind of you know use case or capability that I was trying to to kind of come up with
You need people close to the business with budgets that have high judgment, understand the technology and what it's capable of, and a rigorous process to constantly review where you are getting ROI from deployments.
Box counts 68% of the Fortune 500 as customers.
We are still in the very early stages of what agentic work looks like in the enterprise and what the rollout looks like.
In knowledge work, agents don't always have access to the right data, users are less technical and don't know how to steer them properly, and verifying agent output is much harder than in software.
Maybe 60 to 80% of engineering time goes into work that doesn't make the project differentiated per se — labor-intensive tasks like upgrading library versions, edge case testing, and building end-user features.
A big engineering project that from a standing start could take 6 months, a year, or two years could potentially be shrunk into two months using agentic coding.
At Box, agentic coding means they might be able to do 3, 5, or 10 times more work than before, or do it 3, 5, or 10 times faster.
AI agents have been trained on code from the internet as a plurality of their training data.
In engineering, if you're an engineer working on a project, you already have access to the entire codebase relevant to that project, so the agent also has access to all of that data by definition.
A CEO is the furthest away from the real work happening in the company of any role, other than maybe the board of directors.
Stories about people not getting real ROI from agents often approximate not pushing them hard enough, doing more menial work instead of pushing the limits.
The latest image generation models are perfect at being able to do text and have photorealistic capabilities in a way that wasn't possible 6 or 12 months ago.
In a project tested just yesterday, something that Opus 4.8 couldn't do, Fable finally actually did.
Most knowledge workers are not trained to set up MCP servers, CLIs, compute sandboxes, or manage agent data access — skills that are necessary for agents to be effective.
For agents to be really effective, data needs to be in the right format, skills files and agents.md readme files are needed, and permissions to turn on MCP servers must be granted.
Often companies have to re-engineer processes, migrate data sources to modern systems, and change workflows so agents can be more effective — it doesn't just miraculously work by embedding agents in existing business processes.
one of the problems that traditional enterprises have is our sources of truth are everywhere and many of them are not well-maintained
in one or two or three years from now, I think you're going to see this sort of stratification between the cost of frontier intelligence and the cost of sort of the second best frontier intelligence
if you show me your IT stack I'll show you what you're going to be able to get from agents
there was a brief moment where you could rely on the subsidization from venture capitalists and then get, you know, kind of tokens for your your uh your coding agents
there was definitely actually a period there where where you probably could have like really exploited it and like had like 10 years of software development, you know, paid for by uh by by LPs and and VCs
if the token spend goes up exponentially uh that that should actually be correlated with a good thing which is we can deliver more software to our customers faster
it budgets kind of run at depending on the industry maybe three five you know two to two to 5% of of sort of total revenue
in 2026 there's like you know maybe I have an hour you know heads up of information from from anybody else because like there's some chat thread that that's going on with Silicon Valley founders
you can be as informed as as the best expert in the world right now you get the same news feed as Andre Karpathy, you get the same news feed as as Greg Brockman or Dario Amade and and Sam Alman
I use Salesforce more today than at any point in history, maybe like five times more because I have an MCP server
we're going to move to a world where there's going to be, you know, maybe a hundred times more agents than people using software
for the most part you're not going to go and v code uh a CRM system or an ERP system
you're basically constrained today by the amount of people you have to work through spreadsheets and work through ERP systems and work through, you know, analytics data. Well, now you can actually throw compute at that problem and and all of a sudden you're no longer constrained by the number of people you have on the team
It is no harder or easier to measure ROI with AI than at any other point in history.
Keeping up with the technology is really important and there is no replacement to that.
The best AI implementations rise from employee-up versus the C-suite down approximately 90% of the time.
The person that actually owns the delivery of a particular project is in the best position to know the rate of productivity they can get with AI.
Agents are maybe the most technical solution that has ever been deployed to non-technical people.
Whether agents run wild and grab the wrong data or produce correct results is entirely determined by your technology architecture, how you deployed the agents, and how you trained users.
The CIO role becomes substantially more important with AI because it is the first time IT is responsible for deploying the actual real output of the organization, not just the tools that enable the output.
AI is the first time in history where companies can actually tap into all the unstructured data they have in their organizations.
In the enterprise, agents may have too much access to information, causing data leakage to the wrong people — security through obscurity breaks down.
A contract can't be verified computationally — it has to experience reality, the red lines from the other party, or being taken to court.
often times we're having an AI conversation but it's sort of masking an IT architecture conversation or a data architecture conversation
With the wrong prompt and the wrong limits, you could spend $50,000 and wake up the next day with that bill.
Giving agents access to all data without the right permissions is a security nightmare and will still produce wrong answers because the agent will have too many things to attend to.
There is a tale of two cities: agentic coding has clearly taken off within engineering teams, while knowledge work agents face a much messier, harder deployment environment.
CEOs experience 'AI psychosis' — an instant existential dread or high when first using AI — followed by a pragmatic phase where they realize agents still require significant human oversight and steering.
A new internal role will emerge — an IT/business AI automation engineer — embedded within lines of business but living in the IT organization, serving as a forward deployed engineer for AI agents.
if you were starting your company from scratch, you'd basically design your business processes in such a way where the agent has sort of innate ability to get access to the context it needs to help you automate work
model routing becomes very important where maybe something like a fableesque model gets the planning part of the work and the review part of the work but the in between sort of massive token usage comes from maybe something that is a more cost-effective model for that type of work
maybe you want AI alone to be 5% of your your total revenue so that would be a doubling of the IT budget. Well, to do that, then you ultimately need the line of business to own the budget
if I had unlimited capacity in you know XYZ area what would I what would I do differently in in that area of work
Measuring AI ROI is no different from measuring ROI for any business function: among all choices and opportunity costs, find the most effective use of the incremental dollar.
The strongest early-career AI fluency combines deep domain expertise (e.g., marketing principles) with understanding how AI accelerates that specific domain work.
If an ambitious AI project fails, try it again 6 months later almost every single time, because the rate of model progress means what didn't work before may now work.
we don't have a leaderboard internally. We don't try and incentivize the most number of tokens
I posed this question to our recruiting team the other day. I said, you know, if you could just comb through all of let's just say LinkedIn and you could you you didn't and instead of just sort of stopping at at, you know, the the LinkedIn profile, but you could then hop over to the internet and see like what were the GitHub projects that that person worked on
I'll actually add like and feel free to add anything else you come up with u into the prompt. And so that will be an explicit, you know, sentence or two in the prompt
“they're sufficiently distant from the last mile of work that still has to happen to generate most value with AI”
“if you show me your your IT stack, I'll show you your culture”
“This is not a sort of a one-stop oneshot kind of environment. This is an ongoing budget management process.”
“the right context at the right time with the right guard rails is is still a critical problem for agents to work with.”
Shared with Earmark · earmark-ai.com