AI tokens are a consumable. The Pentagon should manage them like one.

lvcandy / DigitalVision Vectors via Getty Images
The Pentagon must build an allocation structure to precede its purchases of artificial intelligence tokens, writes John Mark Suhy, chief technology officer at Greystones Group.
The Department of War was celebrating in early June.
Nearly half of its 3.5 million employees, officials said, were now using AI at work. It was the kind of adoption number that makes for a good press cycle and a better budget justification.
A few weeks later, researchers and engineers at the Army's Combat Capabilities Development Command received an email asking them to throttle back. The Army Chief Information Officer had announced unlimited tokens in May.
By mid-June, the enterprise pool was empty and limits were back. Whether it gets replenished after October 1 was, according to the message, an open question.
Both things are true at once: the Army successfully drove enterprise AI adoption, and the Army ran out of AI. Those facts are not in tension. The second one happened because of the first.
This is being read in some quarters as a budget story: the Army under-bought, prices are volatile, someone will true up the contract next cycle. That reading is comfortable and wrong. DEVCOM did not overspend.
It did precisely what the department asked of it, and in doing so it surfaced a structural problem in how the enterprise buys and allocates AI. If the department treats that as an accounting problem, it will recur at a moment when the stakes are considerably higher than a delayed staff memo.
The department bought a license when it needed a magazine
Here is the structural error. DoW procured generative AI the way it procures software: seats, subscriptions, enterprise packs. But generative AI does not behave like software. It behaves like a consumable. Every query draws down a finite, purchased quantity.
The Army's enterprise pack came with 100 million tokens for the year. Individual users were allotted at least 200,000 tokens a month, with automatic top-ups when they ran dry. There is no version of that arithmetic that survives genuine enterprise adoption.
It is worth being precise about who ran out first. A research and engineering command exhausting its allocation is not a story about waste. Model-heavy work, code generation, test data, engineering analysis, materials and design tradeoffs, consumes tokens at rates that routine correspondence never will.
Heavy usage at DEVCOM is evidence that the capability reached the people best positioned to get real value from it. The problem is that the enterprise had no way to recognize that and fund it accordingly.
The military knows how to manage consumables. Fuel, ammunition, airlift, and satellite bandwidth are all finite, all contested, and all allocated through mature systems of prioritization: commander's intent, mission-essential task lists, allocation authorities that can surge one unit's share and throttle another's. Nobody issues unlimited JP-8 and hopes it works out.
Yet when the pool ran dry in June, the enterprise had no mechanism to distinguish between an account cleaning up a slide deck and one supporting a program that shapes a fielding decision. The throttle applied to everyone equally, which meant it applied to the highest-value work exactly as hard as the lowest. That is not rationing. That is a circuit breaker.
Adoption metrics created the burn rate
Compounding this: users who had signed up but weren't consuming their allotment reportedly received emails encouraging them to use more. Read that again in resource-management terms. The department made consumption the proxy for adoption, then measured its own success by how fast it drained a finite pool.
This is an entirely predictable outcome of how transformation initiatives get scored. Seats provisioned and queries run are easy to count. Decision quality, staff hours returned, and analytic depth are not. So the easy number becomes the metric, the metric becomes the target, and the organization optimizes for the thing that empties the magazine.
The readiness math is the part that should worry planners
Set the peacetime numbers against the wartime ones. Breaking Defense reported that DoW consumed roughly 20 billion tokens per day during the 38-day Operation Epic Fury. The Army's entire annual enterprise allocation was 100 million.
The steady-state figure and the contingency figure are not in the same universe. And the Army exhausted its steady-state pool in about six weeks of ordinary, encouraged, peacetime use: drafting, summarizing, coding, staff work.
Now run the scenario. A crisis begins. Demand for AI-enabled targeting support, ISR triage, logistics modeling, and course-of-action analysis spikes by orders of magnitude.
Does the pool have headroom? Does anyone know? Is there a mechanism to instantly deprioritize routine administrative use in favor of operational use, or does the system simply degrade for everyone at once, including the cell that needs it most?
If commanders cannot answer those questions, AI consumption is a readiness issue, not an IT issue, and it belongs in front of the same people who worry about war reserve stocks.
What governance actually requires
The fix is not to buy more tokens. It is to build the allocation architecture that should have preceded the buy.
Govern at the task level, not just the user level. Per-user caps are blunt. A useful control plane lets an administrator set limits by workload class (capping routine document drafting while leaving mission-critical analysis uncapped) so that scarcity, when it arrives, falls where it should.
Route to the cheapest model that clears the bar. A meaningful share of enterprise queries are being served by frontier reasoning models that are wildly overqualified for the task. Intelligent routing across a tiered model portfolio is the single largest cost lever available, and it requires no reduction in mission capability.
Instrument before you legislate. Allocation policy without telemetry is guesswork. Program managers need per-user, per-task, per-model visibility in near real time, not a quarterly invoice that reveals the pool emptied a month ago.
Hold a surge reserve. Wall off a protected allocation that routine use cannot touch, released only under defined conditions. This is war reserve stock logic applied to compute.
Stop buying consumption. Contract vehicles that reward volume are misaligned with an environment where volume is the constraint. Structure around outcomes and capped, prioritized draw rights instead.
None of this is exotic. It is the application of resource discipline the department already practices everywhere else to a resource it has been treating as though it were free.
June's email is worth taking seriously precisely because the consequences were absorbable. A capable workforce adjusted and the work continued.
The next time an enterprise pool runs dry, the department may not get to choose when, or who is left waiting.
John Mark Suhy is chief technology officer of Greystones Group.