Tokenmaxxing AI warning compares token counts to StarCraft clicks
A Fortune commentary says raw AI token use can mislead companies unless output is verified and measured for usefulness.
By Maya Lindqvist · Senior Technology Correspondent
3 min read
Tokenmaxxing AI use is drawing a warning from a Fortune commentary that compares raw token counts with a flawed gaming metric once popular among StarCraft players. The argument matters for companies adopting AI agents because higher usage can create more work to check, rather than better results.
The commentary points to actions per minute, or APM, a measure once treated by many real-time strategy players as a sign of skill. In games such as StarCraft, Age of Empires and Command & Conquer, players could raise APM by clicking more, but extra inputs did not necessarily improve decisions or win matches.
The author connects that pattern to Goodhart’s Law, the idea that a metric loses value when people turn it into the target. In this case, the warning is that organizations may treat token volume as progress even when the output still needs careful review.
What is tokenmaxxing in AI?
Tokenmaxxing refers to the idea that using more AI tokens is better, especially with agents that generate long chains of work. Tokens are the units AI systems process in prompts and outputs, so higher token use usually means more text, more computation and more material for people to evaluate.
Anthropic has said agents typically use four times as many tokens as chat interactions, while multi-agent systems use 15 times as many as standard chat. The Fortune commentary argues those figures can reflect cost and activity without proving that the work is useful.
Gaming research cited in the commentary complicates the idea that more actions equal more skill. A 2022 study of StarCraft players reported no statistically significant APM gap between expert and novice players, though the sample was small. A 2014 analysis found that action and reaction times tended to fall after age 24, while older players offset that with stronger judgment and more efficient actions.
Google DeepMind’s AlphaStar project is also cited as evidence against click-rate as the key factor. DeepMind reported that its StarCraft-playing AI could perform worse when allowed to click more, and that in demonstration matches it averaged lower APM than professional human opponents while still winning.
Why verification may become the bottleneck
The commentary says AI teams should pay less attention to raw token use and more attention to observability and verification. It cites an economic theory paper by researchers at MIT, Washington University and UCLA that argues the constraint on growth is “human verification bandwidth,” rather than intelligence itself.
The same concern appears in software development. A 2025 analysis cited by the commentary found engineers spent 9% of coding time reviewing and changing AI-generated code, compared with 4% waiting for AI code generation. In a talk cited by the commentary, AI researcher Andrej Karpathy said instant code generation still leaves him responsible for checking that the code works and does not add bugs.
The piece also uses the OODA loop, developed by U.S. Air Force Colonel John Boyd, as a model for better AI work. The loop stands for observe, orient, decide and act; the commentary’s point is that faster action has limited value if observation and judgment are weak.
Competitive gaming eventually produced a more selective measure, effective APM, which filters out repeated or redundant inputs. The commentary says AI lacks a comparable standard for verified agent output, leaving organizations at risk of counting activity instead of results.
This story draws on original reporting from Fortune.