ºÎ»ê½Ãû µµ¼­¿ä¾à
±¹³»µµ¼­ ¿ä¾à
±¹³»µµ¼­ ÇÁ¸®ºä ÇØ¿Üµµ¼­ ÇÁ¸®ºä
±Û·Î¹ú Æ®·»µå ¹Ìµð¾î ºê¸®Çνº
   ¹Ìµð¾î ºê¸®Çνº 

åǥÁö





  • Why Sensible AI Deployment Matters More Than the Race for More Computing Power

    - Does Using More AI Produce Better Results?

    Companies competitively adopted larger models, more tokens, and increasingly complex agents. Yet assigning more work to AI did not necessarily increase performance at the same rate. The true measure of AI competitiveness was shifting away from computational volume and toward the ability to deploy the right level of intelligence for the right task.

    [Key Message]
    * More computation does not guarantee better performance. Indiscriminately increasing tokens and agents can raise costs, slow processing, and cause errors to multiply.

    * Not every task requires high-performance AI or multiple agents. Organizations should deploy conventional automation, smaller models, and reasoning models according to each task¡¯s difficulty and risk.

    * The cost of AI cannot be measured by usage fees alone. Companies must also consider latency, output verification, system maintenance, security risks, and energy consumption.

    * AI adoption should be evaluated by business value rather than usage volume. Instead of counting tokens, organizations should measure time saved, errors reduced, customer satisfaction, and productivity gains.

    * AI competitiveness comes from appropriate design, not simply from using the most powerful model. The key is to devote sufficient computation to important problems while choosing lighter tools for simpler tasks.

    ***

    Tokenmaxxing, the New Status Symbol of the AI Era
    As the meeting began, the person in charge introduced the company¡¯s newly built AI system. A market research agent collected data, an analysis agent interpreted its meaning, and a critique agent searched for weaknesses. Another agent wrote the report, while a final agent polished the language. Once a person entered a question, multiple AI agents exchanged opinions and produced a seemingly well-crafted result. From the description alone, it appeared as though a professional project team made up of people had been recreated in a digital space.

    The problem emerged afterward. The same question was repeatedly processed as it passed through multiple models, and the agents consumed enormous numbers of tokens while rereading and evaluating one another¡¯s answers. Processing took longer than expected, while the workload of the people reviewing the output did not decrease. Although the company had built an impressive system, producing a simple market-trends report could cost several times more than the previous method. More importantly, the final report was not significantly different from what an experienced employee could have produced by asking a general-purpose generative AI tool a carefully formulated question.

    In May 2026, the journal Nature Machine Intelligence described this phenomenon with the term ¡°tokenmaxxing.¡± The expression combined ¡°token¡± with the internet slang suffix ¡°maxxing,¡± meaning to maximize something as much as possible. It also satirized the belief that investing more tokens?the basic units AI uses to process text and generate responses?would automatically lead to better results.

    The editorial pointed out that companies, technology professionals, and researchers had become caught up in a frenzy to insert agentic AI into their workflows. Anxiety about falling behind created a race, and that race moved toward deploying more models, more agents, and more computational resources. Organizations began asking, ¡°Have we also adopted AI agents?¡± before asking what problem those agents were supposed to solve.

    Tokenmaxxing did not simply mean entering long prompts or receiving lengthy answers. It encompassed model calls beyond what was necessary, excessive reasoning stages, multi-agent architectures without a clear purpose, and automation whose effectiveness had not been demonstrated. In that sense, it resembled a new form of conspicuous consumption in the AI era. In the past, more servers and larger data centers served as symbols of technological prowess. Now, more tokens and more complex collections of agents were being treated as evidence of innovation.

    Yet a car does not travel faster on every road simply because it has a bigger engine. A racing engine can be inefficient and inconvenient when navigating narrow streets or traveling a short distance. AI worked in much the same way. Deploying the most powerful model for every task might appear technologically impressive, but it was not necessarily a sound business decision.


    The Paradox of AI That Thinks More
    During the early years of generative AI, model size was an important source of competitive advantage. Models trained on more data and equipped with more parameters generally understood complex language better and responded more flexibly to a wide range of questions. Later, attention turned toward methods that encouraged models to reason through several stages before producing an answer. Instead of answering immediately, a model could improve its performance on difficult mathematical, coding, and scientific problems by dividing the problem into parts, checking intermediate results, and examining alternative possibilities.

    Agentic AI took this idea one step further. Rather than merely responding to a question, it was designed to set goals, choose the necessary tools, search for information, evaluate the results of its work, and decide what action to take next. Multiple agents could also divide responsibilities and collaborate. One agent could plan, another could conduct research, a third could write, and a fourth could review the output. This structure demonstrated clear potential when applied to complex, open-ended tasks.

    Not every question, however, required deep reasoning. If a customer asked where a delivery was, there was no need for AI to examine multiple possibilities and conduct a lengthy chain of reasoning. There was also little reason for five agents to hold a discussion before reformatting an internal document or finding an available meeting time. For tasks with clear answers and repeatable procedures, a simple rules-based program or a smaller AI model could be faster and more reliable.

    More reasoning did not automatically guarantee greater accuracy. If AI made an incorrect assumption at the first stage, the lengthy reasoning that followed could merely wrap the error in a more sophisticated explanation. Even when multiple agents participated, they could repeat the same mistake if they relied on similar models and data. Cascading errors could also arise when one agent generated false information, another accepted it as fact, and a third built further analysis on top of it.

    It was similar to the way a meeting did not necessarily produce better decisions simply because more people attended. Additional participants could contribute a broader range of perspectives, but they could also diffuse responsibility, prolong the discussion, and obscure the central issue. AI agents did not automatically create collective intelligence merely because there were many of them. If their roles overlapped or their evaluation criteria remained unclear, they simply repeated one another while increasing costs.

    The appearance of an AI thinking for a long time could also inspire trust in users. A complex process and lengthy explanation appeared to suggest that the answer had been thoroughly reviewed. Yet the length of the reasoning and the accuracy of the facts were separate issues. An incorrect conclusion could still be delivered in highly logical language. The moment a company judged AI quality by the number of tokens or stages in a workflow, it began paying for the appearance of thinking rather than for performance.

    The Bill for Intelligence That Appeared to Be Free
    When an employee entered a question into a generative AI interface, the cost was almost invisible. An answer appeared within seconds, and no tangible raw materials seemed to be required. This experience made AI look like a resource that could be used without limit. Behind the screen, however, computing devices in data centers were working, and every token the model read or generated carried a cost.

    In enterprise AI systems, even small costs could grow rapidly with scale. The burden remained modest when one employee used the system only a few times a day. The situation changed when thousands of employees and millions of customers interacted with it. If multiple agents processed a single request, each read lengthy documents, and the output was reviewed repeatedly, one simple question could expand into a large number of model calls.

    Usage fees were not the only hidden expense. They were accompanied by latency costs caused by slower responses, operating costs for monitoring and correcting the system, labor costs for verifying inaccurate output, and security costs created when sensitive information moved among multiple tools. As AI became connected to external search engines, internal databases, email systems, and payment platforms, the number of points requiring management also increased. Every additional stage of automation created a new potential point of failure alongside the convenience it provided.

    The issue became clearer when viewed through customer service. If a high-performance reasoning model handled a simple question about business hours, the value of the answer barely changed, but the cost and waiting time increased. Conversely, assigning an excessively lightweight model to matters involving contract termination or insurance payments could raise the risk of error because customers¡¯ rights and money were at stake. The goal was neither to use the most expensive AI nor to use the cheapest one. It was to select a system suited to the difficulty and risk level of each question.

    The energy issue could not be ignored either. The more tokens AI processed and the longer its reasoning continued, the more computation it required. The consumption associated with one individual question might appear small, but the total changed when companies and users around the world repeatedly called AI systems in the same way. Expanding data centers, securing electricity, and operating cooling facilities were all connected to the physical limits of the AI industry.

    Tokens therefore became a management resource as well as a technical unit. Just as companies managed cloud usage, advertising expenditure, and logistics costs, they also needed to manage the tokens and computing power consumed by AI. A sustainable operation could not be built by allowing usage to expand without limits and checking the bill only at the end of the month. Companies needed to track how many tokens each task consumed and what value that consumption created.

    The questions were simple. Did customer satisfaction rise in proportion to the additional tokens consumed? Did employees spend less time on their work? Did revenue or productivity increase? Did errors and rework decrease? If a company could not answer these questions, higher AI usage was more likely to indicate unmanaged costs than successful innovation.

    Not Every Task Needs an Agent
    There is a saying that to a person with a new hammer, every problem looks like a nail. Something similar happened inside companies as agentic AI attracted attention. Document summarization, scheduling, customer service, market research, employee evaluation, and purchasing approval were all classified as possible tasks for agents. Some organizations even attempted to rebuild with AI processes that their existing systems already handled adequately.

    Different tasks, however, required different kinds of intelligence. Transferring data into a prescribed format or sending alerts according to predetermined conditions was often better suited to conventional automation. A small language model could be sufficient for classifying documents and extracting key information. High-performance reasoning models or multiple agents became valuable only when a task required identifying contradictions across several sources and comparing alternatives under uncertain conditions.

    Failing to distinguish among these tasks created the equivalent of assigning highly skilled professionals to routine administrative work. It was like hiring a top legal expert to number the pages of a document or asking a strategy consultant to enter the same figures every day. Powerful AI generated greater value when it was concentrated where it was truly needed.

    Sensible deployment began with the work, not with the technology. Organizations first had to determine which problems recurred, where the existing process consumed time and money, and how much damage an error could cause. Only then could they select the appropriate tool from among rules-based automation, search systems, small AI models, general-purpose generative AI, high-performance reasoning models, and multi-agent systems.

    For example, finding the number of vacation days specified in company policy required little more than document retrieval and simple answer generation. By contrast, comparing regulatory changes across several countries and evaluating the risks of launching a new product could benefit from an agent system in which research, analysis, counterargument, and synthesis were assigned to different roles. Speed and cost mattered most in the first task, while breadth of analysis and verifiability mattered more in the second. Applying the same AI architecture to both would not have been efficient.

    Classification by risk level was also necessary. If AI produced an awkward draft of an advertising message, a person could simply revise it. In recruitment, lending, healthcare, insurance, and safety management, however, even a small error could cause serious harm because people¡¯s lives and rights were involved. In these areas, organizations had to design not only for model performance but also for data quality, decision records, human review, and procedures for appeal.

    The authority given to an agent also required careful calibration. The authority to search for information and prepare a draft was fundamentally different from the authority to sign a contract or transfer money. As the range of actions available to AI expanded, convenience increased, but so did the potential impact of an accident. It was therefore safer to expand authority gradually?from observation to recommendation, limited execution, and finally conditional autonomous execution?rather than granting full authority from the beginning.

    AI adoption did not end with a single purchase. As with hiring a new employee, an organization had to define roles and responsibilities, evaluate performance, review mistakes, and adjust the scope of the work assigned. AI was not a person, however, so it had to be evaluated according to measurable results rather than the appearance of hard work or a confident tone.

    Asking the Right Questions About AI Investment
    Many companies were tempted to measure the success of AI adoption through usage. Dashboards displayed the number of employees registered, the number of accounts created, the number of prompts entered, and the number of tokens consumed. When these figures rose quickly, the organization appeared to be succeeding in its AI transformation. Usage, however, revealed activity rather than performance.

    It was the same as assuming that doubling the number of meetings would double productivity. Nor could a company conclude that collaboration had improved merely because employees sent more emails. Likewise, frequent AI use did not by itself prove that AI was useful. Employees could consume more tokens while spending more time than before checking and correcting the output.

    Performance measurement had to begin at the task level. Companies needed to compare conditions before and after AI adoption and examine how much processing time had fallen, how much more work each person could handle, and how error and rework rates had changed. In customer service, appropriate indicators could include resolution rates, repeat inquiries, customer satisfaction, and the proportion of cases transferred to human agents. In software development, deployment speed, defect rates, security vulnerabilities, and maintenance time were more meaningful than the volume of AI-generated code.

    Quality also could not be reduced to a single number. A system that responded quickly but was often wrong differed fundamentally from one that was accurate but excessively slow and expensive. Companies had to decide which factors?accuracy, speed, cost, and safety?took priority for each task. Speed and variety might be important during an internal brainstorming session, while accuracy and traceability took precedence when reviewing financial disclosures.

    This was where the concept of ¡°good-enough AI¡± became important. It did not mean a model that achieved the highest score under every condition. It meant one that reliably met the quality requirements of a given task while keeping costs and risks under control. If a 99-point answer cost ten times more than a 90-point answer, companies had to determine whether the additional nine points created enough real value to justify the expense. Conversely, in a field where a one-point error could cause enormous losses, paying the higher cost could be entirely reasonable.

    The performance of an AI system was not fixed from the beginning. It varied according to how users formulated questions, the quality of the connected data, the configuration of tools, and the review process. Before automatically replacing a model with a larger one, companies could often achieve better results by simplifying prompts, providing only the necessary information, organizing the search scope, and eliminating duplicate calls. This was a way to improve performance by redesigning the surrounding workflow rather than changing the model itself.

    Companies also needed a kind of ¡°AI budget rule.¡± They could set limits for the number of tokens, cost, response time, and agent calls permitted for a task, and investigate the reason whenever a system exceeded those thresholds. More computation could be allowed for critical assignments, while a lighter processing path could be applied to routine tasks. Not every vehicle needed to travel in the overtaking lane of a highway.

    Appropriate AI Rather Than Powerful AI
    Behind tokenmaxxing lay organizational anxiety as much as technical anxiety. When executives heard that a competitor had introduced AI agents, they worried that their company might fall behind. Departments proposed larger and more complex systems to demonstrate visible progress. Market trends and managerial impatience could easily take precedence over the actual needs of employees.

    Just as there were risks in failing to adopt AI, there were also risks in adopting it without sufficient review. If a company concentrated its budget and staff on a system of uncertain value, more important digital transformation initiatives could be delayed. Once an organization became dependent on a complex AI system, it could be difficult to discontinue it even if the costs increased. Employees could also forget how to make decisions or lose capabilities they had previously possessed.

    Sensible companies began with small-scale trials and compared the results. They could evaluate, under the same conditions, a process without AI, a process using a single model, and one using multiple agents. The objective was to select the approach that delivered the best performance for the cost, not the most complex system. In some cases, multiple agents produced excellent results. In others, a simple search function proved more dependable. What mattered was comparable evidence rather than technological fashion.

    The role of people also needed to be redefined. If people manually checked every AI output, the benefits of automation disappeared. If they checked nothing, the risks increased. Organizations therefore had to specify the conditions requiring human intervention. A task could be transferred to a person only when the financial amount exceeded a set level, the model¡¯s confidence was low, different agents produced conflicting judgments, or sensitive information was involved.

    The internal structure of responsibility mattered as well. It had to be clear whether the technology department that built the AI system, the operational department that applied it, or the executives who approved the budget were accountable for performance and failures. When responsibility remained unclear, everyone could claim credit for success and blame the model when a problem occurred. Even if AI supported a decision, people were still responsible for deploying it within the organization and giving it authority.

    A sound AI strategy did not emerge only by adding more components. It could also be created by removing unnecessary agents, reducing duplicate questions, replacing large models with smaller ones, and deciding that some tasks should remain under the existing process. The ability to simplify a system did not indicate a weaker understanding of technology. It showed a more precise understanding of technology¡¯s proper role.

    The next stage of AI competition was unlikely to be a contest over who could consume the most tokens. Companies were more likely to advance by producing the same results with less computation, concentrating sufficient resources on important problems, and managing both quality and accountability. Lower costs allowed more users to benefit, while simpler systems made errors easier to identify. Efficiency was not the opposite of innovation; it was a condition that allowed innovation to endure.

    The history of computing had been a story of expanding resources as well as using those resources more intelligently. Each time storage capacity, communication speed, and computing power increased, people treated them as though they were unlimited. Once the scale of use expanded, however, efficiency became important again. AI was entering the same stage. In an era when larger models and longer reasoning were possible, the ability to decide what not to do became increasingly important.

    The questions companies needed to ask also had to change. Instead of asking, ¡°How many AI agents has our company introduced?¡± they had to ask, ¡°Does this task need an agent?¡± Instead of asking, ¡°How many tokens did we use?¡± they had to ask, ¡°What value did those tokens create?¡± Rather than asking, ¡°Are we using the most powerful model?¡± they had to examine whether they were using the model best suited to their purpose.

    The belief that using more meant staying ahead was one of the simplest illusions created by the technology boom. The value of AI was revealed not by the size of the system or the number of agents but by the extent to which it solved a problem. Restraint?deploying sufficiently powerful AI where it was needed and choosing lighter tools where it was not?was becoming a new form of technological competitiveness.

    The winners of the AI era might not be the companies that built the largest digital workforces. They were more likely to be the companies that understood the nature of their work, measured cost and quality together, and clearly divided the roles of people and machines instead of displaying the complexity of their technology. The design that determined when deep thinking was necessary became more important than technology that simply enabled more thinking. Sensible AI deployment was not a strategy of using less technology. It was a strategy of using technology where it could create the greatest value.

    Reference
    Nature Machine Intelligence, May 2026, Nature Machine Intelligence Editorial Team, Stop ¡®Tokenmaxxing¡¯ and Deploy AI Sensibly Instead