ºÎ»ê½Ãû µµ¼­¿ä¾à
±¹³»µµ¼­ ¿ä¾à
±¹³»µµ¼­ ÇÁ¸®ºä ÇØ¿Üµµ¼­ ÇÁ¸®ºä
±Û·Î¹ú Æ®·»µå ¹Ìµð¾î ºê¸®Çνº
   ¹Ìµð¾î ºê¸®Çνº 

åǥÁö







  • AI Steps Out of the Screen and Into the Real World


    The next frontier of artificial intelligence will not be defined by its ability to produce more natural sentences, but by its capacity to understand complex reality and move safely within it. Physical AI?capable of seeing, hearing, touching, judging and translating those perceptions into real-world action?has begun connecting robotics, manufacturing, logistics, healthcare and caregiving into a single technological movement. The moment AI enters the physical world, its performance standard shifts from providing accurate answers to taking appropriate actions and assuming responsibility for their consequences.

    [Key Message]
    * AI is expanding from the digital realm into the physical world. Physical AI does more than generate information; it perceives reality, makes decisions and takes direct action.

    * Intelligence emerges not only from algorithms but also through interaction among the body and the environment. A robot¡¯s structure, senses and movements are essential components of intelligence, making the body itself part of the computation.

    * The competitiveness of physical AI depends on its ability to connect vision, language and action. World models, simulation and data involving touch, force and joint states must be combined so that systems can act appropriately in unfamiliar environments.

    * Factories, logistics, healthcare and caregiving are emerging as major application areas for physical AI. The technology is moving beyond the automation of repetitive tasks toward systems that adapt to changing environments and assist human judgment and action.

    * The defining standard for AI that acts is not the scale of its intelligence, but its safety and accountability. The ability to stop under uncertainty and return control to a human will determine whether physical AI can earn trust in the real world.

    ***

    From Intelligence on a Screen to Intelligence That Acts
    Over the past several years, most advances in artificial intelligence have taken place in the digital realm. Large language models have learned from vast collections of documents to answer questions, while generative AI has produced text, images, music, videos and computer code. Predictive models have calculated weather conditions, financial market movements and the likelihood of disease, while recommendation algorithms have anticipated consumers¡¯ next choices. These systems have demonstrated remarkable capabilities, but their activities have remained largely confined to screens and servers.

    The physical world operates under entirely different conditions from the digital realm. If a word is chosen incorrectly in a document, it can be revised. If a robot grasps an object incorrectly, however, the product may be damaged. Awkwardly rendered fingers in a generated image may remain a quality issue, but an incorrect movement by a surgical or caregiving robot could endanger a person. A chatbot can regenerate its answer, whereas an autonomous vehicle facing a pedestrian who suddenly enters the road must select a single course of action within tens of milliseconds.

    The editorial ¡°From Embodied Intelligence to Physical AI,¡± published in *Nature Machine Intelligence* in April 2026, drew attention to this transition. Its central question was clear: what would artificial intelligence need to move beyond predicting, simulating or reasoning about the world and begin acting intelligently within real environments? This was not simply a matter of installing existing AI in a robot. It was a scientific challenge involving how intelligence combines with a body, how that body interacts with its environment, and how the resulting experience reshapes learning and judgment.

    ¡°Physical AI¡± does not refer to one particular algorithm or product. It describes a broad technological system in which sensors that perceive reality, AI models that interpret situations, control systems that plan movement, and robotic bodies that exert physical force operate in a continuous loop. Detecting an object with a camera is not enough. A system must determine what the object is, where it is located, how much force is required to grasp it, and how the surrounding environment will change once it moves. It must then perceive the outcome of its action and immediately revise its plan if reality differs from its prediction.

    Until now, artificial intelligence has primarily processed information about the world. It must now act within that world and bear the consequences of its actions. The rise of physical AI represents more than progress in robotics. It marks a transition from AI as a ¡°speaking tool¡± to AI as an ¡°acting agent.¡±

    Embodied Intelligence: When the Body Becomes Part of Intelligence
    Understanding physical AI requires an examination of ¡°embodied intelligence.¡± Traditional AI research has often treated intelligence as a process of information processing that occurs inside a brain or computer. Under this structure, an internal model calculates data received from the external environment and produces a predefined output. From this perspective, the body is little more than a passive device that carries out commands issued by intelligence.

    Embodied intelligence challenges this separation. It views intelligence as something formed not only in the brain but also through bodily structure, sensation, movement and relationships with the environment. The act of picking up a cup offers an intuitive example. A person does not merely calculate the cup¡¯s location. The person simultaneously uses the structure of the finger joints, the friction and pressure felt through the skin, the weight of the cup and the movement of the liquid inside it. If the cup begins to slip, the fingers adjust their grip immediately rather than waiting for a complete visual analysis.

    The body is therefore not merely an executor of instructions but part of the computation itself. A bird¡¯s wings and a fish¡¯s fins interact with their environments to produce complex movement naturally. Human muscles and tendons also absorb shocks and maintain balance, reducing the computational burden placed on the brain. The difficulty of the problem an AI must learn changes according to how the robot¡¯s body is designed. Flexible joints and elastic materials can make it easier to absorb collisions and adapt to irregular surfaces. The physical properties of the body assume part of the control algorithm¡¯s role.

    The editorial observed that several academic traditions were converging to explain intelligence that acts in the physical world. These included models that internally predict changes in the world, behavior-based robotics emphasizing direct interaction with the environment, ecological psychology connecting perception to possibilities for action, soft robotics using bodily structure and materials themselves, and approaches inspired by biological nervous systems and evolution.

    Embodied intelligence and physical AI, however, are not completely interchangeable terms. Embodied intelligence is closer to a scientific and philosophical concept that investigates the role of bodily and environmental interaction in the formation of intelligence. Physical AI broadly encompasses engineering efforts to apply such insights to systems operating in reality. An avatar in a virtual world may also be studied as an embodied agent, but physical AI must contend with actual gravity, friction, collisions, material deformation and sensor errors.

    Progress in physical AI will not come from a larger language model alone. The recognition that part of intelligence resides in the body and another part exists in its relationship with the environment is transforming both robot design and AI training.

    A New Form of Learning That Connects Perception, Language and Action
    Conventional industrial robots have demonstrated exceptional ability in controlled environments. A robotic arm in an automobile factory can quickly and precisely weld components arriving at the same location. If the component¡¯s position changes slightly or an unexpected obstacle appears, however, the robot may have to stop. This approach secures high precision by eliminating as much environmental variability as possible.

    The robots envisioned by physical AI are different. Without every possible situation being programmed in advance, they must understand human instructions, adapt to unfamiliar objects and spaces, and respond to changes that occur during an action. When instructed to ¡°clear the dining table,¡± a robot must distinguish among cups, plates, food waste and fragile objects, decide what to move first, and slow down or stop if a person approaches.

    Models that connect vision, language and action are central to this capability. Vision-language-action models do more than understand images and sentences; they generate movements for robots to execute. They interpret their surroundings through camera images, extract goals from natural-language instructions, and determine the direction in which arms and fingers should move or how joints should be controlled. Whereas traditional generative AI predicted the next word, an action model predicts the next movement.

    Real-world actions, however, are far more difficult than token generation. Even when handling similar cups, a robot must apply different levels of force depending on whether the cup is made of glass, paper or metal. If it contains liquid, the robot must also regulate its angle and speed. A movement that succeeds on a desk may fail in a moving vehicle or on a swaying vessel. An object that appears clearly visible may be perceived differently because of shadows, reflections, dust or occlusion.

    This is where a ¡°world model¡± becomes important. A world model anticipates the consequences of an action internally before it is performed. It estimates which direction a box will move when pushed, how a door will open when its handle is turned, and whether an object may fall if the robot moves too quickly. If a robot can compare the likely outcomes of multiple actions in advance, it can reduce trial and error and choose a safer plan.

    Large-scale simulation and synthetic data have also become essential foundations of physical AI. Teaching a single robot to walk by allowing it to fall thousands of times in reality is time-consuming, expensive and damaging to the equipment. In a virtual environment, by contrast, vast numbers of experiences can be generated quickly by varying gravity, friction, lighting, terrain and object placement. When a factory or warehouse is recreated as a digital twin, different tasks and hazardous scenarios can be tested without interrupting actual operations.

    Simulation, however, remains an approximation of reality. A robot that walks perfectly in a virtual environment may fall in the physical world because of slight floor slipperiness, joint wear or sensor delays. Closing this ¡°simulation-to-reality gap¡± requires virtual training and physical experience to be connected repeatedly. A cyclical structure is needed in which real-world data are reflected in the simulation, policies learned in simulation are tested in physical environments, and the results are used for further revision.

    Physical AI is also changing the meaning of data. If online documents and images were the fuel of generative AI, physical AI requires action data incorporating force, touch, joint states, distance, speed and spatial relationships. Such data cannot be collected easily from the web. Their formats also differ according to a robot¡¯s physical design and sensor configuration. Future competitiveness is therefore likely to depend not only on model size but also on the diversity and reliability of the physical experiences a system has acquired.

    Practical Adoption Beginning in Factories and Warehouses
    The places where physical AI is most likely to spread first are not completely unrestricted homes, but factories and logistics warehouses. Industrial environments have clearly defined objectives, can be managed within controlled boundaries, and allow returns on investment to be measured. Whereas existing automation equipment has handled repetitive operations, physical AI is targeting areas in which products and working conditions change frequently.

    In manufacturing, robots could classify components of different shapes and sizes, autonomously correct positional deviations during assembly, detect equipment abnormalities and revise the order of operations. As high-mix, low-volume production expands, the cost of reprogramming robots each time a product changes increases. Robots that learn new tasks from natural-language instructions and demonstrations could significantly reduce production-line conversion time.

    In logistics, packages differ widely in shape, weight and condition. A rigid box, a plastic-wrapped product and a fragile glass item cannot be grasped in the same way. Physical AI uses not only vision but also tactile and force feedback to estimate an object¡¯s properties and adjust its movements. The ability to recalculate travel routes by considering changes in order volume and workers¡¯ locations will also become increasingly important.

    Even more complex environments await in construction, agriculture and energy. The spatial structure of a construction site changes every day, while farms are continually affected by weather, soil conditions and crop growth. Power plants and offshore facilities contain hazardous areas that are difficult for people to access. In such settings, physical AI is likely to be introduced first not as a fully autonomous system that immediately replaces people, but as a technology that performs dangerous tasks or assists workers¡¯ judgment.

    The value of physical AI is not limited to humanoid robots. A humanlike form is advantageous for using stairs, doors and tools designed for people, but wheeled robots, robotic arms, drones or soft robots may be more efficient for particular tasks. Mobile robots may be appropriate for transporting heavy loads in warehouses, flexible grippers for handling crops without damaging them, and snake-shaped robots for passing through narrow gaps at disaster sites. The heart of physical AI lies not in reproducing the human appearance, but in finding the optimal relationship among body, task and environment.

    Corporate adoption strategies must also change. Instead of expecting one general-purpose robot to solve every problem, companies should begin with processes that have clearly defined scopes and allow relevant data to accumulate. They must assess not only task success rates but also downtime, error-recovery capabilities, the frequency of human intervention and the possibility of safety incidents. Physical AI is not simply a software purchase. It is a project that redesigns equipment, physical space, work procedures and workforce operations simultaneously.

    Higher Standards Required in Healthcare and Caregiving
    Healthcare and caregiving are among the fields in which physical AI could create the greatest social value. As populations age and shortages of medical and care workers intensify, the need for technologies that assist with patient transfer, rehabilitation, supply transportation and everyday activities continues to grow. Environments involving direct physical contact with people, however, demand far greater safety and sensitivity than factories.

    A caregiving robot must do more than execute instructions accurately. It must detect changes in a user¡¯s facial expression, posture, voice and movement, and respond appropriately if the person suddenly loses balance. Even when performing the same task, the robot must vary its speed and force according to the user¡¯s physical condition and level of anxiety. Technological success cannot be measured solely by whether an object was moved; it must also include whether the person felt safe and respected.

    Physical AI presents both possibilities and risks in surgery and rehabilitation. A surgical robot could improve precision by recognizing the shapes and movements of tissues and assisting the surgeon¡¯s control. A rehabilitation robot could detect a patient¡¯s muscular strength and responses and adjust the intensity of exercise accordingly. Yet medical data are difficult to obtain, and every patient¡¯s anatomy and condition differ. It is also difficult to collect enough examples of rare but potentially catastrophic situations.

    Medical physical AI should therefore develop by dividing autonomy into carefully calibrated levels rather than moving directly toward complete autonomy. Tasks that a system can perform independently, tasks requiring approval from medical personnel, and tasks that must remain under direct human control should be distinguished. The ability to slow down or stop when uncertainty rises and transfer judgment to a person must become an essential component of intelligence.

    A superior system in the age of physical AI will not be a robot that attempts to handle everything alone. It will be a robot that distinguishes what it knows from what it does not know and determines when control should be returned to a human. The purpose of caregiving technology, in particular, should not be to eliminate human relationships. It should allow technology to assume repetitive and physically demanding duties so that care workers can devote more time to conversation, emotional support and professional judgment.

    Safety and Responsibility That Carry More Weight Than Accuracy
    Errors in digital AI can generate false information or unfair recommendations. Errors in physical AI add collisions, falls, damage and injury to those consequences. For this reason, physical AI safety cannot be secured through a single filter or set of usage rules. Multiple layers of protection are needed, spanning mechanical design, sensors, control devices, learning models and operational procedures.

    The first layer is physical safety. A robot¡¯s speed and force must be limited; it must stop immediately when it detects a person or obstacle; and it must be designed to avoid dangerous movements even if power or communication is lost. The second layer is behavioral safety. Before executing a user¡¯s instruction literally, the AI must determine whether the action could endanger people or the environment. The third layer is operational safety. Organizations must define who supervises the robot, what intervention procedures apply when an abnormality occurs, and how incident records will be preserved.

    The hallucination problem seen in language models becomes far more serious in the domain of action. If AI produces nonexistent information, its answer can be checked. If it assumes that an occupied space is empty and moves a robotic arm into it, a collision may occur. A model that stops when uncertain is more valuable than one that is confidently wrong. Physical AI evaluations must therefore consider not only average success rates but also how a system behaves under worst-case conditions.

    Accountability is also complicated. If a robot causes an accident, responsibility may be difficult to allocate among the robot manufacturer, AI model developer, sensor supplier, operating company and on-site manager. For a system whose behavior continues to change through learning, certification at the time of release is insufficient. The effects of software updates and changing data on safety must be monitored continuously.

    Cybersecurity is also directly connected to physical safety. If an AI system with authority to act is compromised, the threat extends beyond information leakage: equipment, vehicles and robots themselves could become instruments of harm. On-device processing that allows essential functions to continue without a network connection, separation of command privileges, emergency-stop mechanisms and tamper-resistant activity records will become increasingly important.

    Social acceptance will not be determined by technical performance alone. Adoption outcomes will differ depending on whether employees regard a robot as a collaborative tool that reduces hazardous work or as a device that monitors their behavior and threatens their jobs. Companies must explain not only what a robot can do but also what data it collects, who controls it and how workers¡¯ roles will change.

    Physical AI Reopens the Question of What Intelligence Means
    The rise of physical AI is likely to transform the competitive landscape of the AI industry. In generative AI, data, computing power and model scale were the primary resources. Physical AI adds sensors, semiconductors, batteries, motors, reduction gears, materials, robotic operational data and simulation environments. Software companies cannot easily complete the entire system on their own, making collaboration among manufacturers, component suppliers, robotics companies, cloud providers and research institutions increasingly important.

    This development may create new opportunities for countries and companies with strong manufacturing foundations. Capabilities in precision machinery, electronic components, automobiles, shipbuilding, batteries and industrial automation are assets that form the body of physical AI. Competitive advantage, however, cannot be secured simply by placing a general-purpose AI model on high-quality hardware. The decisive gaps will emerge from action data accumulated over long periods in the field, operational knowledge and experience in responding to failures and exceptional situations.

    Changes in employment are also difficult to explain through a simple replacement narrative. Some repetitive and dangerous tasks may be automated, but roles involving the training, supervision and maintenance of robots will expand. So will roles that redesign workplaces for both humans and robots and verify safety and ethical compliance. The experience of skilled workers will not disappear; it may become the essential data that robots must learn. The central issue will be determining who owns that knowledge and under what systems of compensation and authority it is used.

    Physical AI also demands a revised definition of intelligence. The ability to solve examination questions and speak fluently is not enough to explain intelligence in reality. A system must interpret situations from incomplete sensory information, develop plans that account for the limitations of its body, adapt to unexpected changes and stop when danger arises. Intelligence is not only the ability to know but also the capacity to act appropriately and learn from the consequences of action.

    The pace of commercialization should not be exaggerated. A substantial gap remains between a successful laboratory demonstration and daily operation in the field. Performing an elaborate movement for several minutes is different from working for months without failure. Handling one unfamiliar object is not the same as understanding an entire home in which people, animals and furniture are constantly moving. Physical AI is advancing rapidly, but extensive validation will be required before generality, reliability and cost efficiency can be achieved simultaneously.

    The direction, however, is clear. AI¡¯s sphere of activity is expanding from the digital realm into the physical world. The winner of this transformation may not be the company that first introduces the most human-looking robot. Leadership will belong to organizations that understand the complexity of reality, learn the relationship between body and environment, manage the possibility of failure, and create systems that act in ways people can trust.

    Generative AI gave machines the ability to create words and images. Physical AI gives that intelligence senses, a body and real consequences. At that point, artificial intelligence no longer remains an adviser confined to a screen. It becomes an entity that grasps components in factories, moves goods through warehouses, monitors crops on farms and assists patients with rehabilitation. The defining technological question therefore shifts from ¡°How intelligently can AI answer?¡± to ¡°How safely and responsibly can AI act in reality?¡±

    The future of physical AI does not point exclusively toward perfect autonomy. It is moving toward a model in which humans and machines combine their respective strengths to solve tasks that were previously difficult or dangerous. The decisive standards for the age of acting AI will not be spectacular demonstrations, but reliability that withstands real-world exceptions and designs that preserve meaningful human control.

    Reference
    Nature Machine Intelligence, April 2026, Nature Machine Intelligence, From Embodied Intelligence to Physical AI