Is the head of Yushu (a tech company) throwing cold water on things? How long until the "ChatGPT moment" of embodied intelligence?

PanewslabPanewslab

A growing trend is that the embodied intelligence industry is compressing the previously distant timeline of "general-purpose robots" into a single technology cycle in the future.

At a time when humanoid robots are at their most popular, the head of the "first humanoid robot stock" on the A-share market seems to have poured cold water on the industry.

On August 20, at the 2026 World Robot Conference, Wang Xingxing, founder of Unitree Robotics, admitted in his first public speech after the company's IPO that the moment of ChatGPT, embodied intelligence, has not yet arrived. It will be as early as two to three years, or as late as five or ten years.

Over the past two years, humanoid robots have become the hottest tech "fashion item." These futuristic robots are increasingly visible on stages of major galas, at game company exhibitions, and even at the entrances of shopping malls and retail stores. They dance, walk, and wave, sometimes nimbly, sometimes clumsily, interacting with audiences and shaking hands, rapidly moving from the laboratory into the public eye.

But behind the bustling demonstrations, the actual capabilities of these robots are clearly far from people's imagination of "future robots." At least in the more common expectation, robots shouldn't just be performing on stage or serving a marketing purpose, but should actually enter factories and homes, moving, organizing, operating tools, or undertaking some household chores and caregiving tasks. The public's question remains specific and direct: When will robots truly be able to work?

In the field of large language models, ChatGPT provided a highly symbolic answer. While not the first language model, it was the first to make a large number of ordinary users intuitively realize that AI had crossed a certain capability threshold. When a similar moment will occur in the field of embodied intelligence is gradually becoming one of the most pressing questions in the robotics industry.

 

The key to robots entering every household lies in their generalization ability.

At the World Robot Conference held on August 20, Wang Xingxing, founder of Unitree Robotics, pointed to a problem that has been repeatedly discussed in the industry as the biggest bottleneck of embodied intelligence.

In his view, many robot models can achieve a near 100% success rate in fixed scenarios, provided they have sufficient data collection and training. However, the success rate can drop significantly if the object being manipulated changes or the environment is slightly altered.

This is also the most important hurdle before robots can truly enter our lives and homes.

Wang Xingxing offered a fairly specific criterion: if one day, a robot can be brought into a completely unfamiliar home or environment and complete approximately 80% of the tasks using only voice or verbal commands, then embodied intelligence will have reached its "ChatGPT moment." In terms of timing, he believes it might take 2 to 3 years at the fastest, and 5 to 10 years at the slowest.

Wang Xingxing particularly emphasized that the real difficulty is not in making the robot "generally know what to do," but in the last few centimeters or even millimeters.

For example, if you ask a robot to pick up an object, the model may have correctly planned the entire action, and the robotic arm may have moved to the vicinity of the object. However, if the final positional error, tactile feedback, or grasping force is not correctly corrected, the entire task may fail.

This is also where robots and large language models are very different.

The input and output of a language model always occur in the digital space, while every perception and action of a robot must interact with the real physical world. Sensors have noise, actuators have errors, and the position, material, and weight of objects are constantly changing; these errors accumulate as tasks are performed. Therefore, for embodied intelligence, moving from "being able to do something basically" to "doing it consistently correctly" may be more difficult than initially learning a task.

Almost simultaneously with Wang Xingxing's aforementioned assessment, Wang He, founder and CTO of Galaxy General, also provided a more specific year. Wang He stated that with the continuous accumulation of data and further technological breakthroughs, embodied intelligence is expected to reach its "ChatGPT moment" in 2028. He defines this milestone as: robots being able to complete approximately 70% to 80% of daily tasks without specific training for any particular task.

Their judgments are actually very close. Wang Xingxing focuses on the task completion rate of the robot after entering an unfamiliar environment, while Wang He focuses on the direct generalization ability of the basic model without specific training. However, both of them set the critical point at a success rate of about 70% to 80% for unknown tasks.

This also means that the "ChatGPT moment" they're talking about isn't the mobility demonstrated by humanoid robots today, nor is it the 99% success rate achieved by a well-trained, fixed demo. The real change should occur after robots begin to break free from their dependence on fixed scenes, fixed objects, and specialized training.

 

When will robots truly "get it"? The consensus is 2 to 5 years.

Just a month ago at the 2026 World Artificial Intelligence Conference (WAIC), a roundtable discussion posed the same question to six leading entrepreneurs and researchers in the field of embodied intelligence.

On July 19, during the "Intelligent Embodied Forum" hosted by Zhiyuan Robotics and Mifeng Technology, the host asked the six guests a very direct question during the final roundtable discussion: How many years until the ChatGPT moment for robots?

The results were surprisingly concentrated. Yao Maoqing, partner of Zhiyuan Robotics and chairman and CEO of Mifeng Technology, gave an answer of 2 years; Tony Zhao, co-founder and CEO of Sunday Robotics, believed it would be within 3 years; Zhang Zhengyou, chief scientist of Tencent and director of Robotics X Lab, gave 3 to 5 years; Ma Yecheng, co-founder and chief scientist of Dyna Robotics, believed it would take about 4 years; Ren Zhiyi, research scientist of Physical Intelligence, judged it would take 4 to 5 years; and Xu Danfei, professor at Georgia Institute of Technology, gave an answer of 5 years.

The six individuals came from completely different companies and research institutions, and their technical approaches also differed, but their final answers all focused on the next two to five years.

Among them, Yao Maoqing is the most optimistic. His judgment is mainly based on the growth of data scale. In Yao Maoqing's view, the key factors currently restricting the further development of physical AI can be summarized as three walls: "data, representation, and closed loop". Compared with the massive amounts of text and images that can be obtained cheaply in the Internet world, real-world robot interaction data is not only expensive, but also limited by the differences in the robot itself, the task, and the scenario.

A truly ready-to-use robot needs to understand open-ended natural language commands and achieve a basic success rate of approximately 70% to 80% on common tasks. To achieve this, embodied intelligence requires data on a scale far exceeding current levels. He even suggested that in the future, it might require data on the scale of hundreds of millions of hours.

In other words, in Yao Maoqing's judgment framework, the "ChatGPT moment" is largely a scaling problem: as real-world data, simulation data, internet videos, first-person perspective data, and data flowing back after robot deployment gradually reach a certain scale, the model's capabilities may also cross the critical point.

Zhao Zihao of Sunday Robotics, on the other hand, has set a timeframe of less than three years. Rather than simply pursuing increasingly complex robot bodies, Zhao is more concerned with when the robot's "brain" will truly mature. Sunday Robotics has long focused on home robots and basic models. Their thinking is that if a sufficiently strong basic robot model can understand complex environments and tasks, even if the robot uses relatively simple and lower-cost grippers, it may be able to solve a large number of practical problems first.

Zhang Zhengyou's attitude was much more cautious. He gave an answer of 3 to 5 years, but emphasized that this timeframe doesn't mean the industry simply needs to wait for a larger model to suddenly emerge. A key prerequisite for the rapid scaling of large language models is that the Transformer has gradually become a more unified architecture. However, robotics intelligence has not yet formed such a clear technological paradigm.

From high-level cognition, visual perception, and task planning to real-time motion control, body feedback, and safety responses, a robot needs to handle problems at different time scales simultaneously. Therefore, true breakthroughs may require advancements in data acquisition, simulation, real-world deployment, failure recovery, and the simultaneous development of hardware and models, rather than simply replicating the scaling law of large language models. For this reason, when making his assessment, Zhang Zhengyou specifically emphasized that the industry still needs to "work diligently and practically."

Dyna Robotics' Ma Yecheng estimates a timeframe of around four years. His assessment is largely based on commercial deployment. For a lab demo, an 80% or 90% success rate might be sufficient to prove the technology's effectiveness; however, for a factory or service company actually purchasing robots, such reliability is often far from enough. Therefore, the "ChatGPT moment" of embodied intelligence not only means that the robot can understand more tasks, but also that it must be further improved to a level of reliability acceptable to real-world commercial environments.

Ren Zhiyi of Physical Intelligence gave a timeframe of 4 to 5 years. He stated that before joining Physical Intelligence, he once thought this process might take 10 years, but with the development of basic robot models over the past year, his expectations have been significantly shortened.

Physical Intelligence is one of the leading companies on the "bot foundation model" approach, aiming to enable the same model to gradually acquire broader operational capabilities through cross-ontology and cross-task data training. The truly important aspect of this approach is whether robot capabilities can continuously emerge as the model and data scale expand, ultimately producing an effect similar to zero-shot transfer learning in large language models.

Professor Xu Danfei of Georgia Institute of Technology offered the most conservative answer, but even that was only 5 years. Rather than betting on a specific model approach, he focused more on three more fundamental metrics: whether the model can truly learn from the data, whether the learned capabilities can generalize to new environments and tasks, and whether the inference speed can meet the real-time operational requirements of the robot in the real world.

When we put together Wang Xingxing's "at the fastest 2 to 3 years", Wang He's "2028", and the answers of the six guests at WAIC, an increasingly obvious phenomenon emerges: the embodied intelligence industry is compressing the previously distant timeline of "general robots" into a future technological cycle.

 

Is it a clear historical milestone, or a gradual process?

However, although these entrepreneurs and researchers are getting closer in their predictions about the timing, they are not actually using the same set of standards to discuss the "ChatGPT moment".

Wang He focuses on the leap in capabilities of the basic model itself. In his definition, embodied intelligence crosses a critical point similar to ChatGPT when a robot no longer needs to be retrained for each new task and can complete 70% to 80% of daily tasks using only the basic model. This standard primarily measures whether the model can move from "specialized intelligence" for specific tasks to general intelligence with strong transferability.

Wang Xingxing's standard is closer to the performance of robots in the real world. He also considers a task completion rate of around 80% as an important threshold, but emphasizes whether the robot can complete most tasks relying solely on language commands when entering an unfamiliar environment. In his view, many robots today can achieve a high success rate in well-trained, fixed environments. The real challenge lies in whether the model can continue to function after changes in the environment, objects, or tasks.

These two definitions seem similar, but they actually focus on different levels. Even if a model has good zero-shot or cross-task capabilities, it doesn't mean that a robot equipped with it has become a reliable production tool. Wang Xingxing mentioned at the WRC that Unitree has actually deployed robots in car factories and its own factories, and they can complete some simple assembly tasks, but large-scale deployment is still not possible. One important reason is that the efficiency of robots is still lower than that of humans, and they often need to be retrained to face new tasks.

Xinghaitu CEO Gao Jiyang offered an interesting dual assessment of this issue. In the industry roadmap presented at this year's WRC, Xinghaitu explicitly marked 2027 as the "GPT moment" in terms of technology, believing that after Scaling Law drives the basic model capabilities to cross the inflection point, vertical application scenarios will begin to accelerate; however, the real "commercialization inflection point" is placed in 2028, followed by deep penetration in 2029 and widespread deployment in 2030.

However, when discussing whether embodied intelligence would replicate a historical moment like ChatGPT, Gao Jiyang explicitly stated that he believes "there is a high probability that such a moment will not occur." The reason is that embodied intelligence will not be instantly popularized through a software product accessible to everyone, but is more likely to begin in B2B scenarios, gradually unlocking itself industry by industry and task by task.

Technological breakthroughs and instant public awareness are unlikely to occur simultaneously. This is because the explosive growth of GPT relies on everyone being able to experience it firsthand using a mobile phone or computer, while embodied intelligence's initial applications are in production settings such as factories and warehouses, making it difficult for the general public to perceive directly. Therefore, industry turning points will gradually emerge in a subtle and pervasive manner, becoming ubiquitous by the time people suddenly realize it.

These two statements are not entirely contradictory. Gao Jiyang actually distinguishes between the technological inflection point of model capabilities and the social inflection point of industrial diffusion: the former may appear relatively clearly in a certain year, while the latter requires scenario verification, large-scale production, cost reduction and business model maturity to gradually spread to the real world.

This is also one of the most significant differences between robots and large language models. After ChatGPT was released in November 2022, the changes in model capabilities could be communicated to users worldwide almost on the same day. For a software product, as long as the server and computing power can handle it, the cost of adding a new user is relatively limited. Users can immediately perceive the leap in model capabilities simply by opening a webpage and entering a question. Therefore, technological breakthroughs, product launches, and large-scale user adoption were almost compressed into the same time window in ChatGPT.

Robots, however, must enter the physical world through a real body. Even if a basic model suddenly gains significantly stronger generalization capabilities tomorrow, these capabilities still need to be translated into real movements through chips, sensors, joints, dexterous hands, and actuators, while also facing a series of engineering challenges such as accuracy, stability, safety, lifespan, maintenance, and cost.

Therefore, embodied intelligence may ultimately find it difficult to replicate a clear historical milestone like November 30, 2022. It is more likely to manifest as a series of gradually emerging tipping points. Perhaps, even if people look back in the future to confirm the "ChatGPT moment of embodied intelligence," they may not be able to find a single point that everyone agrees on.

This content is for informational and educational purposes only and does not constitute investment advice related to BTCC. BTCC makes every effort but cannot guarantee the truthfulness, accuracy, or originality of the content above.