Physical AI News (August 12, 2026 Issue)
Update Date: 2026/8/12
Executive Summary
August 11, 2026, saw a series of studies aimed at enhancing the 'execution efficiency' and 'field adaptability' of robot foundation models. RynnValue introduces general-purpose value judgment, SLIM-0.5B offers small-scale, low-latency inference, HIL-HARC presents short-duration online reinforcement learning on real hardware, and TempoWAM proposes state-dependent replanning. Additionally, JEPA-WAM, VANE, and WorldSimProbe delved into challenges regarding latent world models, test-time adaptation, and simulator quality assurance. On the business front, individual investor applications for Unitree Robotics' Shanghai IPO have flooded in, highlighting the clear expectations of capital markets for physical AI companies.


*The content of the created article was input into Gamma to automatically generate slides. If you find slides easier to view, please take a look here.
1️⃣ RynnValue: Pre-training a general-purpose robot value model with over 7,000 hours of video
📎 Source: arXiv:2608.09853 / Official Project / GitHub
A research team from Alibaba DAMO Academy and Hupan Lab has released 'RynnValue,' a value foundation model that estimates the remaining time to achieve language instructions from the temporal structure of robot videos. Trained on over 7,000 hours and approximately 3.09 million instruction-labeled clips, it achieves a Kendall's τa of 0.675 on RBM-EVAL-OOD without preference labels. In four tasks using a dual-arm Franka robot, it improved the online RL success rate from 52.5% to 72.5% and offline RL from 63.8% to 82.5% compared to the best comparative methods. It is attracting attention as a foundation that replaces reward design with large-scale pre-training and standardizes VLA post-training and data selection.
2️⃣ SLIM-0.5B: Realizing a low-latency, memory-efficient small robot policy with approximately 0.47B parameters
📎 Source: arXiv:2608.09771 / Official Project
A research team from multiple institutions has announced 'SLIM-0.5B,' a small robot policy with approximately 472 million trainable parameters. Through self-supervised learning using masked trajectory prediction, it integrates forward learning, which predicts future latent states from actions, and inverse learning, which recovers actions from state changes. It recorded 97.5% on LIBERO, 77.45% in zero-shot evaluation on LIBERO-Plus, and an average sequence length of 4.556/5 on CALVIN. The official project page reports an average progress metric of 67.8 across 5 real-world tasks, an end-to-end policy latency of 77.3 milliseconds, and a policy server GPU memory usage of 2.01 GiB. While the real-world values are progress metrics rather than success rates, this is a significant achievement in alleviating computational resource constraints for small robot policies.
3️⃣ HIL-HARC: Improving average success rate from 40% to 75% with 160 minutes of real-world online RL
📎 Source: arXiv:2608.09762 / Official Project
A research team has announced 'HIL-HARC,' a real-world online reinforcement learning framework that combines centralized learning/distributed execution with critic decomposition. It controls continuous arm movement and discrete gripper operation with separate actors, while a shared critic decomposes and evaluates task rewards and grasping rewards. After learning for a total of approximately 160 minutes across three real-world tasks—tennis ball, banana, and pot resetting—it improved the average success rate from 40% to 75% and reduced human intervention during learning convergence to 0%. While it shows the potential to complete final adaptation for each site in a few hours, evaluation of safety stops and certification levels remains a future challenge.
4️⃣ TempoWAM: Reducing WAM calls and failure rates through state-dependent replanning
📎 Source: arXiv:2608.09492 / HTML Version
'TempoWAM' has been proposed, which does not execute action chunks generated by WAM at a fixed length, but monitors task progress to decide whether to continue or replan. A lightweight monitor with approximately 2.27 million parameters was added to the existing WAM, and three tasks on a dual-arm robot were evaluated 30 times each. For simple beverage retrieval, it maintained a 90% success rate while reducing WAM calls by 26.9% from an average of 32.7 to 23.9, and for difficult hand cream packing, it improved the success rate from 50.0% to 63.3%. This is an implementation-oriented method that adjusts inference resources and error accumulation according to the difficulty of the situation without increasing the model size.
5️⃣ JEPA-WAM: Integrating state prediction and action generation for VLA with a latent world model
📎 Source: arXiv:2608.09381
A research team has announced 'JEPA-WAM,' which tasks the same predictor with state transition prediction and continuous action generation in the latent space of V-JEPA. It learns targets with spatial structure that co-encode the present and future, incorporating changes in the fine positional relationships of objects and robots into the policy without generating pixel-level future images. It achieved 79.2% on LIBERO-Plus without large-scale robot policy pre-training, and 86.3% in a configuration using π0.5, and also verified generalization to visual and spatial distribution changes in RoboTwin 2.0 and real-world dual-arm manipulation. Since the state transition prediction part can be removed during inference, the design aims to balance the advantages of a world model with low-latency control.
6️⃣ VANE: Safety-oriented test-time learning that verifies updates with future observations
📎 Source: arXiv:2608.09448 / HTML Version
'VANE,' a test-time learning method that adapts VLAs using unlabeled observations from the deployment site, has been announced. It creates update candidates in a shadow copy isolated from the policy in actual operation and only reflects them if improvement is confirmed through future observations, thereby suppressing destructive online updates. It recorded a 71.2% success rate on SimplerEnv's WidowX, an improvement of 3.2 points over the corresponding TTT baseline. In a single-prompt setting, it reduced backward updates per checkpoint evaluation from 45,824 to 612 (a 98.7% reduction) while increasing the success rate from 67.5% to 69.0%. The evaluation is in simulation, and the effects are not uniform in the Google Robot suite. Further verification including contact failures and rollback times is required for real-world deployment.
7️⃣ WorldSimProbe: Diagnosing the physical fidelity of WAM with over 18,000 cases
📎 Source: arXiv:2608.09298
'WorldSimProbe,' a benchmark that diagnoses action-conditioned world models not by whether the 'video is natural,' but by whether 'the given action correctly became robot movement and the environment changed based on that contact,' has been announced. It establishes five suites: local action calibration, global trajectory coverage, behavior retention by action source, interaction grounding, and interaction dynamics. Using over 18,000 cases from RoboTwin, ManiSkill, and LIBERO, it evaluated six open-source action-conditioned world models and visualized action compression, baseless object reactions, and inconsistent physical behavior. This is an important achievement as a quality assurance foundation before using WAM for planning or synthetic data generation.
8️⃣ Unitree Robotics: Over 8,000 times oversubscribed for Shanghai IPO, with a winning rate of approximately 0.018%
📎 Source: Shanghai Securities News / Unitree Robotics IPO Announcement / Reuters
In Unitree Robotics' Shanghai STAR Market IPO, subscriptions from individual investors exceeded 8,000 times, resulting in a final online winning rate of approximately 0.018%. The issue price was 150.80 yuan per share, the fundraising scale was approximately 6.1 billion yuan (about $900 million), and the valuation at the time of the IPO exceeded 60 billion yuan. According to Reuters, the valuation multiple reached 219 times 2025 earnings and 36 times revenue, reflecting both high growth expectations and valuation risks. While the subscription multiple itself does not equate to technological progress, the capital strength to simultaneously expand R&D and manufacturing bases for robot models, hardware, and new products could accelerate the competitive pace of China's physical AI industry.
Comprehensive Analysis
From the topics on August 11, 2026, it is evident that the competitive axis of physical AI is shifting from "larger models" to "how to efficiently cycle value judgment, latent representation, re-planning, and on-site adaptation at low cost." RynnValue redesigned rewards, SLIM and JEPA-WAM redesigned control representations, TempoWAM redesigned inference timing, and HIL-HARC and VANE redesigned post-deployment adaptation. Meanwhile, as WorldSimProbe demonstrates, naturally generated video and physically faithful simulation are not synonymous. As the success rate of real machines increases, safety stops, failure recovery, evaluation trial counts, and a lack of third-party reproducibility become the next constraints. The extreme demand for the Unitree IPO may accelerate capital inflows. However, to maintain a high corporate valuation, it is critical to determine whether research results can be converted into continuous uptime, application-specific ROI, and mass production quality.
Future Points of Interest
For RynnValue, the focus will be on whether the correlation between value scores and actual policy success rates can be reproduced across robots from different manufacturers and unknown tasks.
For SLIM and JEPA-WAM, it is important to demonstrate the advantages of their computation-efficient designs, including real-machine control cycles, power consumption, and latency fluctuations during long-term operation.
For HIL-HARC and VANE, design criteria are required for detecting online update failures, recovery time to previous policies, and conditions for resuming human intervention.
It remains to be seen whether TempoWAM's adaptive re-planning can maintain computational efficiency when integrated with safety monitoring mechanisms such as collision avoidance and emergency stops.
For Unitree, the turning point for evaluation will be whether they can link the R&D and manufacturing base development funded by the IPO to actual shipment volumes and commercial uptime.


