| Average<\/td> | 7.6 \/ 20 \u00b1 1.7<\/td> | 17.2 \/ 20 \u00b1 0.9<\/td> | 8.4 \/ 20<\/td> | 15.2 \/ 20<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\nAnalysis<\/h2>\n\n\n\nTo better understand the advantages of our residual framework over the baseline, we conducted a series of quantitative analyses. The charts below highlight the specific improvements in efficiency and generalization.<\/p>\n\n\n\n (a) The residual corrects more strongly when the base action deviates from the goal direction. (b) Residual-augmented policies reduce completion time by 9% to 22% on successful episodes.<\/figcaption><\/figure>\n\n\n\n (a) Compatibility with different VLA backbones. (b) Object-centric observation transfers most effectively. (c, d) Residual-corrected rollouts improve supervised fine-tuning.<\/figcaption><\/figure>\n\n\n\nAction correction<\/h2>\n\n\n\nThe arrows show the base VLA action, the residual correction, and the combined action. When the base action becomes misaligned, the residual steers it back toward the goal.<\/p>\n\n\n\n \n \n <\/figure>\n<\/div>\n<\/div>\n\n\n\nConclusion and future work<\/h2>\n\n\n\nObject-centric residual RL combines the generalization ability of VLAs with the precise corrective capability of reinforcement learning. By choosing an observation interface that works in both simulation and reality, we can train the residual entirely in simulation and deploy it zero-shot on the real robot. As a result, the method improves manipulation performance without requiring real-world RL, distillation, or residual-policy fine-tuning.<\/p>\n\n\n\n Beyond direct deployment, the residual-corrected policy also produces better real-robot rollouts. These rollouts can retrain the base VLA and convert task-specific residual corrections into standalone policy improvements. In this way, residual RL can serve not only as an inference-time correction module, but also as a mechanism for generating higher-quality supervision.<\/p>\n\n\n\n Finally, future work includes extending the approach to more cluttered scenes and broader task families. It will also be important to develop more autonomous mechanisms for identifying which task-relevant objects should condition the residual policy. Ultimately, we view object-centric residual RL as one step toward robot learning systems that can use simulation to improve real-world behavior with less human intervention.<\/p>\n\n\n\n <\/div>\n\n\n\n <\/p>\n\n\n\n <\/p>\n","protected":false},"excerpt":{"rendered":" By\u00a0Kinam Kim, Namiko Saito, Heecheol Kim, Katsushi Ikeuchi, Jaegul Choo and Yasuyuki Matsushita Vision-Language-Action (VLA) models enable broad manipulation capabilities by leveraging large-scale pretraining and robot demonstrations. However, imitation learning can cause small execution errors to accumulate over time, pushing the robot into states that demonstrations did not cover well. Therefore, we present an object-centric residual […]<\/p>\n","protected":false},"author":44211,"featured_media":1176050,"template":"","meta":{"msr-url-field":"","msr-podcast-episode":"","msrModifiedDate":"","msrModifiedDateEnabled":false,"ep_exclude_from_search":false,"_classifai_error":"","msr-content-parent":1057371,"msr_hide_image_in_river":0,"footnotes":""},"research-area":[13556],"msr-locale":[268875],"msr-post-option":[],"class_list":["post-1175827","msr-blog-post","type-msr-blog-post","status-publish","has-post-thumbnail","hentry","msr-research-area-artificial-intelligence","msr-locale-en_us"],"msr_assoc_parent":{"id":1057371,"type":"group"},"_links":{"self":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-blog-post\/1175827","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-blog-post"}],"about":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/types\/msr-blog-post"}],"author":[{"embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/users\/44211"}],"version-history":[{"count":16,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-blog-post\/1175827\/revisions"}],"predecessor-version":[{"id":1176087,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-blog-post\/1175827\/revisions\/1176087"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media\/1176050"}],"wp:attachment":[{"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/media?parent=1175827"}],"wp:term":[{"taxonomy":"msr-research-area","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/research-area?post=1175827"},{"taxonomy":"msr-locale","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-locale?post=1175827"},{"taxonomy":"msr-post-option","embeddable":true,"href":"https:\/\/www.microsoft.com\/en-us\/research\/wp-json\/wp\/v2\/msr-post-option?post=1175827"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}} |