Google's DeepMind RT-2: Revolutionizing Robot Control with AI Chat and Web Assistance

DeepMind's RT-2, a vision-language-action model, combines language and image data with robot coordinates to enable real-time instruction of robots. By training the model on images, text, and robot movement data, it can generate both a plan of action and the coordinates necessary to complete a command. The use of coordinates is a significant milestone as it integrates the physics of robots with language and image neural nets. RT-2 outperforms previous models in completing tasks with previously unseen objects and shows promise for further advancements in the field of robot learning. However, the high computational cost of large language models remains a challenge for real-time inference.
- DeepMind's RT-2 makes robot control a matter of AI chat ZDNet
- Robots could now understand us better with some help from the web Popular Science
- How Google's Rt-2 Helps Robots Understand and Act Analytics Insight
- Google's New AI Model Controls Robots PCMag
- Google introduces new AI model that translates vision, language into robotic operations BOL News
Reading Insights
1
12
6 min
vs 7 min read
91%
1,235 → 105 words
Want the full story? Read the original article
Read on ZDNet