
Google's DeepMind RT-2: Revolutionizing Robot Control with AI Chat and Web Assistance
DeepMind's RT-2, a vision-language-action model, combines language and image data with robot coordinates to enable real-time instruction of robots. By training the model on images, text, and robot movement data, it can generate both a plan of action and the coordinates necessary to complete a command. The use of coordinates is a significant milestone as it integrates the physics of robots with language and image neural nets. RT-2 outperforms previous models in completing tasks with previously unseen objects and shows promise for further advancements in the field of robot learning. However, the high computational cost of large language models remains a challenge for real-time inference.