"Microsoft's VASA-1: Creating Believable Talking Faces from Photos and Audio"

Microsoft Research Asia has unveiled VASA-1, an AI model that can create a synchronized animated video of a person talking or singing from a single photo and an existing audio track. The model significantly outperforms previous speech animation methods in terms of realism, expressiveness, and efficiency. Trained on the VoxCeleb2 dataset, VASA-1 can generate videos at up to 40 frames per second with minimal latency, potentially for real-time applications like video conferencing. While the technology has positive applications, such as enhancing educational equity and improving accessibility, it also raises concerns about potential misuse for creating misleading or harmful content.
- Microsoft's VASA-1 can deepfake a person with one photo and one audio track Ars Technica
- Weird Teeth Give Away the Fakery In Microsoft's Latest AI Video Generator Gizmodo
- Ever wondered what Mona Lisa would look like rapping? Microsoft launches VASA-1 AI bot that can make images ta Daily Mail
- Microsoft's AI app VASA-1 makes photographs talk and sing with believable facial expressions Tech Xplore
- Cool or creepy? Microsoft's VASA-1 is a new AI model that turns photos into 'talking faces' Tom's Guide
Reading Insights
0
24
4 min
vs 5 min read
88%
838 → 99 words
Want the full story? Read the original article
Read on Ars Technica