
Android Bench 2.0 pits AI against multi-day coding tasks as GPT-6 Astra leads
Google has launched Android Bench 2.0 to evaluate how well large language models and AI agents handle complex Android development tasks, including long-horizon projects that can take days to complete. The benchmark uses continuous scoring rather than binary pass/fail and includes tasks like upgrading dependencies, adding major features, and building apps from scratch. Early results show GPT-6 Astra leading with a 28% pass rate, while Gemini 3.8 Flash trails at 8%, illustrating the current landscape of AI coding capabilities.
