
"Apple Unveils Groundbreaking Multimodal AI with 30B Parameters for Text and Vision"
Apple researchers have developed new methods for training large language models using both text and visual information, resulting in state-of-the-art performance on AI benchmarks. The paper, titled "MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training," discusses the importance of architecture components and data choices in building performant Multimodal Large Language Models (MLLMs). The MM1 model family demonstrates enhanced in-context learning and multi-image reasoning, enabling few-shot chain-of-thought prompting, and produces competitive performance on a wide range of benchmarks.
