Urgent.News

What's breaking now, across thousands of outlets.

AI

Meta’s Llama 3.1 announcement

Meta, the renowned technology company, has recently unveiled its latest open-source large language model, Llama 3.1 405B. This groundbreaking model marks a significant milestone as it is now considered the world's largest and most capable foundation model available openly. With over 300 million downloads across all Llama versions, Meta has only just begun to explore the potential of this powerful tool.

The release of Llama 3.1 405B has set the stage for innovation, opening up unprecedented opportunities for growth and exploration. This new generation of Llama is poised to ignite a wave of novel applications and modeling paradigms, such as synthetic data generation and model distillation, which have never been achieved at such a scale in open-source models.

Additionally, upgraded versions of the 8B and 70B models are now available, boasting multilingual capabilities, an impressive 128K context length, state-of-the-art tool use, and enhanced reasoning abilities. These advancements enable the models to support sophisticated use cases like long-form text summarization, multilingual conversational agents, and coding assistants.

To promote open-source collaboration and development, Meta has made several changes to its licensing policy. Developers are now allowed to utilize the outputs from Llama models, including the 405B, to improve other models. Furthermore, Llama 3.1 405B and its smaller counterparts are accessible for immediate development on various platforms, such as llama.meta.com and Hugging Face.

During the development of Llama 3.1 405B, Meta trained the model on over 15 trillion tokens, a monumental task that required significant optimizations to the model training process. By leveraging over 16 thousand H100 GPUs, the team successfully trained the 405B as the first Llama model at such a scale. To tackle the challenges posed by the increased model size, Meta improved the quantity and quality of the data used for both pre- and post-training, resulting in higher-performing smaller models.

In response to user instructions, Meta aimed to enhance the helpfulness, quality, and instruction-following capabilities of the model. To achieve this, several rounds of alignment were executed, incorporating Supervised Fine-Tuning (SFT), Rejection Sampling (RS), and Direct Preference Optimization (DPO). Synthetic data generation played a crucial role in scaling the amount of fine-tuning data, with the majority of SFT examples being generated synthetically.

Rigorous data processing techniques were employed to filter the synthetic data, ensuring the highest quality for the alignment process. This meticulous approach enabled Meta to maintain the model's quality across various benchmarks and capabilities, even when extending the context length to 128K.

Written by urgent.news from ByteByteGo's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at substack.com →

More in AI

More from Tuesday 4 August →