Shanghai Artificial Intelligence Lab releases multimodal corpus

The Shanghai Artificial Intelligence Laboratory announced that it will join forces with members of the Corpus Data Alliance to jointly open source and release the "Scholar Wanjuan" 1.0 multimodal pre-training corpus. "Scholar Wanjuan" 1.0 combines the rich content accumulation of members of the Corpus Data Alliance and the data processing capabilities of the Shanghai Artificial Intelligence Laboratory, and will provide high-quality large-model multi-modal pre-training corpus for the academic and industrial circles. The total amount of open source data exceeds 2TB, and has four major characteristics: multi-element fusion, fine processing, value alignment, ease of use and high efficiency.

The open source "Scholar Wanjuan" 1.0 includes three data sets: text, graphics, and video. The text data comes from webpages, encyclopedias, books, patents, teaching materials, examination questions, etc. The total amount of data exceeds 500 million documents, and the data size exceeds 1TB, covering many fields such as science and technology, literature, media, education, and law; graphic data mainly From public webpages, processed to form interleaved graphics and text documents, the total number exceeds 22 million, and the data size exceeds 140GB (excluding pictures), covering news events, people, natural landscapes, social life and other fields; video data mainly comes from China Central Radio and Television and Shanghai Media Group, including news, film and television and other types of program images, the total number of video files exceeds 1,000, the data size exceeds 900GB, and the content covers military affairs, literature and art, sports, nature, knowledge, and video art etc.