Behind AI big model training, a data industry chain is forming

"Miracle" and "aesthetics of violence" are two words that have always accompanied the discussion of ChatGPT. In addition to "huge arithmetic power", "vigor" and "violence" also mean massive data. a16z founder Marc Andreessen also suggested at the Data+ AI conference that the massive data accumulated on the Internet over the past two decades is the most important factor in the development of AI. Marc Andreessen, founder of a16z, also suggested at the Data+ AI conference that the massive amount of data accumulated on the Internet over the past two decades is an important reason for the rise of this new wave of AI, because the former provides the latter with data that can be used for training.

According to OpenAI's disclosure, the text corpus of GPT-3.5 is as much as 45TB, equivalent to 4.72 million sets of China's four great masterpieces, while GPT-4 adds multimodal data on top of the GPT-3 and GPT-3.5 training datasets. And on July 18, Meta, the parent company of Facebook, released Llama2, the first open-source commercially available large language model, with a pre-training estimate of up to 2 trillion tokens.

Having the ability to access massive, high-quality data is seen as one of the core competencies of future big modeling companies, and is also a must for the major giants in the AI arms race. Data is also seen as a key production factor that determines future development. According to the statistics of the Digital China Development Report (2022), the potential of the digital economy that can be unleashed by the data factor will be incomparably huge, and China's data production in 2022 will reach 8.1ZB, accounting for 10.5% of the global share, ranking second in the world ranking, and the development of the digital economy is in a leading position.

However, data as a brand new factor of production also brings a series of urgent problems: how to understand data? How to authorize data? How to explore the value of data? Can data really be traded and circulated? Can data be recognized as an asset in an enterprise's financial statements? How to manage security? In this regard, we talked to Prof. Xueyun Zeng, Vice President of the Institute of Science and Technology of Beijing University of Posts and Telecommunications, and asked her to answer the relevant questions in depth.

The following is the transcript of the conversation:

Tencent Technology: ordinary people may be concerned about where the data for large model training comes from? Is there any use of my personal data, and will there be a problem of corroboration of these data?

Prof. Xueyun Zeng: These data calculated by big models are personal data. Personal data, compared to enterprise data, has an ownership problem. In principle, I am the master of my data. For example, the data generated on social software, in principle, the company owned by the social software can not use my personal data, although these companies have already practically controlled the data by default authorization, but how the specific data is used is regulated by the Personal Information Protection Law.

So if it's going to be used for big modeling calculations, how should it be used? Technically need to be anonymized, in the operation of the need for a market entity, that is, to give a certain company a legitimate right to operate these data, in other words, to find a market-oriented body of these data. When this market-oriented subject gets the data, it needs to invest manpower, time, intelligence and capital to produce the data, which can be called labor input. After the labor input, the data information belonging to the individual is derived into a kind of regenerative data of the company, or called secondary data. Then, the secondary data generates process data, and then data products and data services. At this time, the primary individual data with individuals as data owners is changed into data products and data services of the company. This is a process of productization.

Tencent Technology: Is it possible to understand in this way that Internet companies acquire individual data through authorization, and after the processual processing of these companies, it can be turned into some kind of data assets of this company?

Prof. Zeng Xueyun: It can also be understood in this way, we individuals generate a large amount of data on the Internet, just like various natural resources in nature. For example, the land can grow a lot of flowers, plants and trees, and there can be a lot of resources growing. This kind of resource is a public resource that can be developed and utilized, but not directly traded. What is generated after utilization and processing is the assets of the enterprise, which is permissible, and we should also encourage the development of data production factors in this way.

Tencent Technology: From an individual perspective, how can we protect our personal data and let them flow the way we want?

Prof. Xueyun Zeng: In the age of artificial intelligence, it's getting harder and harder to protect people's privacy. This is because everything people do is being recorded, geographic movement, living, working, eating, and living. Once recorded, this information, which originally belongs to us as individuals, can no longer be controlled by the actors. Therefore, the risk of privacy leakage at this time is very high, and the task of data protection is also very heavy, and the difficulty of data protection is also very high.

How can people safeguard their data rights? Actually, there are some commercialized approaches in various countries. The first one, like Japan, uses a data bank, which means that everyone can deposit their data in a data bank like a bank deposit. The data bank, which is a custodian of data, can itself act as an original developer of data value, and then individuals can also receive a certain amount of income. This is to say, it can allow a part of this part of the people who are willing to disclose and utilize their own data under a certain limit, can have a business model to solve the data protection problem in a self-selected way. That is, constructing a legitimate data flow, a legitimate model for the development and utilization of data, that's one piece.

The other part, that is to say, I personally do not want to, then do not authorize the data possessor. In the case of non-authorization, the state has to strengthen data protection. If who wants to illegally go to develop this part of the data, then we have to carry out disciplinary action, to carry out legal regulation, you can use blockchain technology to track such behavior. For example, our data has not been leaked, was leaked to where, go to the data flow tracking. It is also possible to carry out the tracking and analysis of data lineage, and there is now data lineage technology. Probably means that the data where it comes from, where it goes, data lineage analysis is actually a kind of data correlation analysis, as well as the traceability of data, with the word lineage is a very graphic description of the ins and outs of the data. Everything is being recorded, so this data and technology that records other people, it can also be recorded, it can also be publicized, it can also be penetrated.

Our Civil Code has special provisions for the protection of personal information in the chapter on personality rights. Article 127 of the Civil Code highlights the property attributes of data by juxtaposing it with network virtual property. In local legislation, Article 12 of the Shanghai Data Regulations directly reflects the "separation of human and property" rights allocation model. The article states, "The city protects the personality rights and interests of natural persons in their personal information in accordance with the law." "The city protects, in accordance with law, the legal or agreed property rights and interests of natural persons, legal persons and unincorporated organizations formed in the use, processing and other data handling activities, as well as the legitimate property rights and interests obtained from data innovation activities in the development of the digital economy."

On August 20, 2021, the thirtieth meeting of the Standing Committee of the thirteenth National People's Congress voted to adopt the Law of the People's Republic of China on the Protection of Personal Information, which will come into effect from November 1, 2021 onwards. The details can be found online. The juridical nature of personal information in the Personal Information Protection Law is also the protection of personality rights and interests, with little reference to the property rights and interests of personal information.

Tencent Technology: What exactly is the high-quality data that plays an important role in the training of large models?

Prof. Zeng Xueyun: Data should be the entire record of human economic, social, production, operation, business, and even military activities. Such a record, which is produced in various industries, fields and aspects. As far as raw data is concerned, it has high quality and low quality. For example, the financial statements of listed companies, financial data, is a kind of high-quality data, and it is a kind of structured data. Because such financial statements and financial information are audited by the society, audited by certified public accountants, and there is the Securities and Exchange Commission (SEC) to regulate the disclosure of information, so it is high-quality data. Another example is that the thesis data in the China Knowledge Network is also high-quality data. However, the data generated on the Internet is unstructured and non-standardized. Such data is a kind of raw, messy, unstandardized data, which needs to be cleaned at a granular level before calculation, so high-quality data usually has a processing process from unstructured to structured.

Tencent Technology: Since high-quality data can be produced continuously, why is there such a saying that "high-quality data is running out"?

Prof. Zeng Xueyun: I think the ability to produce and process data cannot keep up with people's demand for data, and the productivity of the entire supply chain value chain for data production and processing is still relatively weak. Because we know that data is constantly exploding, but high-quality data is running out, it just means that in the process of going from data to high-quality data, we lack a kind of productivity, a kind of integration capability. At this time there is a need for data vendors, we now have many data vendors, only in the direct utilization of data, but for the production and processing of data, for how to produce high-quality data, this piece of capacity or business model design is still very insufficient.

In fact, OpenAI's GPT-4 uses a large amount of data produced by the previous generation model, GPT-3.5, for training. the founder of OpenAI also said in a recent interview, "Synthetic data is an effective way to solve the shortage of data for large models. The key to this is having a system in place to differentiate between what's available and what's not in AI-generated data, and constantly providing feedback based on the effectiveness of the trained model." This company is not just able to raise money, can dominate a lot of arithmetic so simple, for the data of the product technology capabilities, is also one of the core competitiveness of this company.

Tencent Technology: In order to improve the productivity of high-quality data, what are the necessary aspects of industrial design?

Prof. Zeng Xueyun: On this issue, first of all, we need to understand what data is? What data do we have? And what to do with that data? That is to say, to produce high-quality data, it is not a matter of having the production capacity to have high-quality data, nor is it a matter of having the will to produce to have high-quality data. It must require an understanding of the data at the source, what problems in society are to be solved with the data? Where is the demand side of the market for data. Then, from the original data to the demand side, how to produce in the middle? This series of problems need to have industrial design in it, and the current overall thinking is not enough.

Tencent Technology: industrial immaturity is one aspect, does it also mean that this industry is still a blue ocean?

Prof. Zeng Xueyun: very early a blue ocean. Earlier there were some violations of the direct sale of data, and then the national legislation can no longer directly buy and sell data itself, no longer to trade raw data. Data can not do raw trade, it should be the result of their own production inputs to do trade, rather than that possession of what data, I go directly to sell data, which is not allowed.

2022 (December) introduced the "twenty articles of data", "twenty articles of data" which puts forward the requirements of data tenure resettlement, the ownership of data, the right to operate, the right to benefit from the resettlement of multiple tenure, which refers to the data to be carried out this hierarchical classification management. This is the top-level design of data governance, an overall blueprint. It can also be said that it is the beginning of the standardized development of the future data industry. At this time, people realize that data is not a whole, and to understand what rights and interests data actually have, which is also the original jurisprudence-based research advanced to economics-based research. To go to establish a data market, the market must be economic behavior. This economic behavior, to use a lot of economic tools, economic theory, so now from the study of data science, the state of data governance, to the academic research on data, industry to the use of data is a blue ocean, are a just beginning state.

Tencent Technology: so it seems that data can exist as some kind of asset of the enterprise, data belongs to which type of asset?

Prof. Xueyun Zeng: Data classification is a very hot topic in academia. In most cases, people will think that data is intangible, invisible and untouchable, called intangible assets. But in fact, from the point of view of ITU's classification, data is closer to inventory assets, because data also involves a process of production and processing. And the data itself is a kind of electronic tangible assets, why is it electronic tangible assets? Data it will occupy physical space, a lot of data itself also has a physical form, it is in the network side of a physical form. Picture, can see this electronic picture; sound, can hear the sound, portrait, can see the portrait, so the data it is digitalized tangible assets.