<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>Lab</title>
        <link>https://paragraph.com/@lab</link>
        <description>Trying out ideas</description>
        <lastBuildDate>Sun, 02 Aug 2026 08:09:36 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <image>
            <title>Lab</title>
            <url>https://storage.googleapis.com/papyrus_images/facd4003d2938e23a9f1411f497dd78f</url>
            <link>https://paragraph.com/@lab</link>
        </image>
        <copyright>All rights reserved</copyright>
        <item>
            <title><![CDATA[TOWARD DECENTRALIZED AI PART 3: RETRIEVAL AUGMENTED GENERATION]]></title>
            <link>https://paragraph.com/@lab/toward-decentralized-ai-part-3-retrieval-augmented-generation</link>
            <guid>keNBxxUjP83rrl53QgRH</guid>
            <pubDate>Tue, 05 Mar 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[IntroductionIn the Toward Decentralized AI series, we’ve already covered Data Collection and Model Building, two of the better-known / more obvious s...]]></description>
            <content:encoded><![CDATA[<div class="relative header-and-anchor"><h2 id="h-introduction"><strong>Introduction</strong></h2></div><p>In the Toward Decentralized AI series, we’ve already covered Data Collection and Model Building, two of the better-known / more obvious stages of AI. However, especially in recent years, another area of improving AI performance has been popularized: retrieval augmented generation. If you’ve interacted with any generative model, you’re probably familiar with the lacking, outdated or ‘made up’ knowledge issues. This was especially a concern during the first launch of ChatGPT / GPT-3.5 and the other models around that time. Since then, we’ve started seeing more and more models with internet access, PDF/text uploads, or narrower knowledge focus. All of this ties back to retrieval augmentation, which, in its simplest form, means bringing knowledge in before generating something. As seen in the figure below, before the release of ChatGPT, research and advancement in RAG was relatively slow after its inception in 2017.</p><figure float="none" data-type="figure" class="img-center" style="max-width: null;"><img src="https://storage.googleapis.com/papyrus_images/ec899dac91a4f2e06fc25a7ef8b7077b.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><p>RAG combines the power of large-scale pre-trained language models (such as GPT-4) with the ability to retrieve relevant information from external knowledge sources (such as Wikipedia) on the fly. RAG models can dynamically query the knowledge sources during the generation process and use the retrieved information to enrich and guide the output. RAG models have shown impressive results in various tasks, such as answering questions, text summarization, dialogue generation, and more. Below, you can see how the WikiChat project implemented a RAG workflow and improved model performance.</p><figure float="none" data-type="figure" class="img-center" style="max-width: null;"><img src="https://storage.googleapis.com/papyrus_images/b52b5fb6095f701d7fbefd34ce94d87c.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><figure float="none" width="437px" data-type="figure" class="img-center" style="max-width: 437px;"><img src="https://storage.googleapis.com/papyrus_images/cbf54d71949e3b371eb0e5606368e353.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><p>However, RAG models also face several challenges and limitations, such as the cost and scalability of accessing and storing the knowledge sources, the availability and quality of the knowledge sources, especially for specific domains or niche topics, and the interoperability and compatibility of the knowledge sources with different AI applications. Many of these problems are still being investigated by various teams since both the technology and the applications are evolving rapidly.</p><p style="text-align: start">In this article, we will explore how decentralization and blockchains can help address these challenges and enable a more efficient, robust, and collaborative way of creating and using knowledge-driven AI. As always, our approach will be strictly focused on AI’s problems that would make sense to be solved with decentralized methods rather than making up problems for decentralized tools to solve.</p><div class="relative header-and-anchor"><h2 style="text-align: start" id="h-rag-is-here-to-stay"><strong>RAG Is Here To Stay</strong></h2></div><p style="text-align: start">RAG models are important for several reasons.</p><p style="text-align: start">They overcome the limitations of fine-tuning, which is the conventional method of adapting pre-trained language models to specific tasks or domains. Fine-tuning involves updating the parameters of the pre-trained model using a task-specific dataset, which can be costly, time-consuming, and prone to overfitting or catastrophic forgetting. Moreover, fine-tuning does not guarantee that the model will have access to the most up-to-date and accurate knowledge, as the knowledge embedded in the pre-trained model may be outdated or incomplete. RAG models, on the other hand, do not rely on fine-tuning but rather on retrieving the relevant knowledge from external sources at the inference time. This means that they can access the latest and most comprehensive knowledge available and adapt to different tasks or domains without re-training. RAG models can also handle out-of-distribution or rare cases that may not be covered by the pre-trained model or the task-specific dataset by querying the knowledge sources for additional information. Below is a comparison of fine-tuning and RAG performance from the {paper name} paper:</p><img src="https://storage.googleapis.com/papyrus_images/868a614b2d2c3acb03374a04ceb7814f.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAcCAIAAACPoCp1AAAACXBIWXMAAAsTAAALEwEAmpwYAAAF/0lEQVR4nJVVMajk1hUVbDXVL37xYZYpPvGCiilUTJhChYLi+JuB/ELFQKbxss3Y/GIKxZnYKLFwBF+NwCpeYeWvH7GKF34EURKt4TVio0Zh5Wzk7CoGsSuIBRHYgt0Huw9HjUJ049n1z6TIqR5P6J577zn3PoExluf5/b/cz/M8y+6VZfmrjz+WZfnV7786nU7X67UoiqqqLpdLVVVnAxRFkWV5NpvJsjyfz09OTk5Pf6iqahzHf/3ss+wb5HnOORc45/kAuMrzPEkSx3Fc19V1HSFk27bjOBhj27Z1XTdN07Is0zQNw4CDZVkIIdd10zR9OU5RFF3XCf2Atm3Pz88vLy+3263jOAgh0zSBBiEENxhjz/vwN8Hvojv0o49uB0EArPAJY4wQ2m5/8qH3yyj6BML2ff8fgrqudV0nhFBKowGUUozxZrMJ/43fv/3jt9//xfuQ8jvvvgPV6LpuWZZhGPaAs7Oz8/Pz7XZrGO8yxrque0HAGEMI5XneNM3fv/ji0ePHTdMQ8mtJknRdT5Jks9nIsux5HiFE1/XxeGwYP/N9Pwx/+z1F2Ww2cRxblqUMZ4wxRH9B0HWd7/uO40DH27bt+z5N0+l0KopimqaU0tlsxhjr+z7Pc0EQwjDs+55zDsSQpSAItm1DwG8R9H3fDGAD4HNZlmdnZ47jdF2XZZlhGJxzuNc0rSxLINB1PY5jEFLTNDjvqaAYAAaAQFmW3bp1a7Vage1WqxVUVpbl8fFxlmXw43q9hqCc8+12+z8JEEKWZYEZqqrquo5SenR0NJ/Pi6LwfV+WZbhPkkQQhCiKuq5rmmY6nbquC8RHR0fQIkjxBQHnPEmSu3GcJAmlNEmStm0JIdPpdDabRVHk+z4wNU2TJMl3h8umaYqimM1mGOO2bbMsEwTBNE3G2FWCncgYY9f9IEmSvu/v3bs3m81UVa2qilKqqioUXhTFZDJJ0xQyUxTlzp1PdiKbprlf5JexE9m2bcju87997rouuKiu6+32p0VRAIHneSA4eB3urxJwzmEsPc+zbRuESpJEkqTpdJqmaRRFkiRB4XmeHx4eQpWc8/l8DpZt2/bw8BD02EOQ5/mnn/4ZjFRVFef8D1GkKMpqtXrw8MHl5eXrJ69XVdW2XyVJIstyFEVffvXlo8ePFosFIeTp0ydFUSwWi4vbF0+ePrmqAWPMdV1rAByGDRGORqPJZBKGIcZ4MpkAMaX02rVrYRhyzuu6FkURZqUoitFoZFlW13VXCaDweHBROqCu67Isl8vler1mjCVJslwuQYOiKA4ODqBFXddpmrab6tVqtX8OoIjnz59xzr/+59fPnj8DkS3Lsm27GrDbME3TbLdbELbrOowxDB3nnBBSVdV+DUzTfOvNtzabjTEgCAIYqPF4nKap7/vXr1+HSaaUCoJACAFhj4+PQdi6riVJwhjvIWCMVVVVD4B8m6YJguD09HS5XN79411CyOnpaZZldV2DBr7vw6Cpquq6btM0eZ6Px2PLstq2hWZ+a9Acx9F13bZt0zRd13Uch1IqiqIkSVmWUUpPTk6apoHWHRwcUErhx/l8jhDaDZplWfs1AOm7bwCBDMNACIGJXdeFFlVVtVgsdpOMEALBwYpwv2fZ0QFpmoKXsiwD52ialqZ/StP05q2bOwJJkl52EegBb4Pv+/sJwjB0XdfzPIQQIcTzvDAMv/PKK4qilGVJCFFVtSzLruvSNB2NRkEQwDYFDYB4Mpns36aMsSAICCFxHBNCwIKEkMlkIopiGIa+7yuKUhQFYyyKIkEQMMac87IsFUWBNj54+PAHr72GMd6zTcHCCKEwDC9uX4CXsyxTVVXTtH80TRzHi8Xi5Sdz98homgaCM8ZWq1UURf/3Nm2apixLSA38/vP33oMkOOe+78MGZYyZprl76a5W4DiOYRiEENM04Yc4jieTyY0bN9I0DcNwPp//dwVd14mieHH7AiZcEATHcfYPWhzH8JbFcZymaVEUURT9aLXSdf1+nhNCbr7xRpZlRVFQSmVZDoKgLEt4t3fn9fpNkHCnwb8Av4Cb0Z3eVTYAAAAASUVORK5CYII=" nextheight="888" nextwidth="1000" class="image-node embed"><p>RAG enables a more natural and expressive way of generating content, as they can incorporate factual, contextual, and diverse information from the knowledge sources into the output. RAG models can also generate more coherent and consistent content, as they can keep track of the information that they have retrieved and used and avoid repetition or contradiction. They can generate more personalized and engaging content, as they can tailor the output to the user's preferences, interests, or needs by retrieving the relevant knowledge accordingly.</p><figure float="none" width="681px" data-type="figure" class="img-center" style="max-width: 681px;"><img src="https://storage.googleapis.com/papyrus_images/a41e96c4e91c652410e455e319bcecf1.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><p>RAG models open up new possibilities and opportunities for AI applications, as they can leverage the vast and rich knowledge that exists in the world, both structured and unstructured, to perform complex and creative tasks that were previously difficult or impossible. For example, with RAG, one can generate summaries of long documents, synthesize information from multiple sources, answer open-ended or multi-hop questions, generate dialogues or stories with characters and plots, and more.</p><div class="relative header-and-anchor"><h2 style="text-align: start" id="h-its-not-all-figured-out"><strong>It’s Not All Figured Out</strong></h2></div><p style="text-align: start">Despite the advantages and potential of RAG models, they also face several challenges and limitations that need to be addressed.</p><p style="text-align: start">Doing RAG with publicly available knowledge is unnecessarily costly because everyone is finding knowledge, generating embeddings, and storing it on their own, repeating the same work done by others thousands of times. RAG models typically use a two-stage process to retrieve the relevant information from the knowledge sources: first, they use a dense retriever, which encodes both the query and the knowledge documents into vector embeddings, and then performs a nearest-neighbor search to find the most similar documents; second, they use a sparse retriever, which uses a keyword-based search to refine the results and extract the most relevant passages. However, this process requires a lot of computational resources and storage space, as the embeddings and the indices need to be generated and maintained for each knowledge source. Moreover, this process is redundant and inefficient, as different users or applications may use the same or similar knowledge sources and thus repeat the same work of finding, embedding, and storing the knowledge.</p><p style="text-align: start">While general-purpose knowledge sources are easier to find, specific domains or niche topics are more challenging, as a single person or team cannot gather data on them easily. RAG models rely on the availability and quality of the knowledge sources, which may vary depending on the task or domain. For general-purpose tasks or domains, such as trivia or news, there may be abundant and reliable knowledge sources, such as Wikipedia or other databases, that can be easily accessed and used. However, for specific domains or niche topics, such as medicine, law, or art, there may be scarce or unreliable knowledge sources that may be difficult to access or use. For example, the knowledge sources may be proprietary, paywalled, outdated, incomplete, biased, or inaccurate. Moreover, a single person or team may not have the expertise, resources, or incentives to gather, curate, or update the knowledge sources for these domains or topics.</p><p style="text-align: start">Other problems include the privacy, security, and trustworthiness of the knowledge sources, the alignment and compatibility of the knowledge sources with the user's goals and values, and the ethical and social implications of using the knowledge sources for AI applications. RAG models may also encounter other problems or risks related to the knowledge sources, such as:</p><ul><li><p>The privacy and security of the knowledge sources, as they may contain sensitive or personal information that may be exposed or compromised by malicious actors or hackers or by the RAG models themselves if they are not designed or regulated properly.</p></li><li><p>The trustworthiness of the knowledge sources, as they may be subject to manipulation, misinformation, or disinformation by the creators, providers, or users of the knowledge sources or by the RAG models themselves if they are not verified or validated properly.</p></li><li><p>The alignment and compatibility of the knowledge sources with the user's goals and values, as they may reflect different or conflicting perspectives, opinions, or biases that may not match or respect the user's preferences, interests, or needs or may harm or offend the user or others if they are not filtered or customized properly.</p></li><li><p>The ethical and social implications of using the knowledge sources for AI applications, as they may have positive or negative impacts on the individuals, groups, or societies that are affected by the AI applications or by the RAG models themselves if they are not monitored or regulated properly.</p></li></ul><div class="relative header-and-anchor"><h2 style="text-align: start" id="h-decentralization-can-help"><strong>Decentralization Can Help</strong></h2></div><p style="text-align: start">Among the problems teams face while implementing retrieval augmented generation solutions, there are many problems that would be solved or at least find better solutions by putting some decentralized proccesses or methods in place. Here are some examples.</p><div class="relative header-and-anchor"><h3 style="text-align: start" id="h-global-shared-vector-database"><strong>Global Shared Vector Database</strong></h3></div><p style="text-align: start">A global public vector database that stores embeddings in a decentralized storage solution like Arweave can enable us to only create embeddings for a knowledge once and store it permanently, and everybody can access it. Decentralized storage solutions, such as Arweave, provide a way of storing data on a distributed network of nodes that is permanent, immutable, and accessible by anyone.</p><figure float="none" data-type="figure" class="img-center" style="max-width: null;"><img src="https://storage.googleapis.com/papyrus_images/41216e6ebee61f23e336156cc78d2633.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><p style="text-align: start">A global public vector database, built on top of such a solution, can store the embeddings and the indices of the knowledge sources that are generated by the RAG models or by other users or applications and make them available for reuse by anyone. This can reduce the cost and redundancy of finding, embedding, and storing the knowledge and increase the efficiency and scalability of the retrieval process. Moreover, this can also ensure the privacy and security of the knowledge sources, as they are encrypted and distributed across the network, and the trustworthiness of the knowledge sources, as they are verified and validated by the network.</p><div class="relative header-and-anchor"><h3 style="text-align: start" id="h-dria"><strong>Dria</strong></h3></div><p style="text-align: start">Dria can be described as a multi-region, decentralized, public vector database. It functions as a collective memory hub for AI. It allows anyone to upload their unstructured or structured data, transforming it into easily queryable decentralized vector databases. This facilitates permissionless access to AI knowledge. All data is turned into vector embeddings and stored in a decentralized way, making it globally accessible to users and developers for various use cases.</p><img src="https://storage.googleapis.com/papyrus_images/f5537e4ad1de96acff88b65868f5da72.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAZCAIAAADfbbvGAAAACXBIWXMAAAsTAAALEwEAmpwYAAAET0lEQVR4nK2VXWzbVBTHLZ4QQjzshSd4QqsQFEZfNglatWISE6q0B4ooQnsFMZDWN5gY1V74GJWmVaqQRqWVaail6ZatVZSpQu2azm26hHz0Jk4cJ8537Dgfjp1eXzfXbowSsyztQmkLfx1dXR/7+ufzP9c2UUrEC1E2ByIpdyjupBJOOu1JJDbYrC8dX48zS3RsNdocY+47Gw/n5pfsNgqwXJArMAUuIkiFKtpSEcKaviNvqaUqKldRqYqKFaUxyohAklQtlCs5IR9OFRh+ybL4yfsfjV8ae/u1kxc+Pn/165/OvjM4+unFy59/80H/2f63+i5+9tXlkdGr346NXbpyb/ruuQ/PjY+Nk6vk4ODg7OysYRiapukN7dQNo24YhI4xRiqSocSXVVm13rz95vFu4rFeIJ5vTV5+8aVWniCI55559tVXuk68foIgiJELIwRB9Pf3G4ah67rRJsIwjPpOXccaRqoqI1VGYr4IxSpGWNvW5IJksywwgH5wf9m95gr5g+l4mlwmKV/A6/ZYLXeqFVnDjUeu1WrJZLJWq7Xf/W+AyTADq1jDmoY1c6JA5Y3u7lMnTxEE8cX5L7uOdw0NDfX19g0PD/f19p5570xPT48oik8/+F5Au+oN9xr2GYYhiqLDsWqz2ez2++Fw+Nat3xYXF5eWl30+PwABp3NjxeFo3bojowOgpUwmu+IgyXXnutMVY+O5HJ/NcZsg6P7T+5Bc23jkWnGQmwCUmxX8k/YDoKZkWTYMY2BgwOxtKEwDAIrFUq0pCJWnfT8owFStuX50dPTYsWPvnj6dz+ezuSxC6F8XHgigN2xtNL9T3pzsHB2AJLGUS4EH7kd/3Fty2Cma9rtcAa834PXy6azH5QH+gCBUiiX5qABFsVrvWizWqalfb0zdnPn9tmXOOjl5wzJnNaMsiltwO8tVjgjQMI7QTCKe4rg8z/GZVIbj+HQyw2X5RDwlCOXl+fWfJ6e8MW4fr/YD1FRVkaRqpaJIkhlIkqEoS4Io8mIxX/GuU26vhxMbFu3t0kEALdUfj2YcSgcC6LtfUb1New5byQ4AjDWGiQIAKIoShALGGkKqGZlMlqJCFBUCIBAO0xAqZl4QCgAEOI5vvxhCBWOtcwWWOet3319ZsNkhRGZUIdpWa17f5g8/jtE0w3E8x/HmKVXFyWT62vjEtfEJj8ePsSbLWxAiWd71bj8B7CkTY6099OY/pFVre978ZuxO7hyuB/9FTwAIIdim/T9hhwOYfQcAXL/+y/T0zPT0zMTEBEmuIaT+PwBTpu9ac9R1XWkUpLSMbl/T8tq0vr1tzY2EOvcAQgWAIEWFGCbq8/kZJtradjE2HgRUOEybZyFUTHYymZqfXyDJNZvNbrPZAAj6fP4IHem8TRFSaZqJMNEYGzdHQShU5SqESiaTDYUjLJuMRmMRhhWEotzMi2KFosIxNr4JghEmyrLJTCZbLJbaAX8B/hRi1mZNA0gAAAAASUVORK5CYII=" nextheight="782" nextwidth="1000" class="image-node embed"><p>Dria is designed to store information in formats understandable by both humans and AI, ensuring universal accessibility. It supports different data formats and is fully decentralized. Each index in Dria is a smart contract, enabling permissionless access to knowledge without relying on centralized services. It also features an API for knowledge retrieval, implementing search capabilities with natural language queries.</p><p style="text-align: start">Dria simplifies knowledge contribution and enhances access to AI knowledge. Users can effortlessly retrieve information from the Dria Platform by using the search bar. In addition to the platform, Dria provides native clients to facilitate the development of Retrieval Augmented Generation (RAG) applications. However, the most valuable aspect of Dria is its commitment to open and permissionless access to all knowledge within its platform. Developers can access any knowledge and retrieve information on their local devices using open-source Docker images released by the Dria team.</p><p style="text-align: start">Dria has ~2000 datasets from various domains, including full articles of <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://dria.co/knowledge/uaBIB4kh7gYh6vSNL7V2eygfbyRu9vGZ_nJ6jKVn_x8"><strong><u>English Wikipedia</u></strong></a> and <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://dria.co/knowledge/yQ8NzzdDUjQdVrcVIMcng5oxHI_OLCyMMzmX1KVOI8I"><strong><u>Wikipedia Simple</u></strong></a>. This ensures the Dria’s knowledge hub is useful for a large variety of tasks and use cases.</p><p style="text-align: start">Here’s a diagram showing how Dria saves time and money throughout various steps of the RAG process:</p><img src="https://storage.googleapis.com/papyrus_images/2ad87d0a85ee3d0a44781076b11989b5.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAWCAIAAAAuOwkTAAAACXBIWXMAAAsTAAALEwEAmpwYAAADk0lEQVR4nLVVbWscVRS+MYpUGwzW2BDTgLH6FyyI2kaiolIkv6EfI/itf6SfJKCm9ZUWl+0b24V8CAtZAlmS4Lor2cwmM5Pdmdmdl537svN+5d4bJ5PNVFHow90L98zZ85xz7nlmAH3GANlDwjHikXBLs9nc2tqq1+uKomxvb0uSRCmNouhsxJEI4J8fp7Bt2+BACPX7fdu26VOQTyCs6+vrKysr4hhxUJ77zs7O2f93u11VVcMwFKUIZ4xxsVjMcpwiaDabhUKBEEIp9TwSBAGltFKplEolEQVjTAgRzuVyuVAoeJ7HnRkopaqqLi0tCVbhdtIiQoiiKI1GwzTN1BiGYafTkdqSIEvh+36325VlWWSThkMIVSoVQXCqAtFlSZIO5ANN09LbQwhpmmYYBoQw2586R41DOKcEjuPkE4gmjGRKKQ2CQJQ/YnRs23GckUGKosh13adecr/f39vbS49i1zSt1WqNGC3LqtVq1Wo15RZ2x3HW1tbyCZIkCcPQ9/3kNCKO7GCkchnxHLHnEGQjZpEkie/7lmVBCF3XxRjnugnPLFOO0OgZnSXcFUK4sVFttVq/1+u7u7uEDJn5757xH43CKOQrRwdpIMPQ9/f3O51Ou92W2myoVFU9aLcxxiPEA8fRdM00Tdu2LMvq93vEOx7ZnDtgI+8lYRAOHGTohm3aPd3QNb2n98y+aegG01fMMmg0/rx/v/SktHb3bnH1+x/v3P7lhzu/rq7+/N23P927V3z4sFz47bFhmDlCg26g6z32gAeKoghBptshGWJMMCa+x6Z7efkmANOvv3Z5duqt6cm5qVfnz4OLALwCwCQAE3wHDx48ESLNECQUI2I7A9sZuC6EECGE4zgRPQ2jJI4T32MS+Wr55jkwfWnmnanJ+fMT8xcuvD0+NjsOZsbBzAtjb7wEZgB4+dGj8hkCVgE+lBVpv12rbdfrf6jqESEB8fzjRYauiyilN258DVjCU8+Bi89PXH7x3JtgbA6ASwDMAjAnKigWH+cQEOJDhDw/2KhuHsoK8XzbHjjOAEIEIeY7I7h165sr737y8eLS4kdfLi5cX7j6xbUPPr/6/mcf8rVw7fp7Vz7d3NxKvxYnBAjiTtfQdMPj+SIsXp3DOGZDzbY4oTH9rzghGJLgQD46lI9kpWNZtmU7LnvBZXSRUBpTPvT/gnyCJGTXOCSB7wcsYb5YUB6XrWNl/d8KnhH+Ak4mep/RiOtkAAAAAElFTkSuQmCC" nextheight="677" nextwidth="1000" class="image-node embed"><div class="relative header-and-anchor"><h3 id="h-crowdsourcing-knowledge"><strong>Crowdsourcing Knowledge</strong></h3></div><p>A public access decentralized vector database can enable efficient crowdsourcing mechanisms that make it easier to feed niche knowledge to AI applications. A decentralized vector database can also enable efficient crowdsourcing mechanisms that can incentivize and facilitate the creation and curation of knowledge sources, especially for specific domains or niche topics. For example, users or applications that need or use the knowledge sources can pay or reward the creators or providers of the knowledge sources using cryptocurrencies or tokens that are issued and managed by the decentralized vector database. Alternatively, users or applications that have or produce the knowledge sources can share or sell them to other users or applications using marketplaces or exchanges that are built and operated by the decentralized vector database. This can increase the availability and quality of knowledge sources and foster a more collaborative and diverse ecosystem of knowledge-driven AI.</p><div class="relative header-and-anchor"><h3 id="h-finding-the-right-knowledge"><strong>Finding The Right Knowledge</strong></h3></div><p>With crowdsourcing a large number of datasets, it quickly becomes much harder to find the right knowledge. To ensure the knowledge feeded to the AI model for RAG is in fact improving its performance, there should be a mechanism in place that finds and chooses the right knowledge based on some pre-determined criteria. This mechanism acts almost like a librarian bringing the relevant books based on one’s study topic, checking a set of high-performing datasets to see if they are relevant to the AI’s tasks. In a decentralized RAG system, such a librarian can find the necessary smart contract and retrieve the knowledge quickly, providing easy access to significant performance upgrades.</p><div class="relative header-and-anchor"><h3 id="h-knowledge-assets"><strong>Knowledge Assets</strong></h3></div><p>Knowledge assets (KWAs) are a novel form of tokenized real-world assets (RWAs) that represent units of knowledge on the blockchain. Unlike other RWAs, such as art or real estate, knowledge is intangible, dynamic, and subjective. Therefore, KWAs require a different approach to valuation, verification, and exchange. KWAs can be created by anyone who has some knowledge or expertise in a certain domain and can be verified by a network of peers or experts. KWAs can also be rated, ranked, and categorized according to their quality, relevance, and usefulness.</p><p>KWAs can be used for improving retrieval augmented generation (RAG) applications, which are natural language processing (NLP) systems that leverage large language models (LLMs) and external knowledge sources to generate text. RAG applications can benefit from KWAs in several ways. First, KWAs can provide a rich and diverse pool of knowledge for RAG applications to retrieve and augment their LLMs, improving their accuracy and performance. They can enable RAG applications to access and use proprietary, private, or dynamic data that are not available in public knowledge bases, enhancing their customization and personalization. KWAs can also incentivize RAG applications to produce high-quality and valuable text, as they can reward the creators and users of KWAs with tokens.</p><p>KWAs can facilitate the development and deployment of RAG applications, as they can reduce the cost and complexity of retraining LLMs for specific domains or tasks. Instead of retraining LLMs with large and static datasets, RAG applications can use KWAs as dynamic and modular knowledge units that can be easily updated, combined, and exchanged. KWAs can also enable RAG applications to explain their reasoning and sources, as they can provide metadata and provenance information for each knowledge unit. This can increase the transparency and trustworthiness of RAG applications, as well as their compliance with ethical and legal standards.</p><div class="relative header-and-anchor"><h3 id="h-ai-knowledge-blockchain"><strong>AI-Knowledge Blockchain</strong></h3></div><p>There can be an AI-knowledge-focused app chain/L2 where different apps can generate knowledge assets (share query revenue, stake for trust, tokenize any knowledge, etc.) and make them tradable globally. This increases the interoperability of RAG-dependent applications and makes it easier to bootstrap apps. A decentralized vector database can also enable an AI-knowledge blockchain, where different apps can generate knowledge assets that are based on the knowledge sources and offer them to the market. These knowledge assets can be exchanged or traded on the platform, providing a way for creators to monetize their knowledge and for users to access a wide variety of specialized knowledge. The AI-knowledge-focused blockchain could also provide a framework for establishing trust and verifying the quality of the knowledge assets.</p><div class="relative header-and-anchor"><h2 id="h-conclusion"><strong>Conclusion</strong></h2></div><p>While RAG has shown impressive results in various tasks even with its current state, there are some areas where decentralized infrastructures, applications, and assets can help RAG become more efficient and effective through higher quality knowledge and lower costs.</p><p>In the next article of this Toward Decentralized AI series, we will focus on how decentralization can play a key role in AI safety.</p><div class="relative header-and-anchor"><h2 id="h-references"><strong>References</strong></h2></div><ul><li><p><a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="http://dria.co"><strong><u>dria.co</u></strong></a></p></li><li><p><a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://arxiv.org/pdf/2305.14292.pdf"><strong><u>https://arxiv.org/pdf/2305.14292.pdf</u></strong></a></p></li><li><p><a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://arxiv.org/pdf/2312.10997.pdf"><strong><u>https://arxiv.org/pdf/2312.10997.pdf</u></strong></a></p></li></ul><div data-type="callout" type="info"><div class="callout-base callout-info" data-node-view-wrapper="" style="white-space:normal"><img src="https://paragraph.xyz/editor/callout/information-icon.png" class="callout-button"><div class="callout-content"><div><p>This article was originally posted on the <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.firstbatch.xyz/blog/">FirstBatch blog</a>. It was reviewed by <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/batuhan-aktas-38692b148/">Batuhan</a> and <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/kerim-kaya-552878129/">Kerim</a>, edited by <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/ilkyazyesilserit/">İlkyaz</a>.</p></div></div></div></div><p></p>]]></content:encoded>
            <author>lab@newsletter.paragraph.com (omer)</author>
            <enclosure url="https://storage.googleapis.com/papyrus_images/b6f7e814f05acd4f1e63749ffeecbebf.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[TOWARD DECENTRALIZED AI, PART 2: MODEL BUILDING]]></title>
            <link>https://paragraph.com/@lab/deai2-model-building</link>
            <guid>dY7uCAwae8yj4EnevOrb</guid>
            <pubDate>Wed, 14 Feb 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[This is the Part 2 of the Decentralized AI series. Focusing on model building.]]></description>
            <content:encoded><![CDATA[<div><div class="callout-base callout-tip" data-node-view-wrapper="" style="white-space:normal"><img src="https://paragraph.xyz/editor/callout/tip-icon.png" class="callout-button"><div class="callout-content"><div><p>This is the Part 2 of the Decentralized AI series. Check out the Part 1&nbsp;<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://paragraph.xyz/@lab/preview/aSkv5CGs8rWiIqZaOh8Q">here</a>.</p></div></div></div></div><p>Building models is, obviously, one of the most important stages of the AI pipeline. While the research and the progress have been going on for many years, the recent boom of consumer-facing AI applications and accelerating LLM adoption increased the focus on the development of new and better AI models.</p><p>Taking the AI models to their next level comes with a set of challenges such as:</p><ul><li><p>The need for large amounts of data and compute resources. They are often centralized, costly, and inaccessible for many users and developers.</p></li><li><p>The lack of privacy and security for the data and the models can expose sensitive information or lead to malicious attacks or misuse.</p></li><li><p>The difficulty of verifying and validating the models can result in errors, biases, or unfair outcomes.</p></li></ul><p>While throwing more resources at the problem may sound good, it doesn't typically address the efficiency, scalability, and accuracy-related issues that developers encounter.</p><p>In this article, we will explore some of the use cases and examples of how decentralization can be beneficial to AI model building, and what are some of the existing or emerging projects and technologies that are working in this direction. We will be approaching the existing problems by only applying decentralization concepts where they are needed instead of starting with the solutions we have and looking for problems.</p><p>Computing resources needed for the state of art models are significantly high. For this reason, a significant portion of the focus so far has been directed to decentralized compute protocols to enable a more efficient and democratic way of training AI models by allowing multiple nodes or parties to collaboratively train a model on their own data and compute resources, without sharing or transferring the data or the model. This aims to reduce the cost and time of training and increase the diversity and quality of the data and the model.</p><h1>Decentralized Compute for AI Training</h1><p>A major challenge in building AI models is the need for high-performance and scalable computing resources, which are often centralized, expensive, and limited for many users and developers. For example, running a complex or large-scale AI model can require specialized hardware or cloud services, which can be costly or unavailable for many applications or scenarios.</p><img src="https://storage.googleapis.com/papyrus_images/e9bd010738eb91ec13841ab0f0dc616e.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAATCAIAAAB+9pigAAAACXBIWXMAABYlAAAWJQFJUiTwAAAFzklEQVR4nG2UfUwTdxjHuxD9o4lZmBu6TvsHjCZbRVmsW6JLnG8QwQlxU7dVAXXFl03nQClvBWFtKViwSgtSqLVypXLlaHseV+6qtLycrVQJoFQMBWopBV//qQvcFu3SHgMxPvnle8/dH9/PPc9z99CCweDrN6/j1q5Z+enKr+LXrVixgsNZz2LFRkdHJycnMZmrmczVbDY7KuoTDmf91q1bqNvIyEgOh7Nr1y4Wi2UwGILBIEmSwfcFjbq4XEMEQXgeP55+9szn909OPfX4Jn3+KY9vMpz4qcTjm/R6PZ6xUc/YqPvRsHd8zDM2Ojnh8fsmJic83vGx50+f+H0Tz58+eTI9RSHnACRJjrofGdrxcxfrzDYHbCGgto6w2lCro6W9E+m802K2Ybf7JEX5deISK3rDBGhQUGsE1Cio1avrUVCHNGsbFTKN/IJcfE4uPvcuYPjh0LUWKKe8Gut0wrfsRqwbtTqMeA+lFMxi7ys4eTyPl9aokAmzT5wX5EBXG1BQZwI0suK8w0lbeCmJGUnb4j5e+jUz6u9XrxYA457xCe84CLcVVNXcIvraunrhW3Zzl3MOQMGw7g7nYP7J42czftIp5fVSsVxcrL4khdTKPF5aNI226gPagX17aUuW7UxI7DBBgUBgDuB2j6xbt3bo/n1dq0lwodZs610wxRcB2rp6lRdllYKz0vxszaWKklNHJblZNWUlfxXmbdi0+cPPouO++XbZqpiVMV9Ky8T//vN/i7xez46E7cMPXY57/S3tnZQd1Gabn4ER627rsMMWQgWAmdz9h35M4XK5JYLCktwzkqK8Exlp323ZvnvPvk2bt63bsPGLeE4mj9egqCbJ2TnA1JR/Z9LO4Ycu58AQag1N2NJzD7YQZlsvlRvxnlYsNAbU6lCUi9VVZQ2KakEuf/eevUd+5YlKSyqEpWnp6Zs2byvIza0QlhJmGAV1Cy2anp6OW7um02btsodaP9+QtxW2ECDSUQNAQmHZoYwMHi+zgM/n8XgbN289fvSYVFR6TlDIz86qkpSVFgnk5yVGXeNCBS9fvtyd8v3IyLBzYMiI97zdfTR8zLbeyzpo096Dp8XSP87kHMvM/PP06YMH0mtllVJRqSCX//Mv3B/27r+iuGQCNDqVUqeqR5q1iwCJiQn22wThvDvvSDGUTa2o1VEkqz2cV3yqtKJSpS2TlENXG5oalHWyylpZZaVIKCwuzsnK0qmUcNM16pNt1zctAoyNjsbGxrpcD/oGhwATpgJhwISp9QhgwkSKhosa3WG+oFh2ufCCIr9KcZafn56efl4ksrRcx/S6dn2TBQLxFh2mD1lTALjp2rsVQBDkm/DeG7zfaMDqm42NcLu4RnUZhCDchnTeuaAGNCakEW7X453qBpWouOhKjRwFtZQdCuoMjer5nAIYdY0zszOL/uT+/j773Xsq8EaRrFZar+GXX/zht6zdvN/rr7fCltBgqNnAej31vtSSCKuOWhXzT2Dt1UUVBIPBmdkZv28C7bDmV8klSjWIWkGkAzR3qUAEau9qxUKfaSvWA5q7QKAJ1l41AZqQUdO19593ACRJBgKBwYH+G7hFWF0rqlbcInqNqIW4O2hCbxLO/tY23HrbidmI/odup8PeZ+8edDq6LeYO9AasB7tvYpBWM6cWs6EZ6L6JoQb91JR/oYL5hR4IBLxeL0mSbvcISc6GNZQHg2+czt4jRw4dTOMeO34sOztLIikLrxl3IBBwuVwkST4YehAIBMY94zOzs263+8WL5yGA2+1Wq9U2mw0AAARBDAYjCII4jlOK43hdXR0AADiOy+Xy1NSUHQk7EhMTkpOTuFwujuPqcBjCgSAIGA6T0YggSFVVFUEQNARB6HR6fHx83Jo4Go3GYrEiIiLodPrSJUuZTCaDweBwOBEREVFRUbEsVkxMDIPB+DwmZvlHy9lsNiMcbDabyWTS6XQGgxEZGUmj0VJTU0EQLC8vHxgY+A/cmpzi33VWvwAAAABJRU5ErkJggg==" nextheight="1350" nextwidth="2220" class="image-node embed"><p>Decentralization can enable a more accessible and affordable way of accessing compute resources by allowing multiple nodes or parties to share or rent their idle or unused compute resources without relying on a central provider or intermediary. This can lower the barrier and cost of accessing compute resources, and increase the availability and efficiency of the compute resources. Currently, GPU concentration and, as a result, AI training are concentrated in the hands of a few companies, which creates a monopolization risk in the near future.</p><img src="https://storage.googleapis.com/papyrus_images/53feeace15b1cc259f7248b900337d04.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAATCAIAAAB+9pigAAAACXBIWXMAABYlAAAWJQFJUiTwAAADnklEQVR4nK1UT2gcVRz+3cSTF8GLFEHsQUGvVaoIWm0RpVVTD4rSmx68iCANpUJBRPTScw6xNDTWGAu5SElKhXbTZptsNrub3e1mZ3d23r75+3Yy2e28fW/emzcyu5s0WJpWzMd3+Bge873v936/HzT19o3MyuJSMZevrBWqDtkUUtI+Z3ybjDGekvb7Q8EYHYmUfA9yzmG9snF1PnP9RjazlF+4nimWaz16r9cPd0iFYIL3IxbJoeBUiH7EWcT6nO0++SBDTkENIOM4llIpJaQUUsp4xDgWXbzmesT3/ZVcFWPL73XduxWCDM/3gyDYOfkwjgySh0L1zDWMTddxcrkaMrDT6bh3y8TQLdf1fT95FFKDJEloGBYyN31C0l/u5fefAWEYVup1s1RYOPiin8uqJImFULsQyxGEEDtyW6RV3RuQlj5JaLNxDZ4IsjfTBFLuZwLO+VavZ1RqJ+Gl/N9ZqlTE+T5WCSjtb2562ew6wKHJX35A184yHqv9A4hBQdZLGsDhy+PvNCeA9gWlPaUSpeSA/ysNhCFFSL91Kw/wxvTpo2gK3nrv3PiZySRJhBDDJxnOx2O+6r8TyFGCBsDh6dNH278BwIkXXvl6/MzE0u2ihQ0ax3TgtA014GMnkDK2bbyYyQG8mSa4BPDMqaefPwXw9sWzXzhXjyz//FPh4oWVYtWyidMJpEqkjNNheRSS3QmKhTrAa5e+O5IaPPXpkwc+Bzg29e3raBqWDr06e+wDgI/nzn/j/PV+o5SxjQqNFUvnIqL9iHERy0jIOBIqkrGQ6SClS0bGowTbj/zuzLmP8CzAc18efPkrgLHfvz9uzR3Inzz+59hnAGNzP36IpsCYfNac+WRt4tfVK1fmL/yBV2aMtYXLs4v6+h27Mu9ptzvNnNXI+VYVbyynBsPdQIifvVPxcLPr1orrjWq1mVutebi55dR8bYPo+mpBI5a+5W50cLlna1uVklaulvLVrlkKHG15tem2tZ5d9nE59BqkVQrsumcU7u8izjkhjm6gFjJNE0spOsRpY7NlWDpCCGPbxrqBEHZM29WaRsuyWshwXKttug29ZdsYm6lG2MGmbTteU0eMR/cNpIx13QiCwHFspRLGWL2uhSENgs2GphFCTNMMgsD3O5ZpthHqEGKZpq7rhqH76d7eIh5pI8Q573W7GLdpGI66aGiQdvqw67cDMcYe/D64ymj0lFKGYXieN9SjzhnoYeMMDf4BUtFqUryNaTMAAAAASUVORK5CYII=" nextheight="1268" nextwidth="2096" class="image-node embed"><p>But there are some decentralized solutions that aim to tackle this market &amp; dominating players. Two examples are Bittensor and Gensyn AI.</p><h2>Bittensor</h2><p>Bittensor is an open-source protocol that powers a decentralized, blockchain-based machine-learning network. The project aims to let developers train machine learning models collaboratively and get rewarded in TAO according to the informational value they offer the collective. TAO also grants external access, allowing users to extract information from the network while tuning its activities to their needs. Ultimately, Bittensor’s vision is to create a pure market for artificial intelligence, an incentivized arena in which consumers and producers of this valuable commodity can interact in a trustless, open, and transparent context.</p><img src="https://storage.googleapis.com/papyrus_images/106eb231f0f8835caa54dac702532254.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAARCAIAAAAzPjmrAAAACXBIWXMAAAsTAAALEwEAmpwYAAADjElEQVR4nIVUMW/bRhTmz6gz2eogA0VsTw08hGKG0MpgVVlkp0gsy4vYQwEeaBQgI6OwSHkweVIH8iigEXFpUNEnDw5wKRAUhDPd1oDoJmdVl6YEDKSAZxX1wYacOPI3He8e33ffe987KbkKxpjreWHY7UWR63lyQWnadi+KwrDreh5jTIT5QYAxDsOuHwSqquoQXsZQSicTSpMfnHNKKcaYXmBhcckPAsZYP479IKCUcs4ZY2JNCGGM3ZXlpm3/9uoVIQRj3Isizvn1BIPBYau114uifH7+0aNvXc9bWFxCCG1Uq7NzOcdx+nH85viYMYYxlgvKXVnuRdHXd+7oEDYaOzMzt4Tc6wk45+125/GTJ5TSWm3LtKzXv7+WC4pIt1oq9aLoUgEhpK5pAADOuaqqQrS6UuxF0TQF4meEEADf6xD6QVD6ptxud1qtvcrauuM4om6iRDqEOoRh2K1UKk3b9gO/srY+TYEApdS0rCzL0jStbm42bbu6uck5Pzs7M4ztfhyzc+gQjs5R175rNHZqta1+HI/HY9fzMMY3EDRt27Se7jZ3Xc+ra5rreQghHULD2KaUih44jtO07R93d9vtjmgAQsi0LNOyCCHTCPpxTAgxDEOHMEkSdaUoKr5aKoVhtx/Hlz0wjG3Rg9VSCWPMGKtUKj8/e3ZzicKwK0mSXFBcz5MkyXGccvmhWAiPi7bncl/OzuUIIZIkbVSrpmVJknSzAkqp4ziKcq9cLiOEZudy++6+DmE+P29a1mBwmKZpkiT77r6qqsXiA4yDfH5eh9C0nn51+7YO4TQCxpgYn0Zjp9HYIeR5XdMIIb0o0iHsx7GYVWFK07Jarb3B4BAA4AcBIQQAII4mc16Zg7qmbVSro9Ho7BzZ//gnyzJhmA//fsAYV9bW2+3O+eZfJyfvsguMx+Msy4bDISHkegKBN8fHdU1bWFxaXl6+r97/qdPBGLueJ+54cHBw9PLo17hPCHE9r1bbEk+LYRhfzMwAAISJP0vAGEvTVFVV6QKMMdOyAABC4qTcJEkQQpRSPwjqmibi5YLy0Y0/ViBoXvzyghAihpZSWtc0YcTJGEJIsfhge/sHSilCSC4oC4tLq6XSZwk45wAAYf/hVfzx9q0orrpSrNW2FOWeHwSj0Wg4HJ6cvBMxf79/f3p6mqZ/TnMRpVQ8Bp/Kurw1Ic/Fq/fpKWNsMDg8enk0qeA/mKbirjJvJ5cAAAAASUVORK5CYII=" nextheight="605" nextwidth="1115" class="image-node embed"><p>Bittensor aims to solve some of the major problems of AI model training, such as data quality, data diversity, data privacy, model bias, model transparency, and model scalability. By leveraging distributed ledger technology, Bittensor enables:</p><ul><li><p>A novel, optimized strategy for the development and distribution of artificial intelligence technology, where models can learn from each other and share their insights without compromising data ownership or security.</p></li><li><p>An open-source repository of machine intelligence, accessible to anyone, anywhere, thus creating the conditions for open and permission-less innovation on a global internet scale.</p></li><li><p>Distribution of rewards and network ownership to users in direct proportion to the value they have added, creating a fair and sustainable incentive mechanism for AI development.</p></li></ul><p>Bittensor is still in its early stages of development, but it has already gathered a good amount of attention from crypto and AI communities, reaching billions of dollars in market cap.</p><h2>Gensyn AI</h2><p>Gensyn AI is a company that aims to provide a decentralized machine learning computing protocol. This protocol connects hardware that performs machine learning tasks, such as GPUs and CPUs, and makes them available to engineers, researchers, and academics. Gensyn AI claims that its protocol can lower the cost, increase the scale, and ensure permissionless access to machine learning computing.</p><p>The high cost of machine learning training is one of the main problems that Gensyn AI is trying to solve. According to Gensyn AI, traditional cloud providers like AWS charge high margins for renting GPUs and other hardware, making machine learning training expensive and inaccessible for many users. Gensyn AI’s protocol eliminates these margins by allowing users to access hardware directly from other users without intermediaries. Gensyn AI says that its protocol can offer up to 80% cheaper computing than AWS.</p><p>Another challenge that Gensyn AI is addressing is the limited scale of machine learning training. Gensyn AI argues that cloud providers cannot meet the growing demand for machine learning computing, especially for large-scale models that require massive amounts of data and computation. Gensyn AI’s protocol leverages the underutilized devices in the world, such as consumer GPUs, custom ASICs, and SoC devices, and connects them into a global supercluster that can provide more GPUs than the cloud. Gensyn AI uses advanced techniques such as model and data parallelization, distributed computing, and fault tolerance to enable efficient and reliable training over the network.</p><img src="https://storage.googleapis.com/papyrus_images/08e5c6d5f158436db46bf3044bd8a237.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAALCAIAAACRcxhWAAAACXBIWXMAABYlAAAWJQFJUiTwAAADrUlEQVR4nE2SS2wbZRSFf4lNVywDO3a03VDYdAFCUAkBQiKqqiLeTUsipQ2J2zzq4tTPepJpPLIT24ONnTguHjdmHI9NBsfxxIw9Gc+LmdjUVkwQCUoTZ+JXW7EHIWSbEqSjf3WPzv2/c0GzILJBbwKxFokIE/TkAh4eW6RQx4/fOMVwsEhEuNA8G/TxWCCNImzQK+NYTea6eigxaiXzwRUMPIeA045nz8xcPAXrP5yvKlzj6QxoFkQZx8JGrYxjMdiUQO4kECuFOiIWHWmfzi6gUchAwCYm4EnO3Q0btWGjtuusiuzD8jq5EgXgy7MfXXVgpO0+9uY7AwBofghSj8q8KuXbAQdc9tvb4zHYRMDmGGxigt4Ual8y6wjYHLFM0j5XCrXhkJ60T4WNtzD9zShkOBSYTgCj7qTdMyHwzCf+o+G//m48/lNd++PrE89/Zh2OP/lNOBTYdkBN5vZZmva5lkxfea4PrjptBGwhYPMKMnVPNxaFjARsYkN+HDKQdoiATVIkqEptZ1Vk1V/WfYElAEZvvDc2og1As66XXxw82dM3OrD4ZPt/Ac2CSPtdTNCz6rStIFMpFEmjdtrnSqN2CnUkECuPLTIBDxNoDzABT5XL1mS+KuX3HmQYfrnnpOXi2UsXBq65qanblOby0LsTX0y0KmJN6nTQBdoqKa3SZr0oRSy6hZsjBGz5DtLjkD5imbynGyPt0zhkoFBHSD9B+93NgqhKrCrl9wpZpYi/cvr6uHuIeIyOT98itiOmRJ+3b7RVkY4D6grfLIgNRWgogiowW8l4x8+qArNNkXvMehdjG0t79/zxFYm53d3UnMbe//mlhaTl3Jkr4/2jOvT9TR/WKsm1DknQUIRdek2KBEtkbCsZ34xiO/TaVjJeJqNlMrpDpyqpRJkkdui10vfLv1LJSirxX0BN5nYEqlIlzefNN17rxap3/JJ28lTfQYY5KuSPOquAusLvs7SyHKZQB+1345A+5bSlUTsO6VedNg4LdIqxZryzMdiU8c6mUaRbcvvrEn/488aD3cRLb9sAGDn/qabn3BAAE3Ig/qhcqEviv4gaitDBuqEKjCowe0xGFZh9luZC87Tf3a2XC83LOJacs5XJaJukxB5IuYqcJItzxt6rr4J+AEYAuAbAIAAjd8FlIgdvK6ma9LTko/bL12S+rhxrn6W7OuDaUiW2KuS6ZA9/2thXspxyP0JPExoDPjwz+8bMC73uE2+5kdeR+McQljCSiuv3Iv0P4ey0Y6oNt78AAAAASUVORK5CYII=" nextheight="754" nextwidth="2184" class="image-node embed"><p>They believe that machine learning training should be decentralized and uncensored and that anyone should be able to own and control their own AI models. Gensyn AI’s protocol allows users to train models without interruption on any available GPU on the network without relying on centralized authorities or intermediaries. Gensyn AI also uses cryptography and privacy-preserving techniques to ensure the security and integrity of the data and models on the network.</p><h1>RLHF</h1><p>Reinforcement Learning from Human Feedback (RLHF) is a technique that trains an AI system to optimize its behavior based on human preferences. RLHF is needed because many AI tasks, especially those involving natural language processing, are difficult to define or measure using algorithmic criteria or metrics. For example, how can we evaluate the quality of a generated story, a conversational agent, or a code snippet? These tasks depend on subjective and context-specific human values and expectations, which are hard to capture in a predefined reward function.</p><p>RLHF works by collecting human feedback on the AI system’s outputs and using it to train a reward model, which predicts how good or bad output is according to human standards. The reward model is then used as a surrogate reward function to guide the AI system’s learning process using reinforcement learning. This way, the AI system can improve its performance and align its behavior with human preferences without requiring explicit rules or references.</p><img src="https://storage.googleapis.com/papyrus_images/c99be2b4900c238c1748b38f8c18de8d.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAYCAIAAAAUMWhjAAAACXBIWXMAAAsTAAALEwEAmpwYAAAFz0lEQVR4nJWUfUwTdxjH74/9OeOmZmoWFhdZdDCJk/mCypxmbm44jcAcM+gUZ3Eb6kBCgQRaWquWvlDo0RcOr9AihcK1ULEUCrTrSmEHLRQ4oAqsOrZUbaDDQoct3gKFFvH9yfePXy7PPZ/n933uOQCfixmv13fA8cf468eM1/tM4TgO+DKmpqaoly5nZGZlZGad/zWlslI2+9prMZ6SLwBfy0N9qJBBLLiUUngllU0+Z+loxXHc4/G8YvWhzs72WmVHnaqrUdPVqOmou9leq7T1YrMA71wVBKLsWAPsWgvsCgK2rwIIXwfdud37io7NeL1Q8sVtwR9EbNi4fsXK1QAQsWHjwYidMtrVqYmH8wAFfGXLMuAjIKD4vStHBkwvZczguGd6uvoqo/1mfYeq4auduyLDNrdUyDpUDXUgb8LhCADCVwBhC9XDACByHdAgK/TP/5l2+R66JyflDLapXq3M52YfiWUlnLY0t4x096h4RQEAAuVsWfYE4LP1QJNc+OQHtrT0v06nTqfzTE8rWHmirOyf9u4nxR3jJl0oo9IYxPRSMsU17gR87xvU5dvXAJueD4AgCIZhp9Ppq+57mHDqVGRkJI7jUgqtOJtEijuWtO9z8rdxSi4o4/ElVJprfBzwu6ypFnwZCmxZNquP3wR2BwENMtAPSE1NjYmJEYvFHA5HLBbT6fTk5OQ9c9Xdk5NSCtXS3GKoqhakXJSQcroaNbdQVJkPzlq0cOlZxoBZX8JILssnlhekw7nnBky6xUP2eDxutxvHca1Wm0unF4KFGo3GN+cmUYmMRldyClQ8gYrHV7A5lbTLhqrqwKK53W4+n5/H4UorqkSlZRBcwuNDLpdrsek+3yViMZvFwnG822yGYdg/jxkcfzTXgcfjebSQHADYbDYMwwb6+zGsD8P6rNZBo9E4PDy8ZLYgCErEYn/R7GyS3W5/3jf2xK/CarVaLN2WhcAwTK/XY9jsKs48nl97EARlMtni21gsFjKJ5PNNWavMpefCxbBCXptLz9Vomn2ZAUB7W7vZbEZR1HfQarXd5m5/dxqNRigQPN2sVqsthiAcx6OiDwAAEBT8bkh4yBvLASIpbS7ZOwtwu91FMvCyMItxjcISXaIX5zBE1KsQmSvNheTgwC1s0jWJIIjNZrPb7aOjo46FuHvnztj42MjtEUQjPSuMTyk9cx5OSOTNHhJ58RZb17xFjjFHroQcuid495Gt26LC1oWv2XogdG3I23lSqrofkd2QCgVFEFTE4XCys0koiprNZr1eb/hd/4/D0draivX1IyoxkRebwf/OLyL/6OBd0zzAft8O3+T23Df2PWjLgJKSWMcz4J97H6Dd91rlf5Qq1IhQUATDMARBfD4fwzCz2dRqNLah3dcoOe1txvJymVIlLoTjQSggbtGxu6M9AUBpA6hAkagzJ37IORydsv9ETrRUTUf/rFV0livqq9ksDpPJ5HA4EARhGGaxWAYGb8VGAF8Ev9XX2z80PNJUX8EiRbHIhwLKPmgbWrDIft9e1syTtVV/uG9fOu/HtEICUXiOzo0zYCUKU0Vtg7wgH+RwOHw+//r161ar9a/Rv4sK2bGRy92usZ6eng6TWYVI0o7uTP/+U7/Sju4eNLcuvgFPP9wIIvzjpG+ifok8Tj6kH6zrvqer6SxTqBE2g83OYzOZTARB1Gp1amoqhULN54ny88HmJs0NlbqqWJAQHnJmxya/Tm8N7W8zBADscmrzoNwwfINWRiSJLlAkqY2YQjOggFSs2gb5bzo9giB1dXU1NTVisTh5LjIz088mJlZWVhra20UMZvTq92Lefd+vw+8EmZua5gHT/01L5CJeWT6vrEBULZQoYBECFUlBgZR7rUqo0+sIBMLJkycJBAKNRnM6nb5F84Xb7Z6YcpUwwdhVm+PWfuJX7KrNpvqWwCa/IFwuF4qiRqMRRVEMw575V3A6xiy6tiWamngYACxuakm8tIMX5/wP0leuVhI8ZCgAAAAASUVORK5CYII=" nextheight="1571" nextwidth="2080" class="image-node embed"><p>One of the challenges of RLHF is to gather high-quality and diverse human feedback that can cover a large and complex output space. One possible solution is to leverage the power of the community by crowdsourcing human feedback from multiple sources and aggregating them in a principled way. For example, Hugging Face has created a platform that allows users to provide feedback on language models and their outputs, and to share their feedback with other users. This can create a virtuous cycle of data collection and model improvement, as well as foster collaboration and trust among the community.</p><p>RLHF is effective in various domains of natural language processing, such as text summarization, natural language understanding, and conversational agents. RLHF has also been applied to other areas, such as video game bots and image generation. RLHF has enabled AI systems to generate more diverse, creative, and human-aligned outputs and to overcome some of the limitations of traditional reinforcement learning methods</p><p>Decentralization can offer some solutions to these challenges by leveraging the power of distributed networks, communities, and markets.</p><p>One example of a decentralized concept that can be used for RLHF is decentralized autonomous organizations (DAOs), which are self-governing entities that operate on a blockchain. DAOs can coordinate the actions and incentives of multiple stakeholders, such as AI developers, human feedback providers, and end-users, without the need for a central authority or intermediary. DAOs can also enforce smart contracts that specify the rules and rewards for RLHF, such as how much feedback is required, how it is verified, and how it is compensated. This is a more efficient and economically sustainable way of crowdsourcing feedback.</p><p>Another example of a decentralized concept that can be used for RLHF is information markets, which are platforms that allow participants to trade on the outcomes of events or questions. Information markets can elicit and aggregate the collective wisdom and opinions of a large and diverse crowd, which can provide valuable feedback for AI agents. Information markets can also incentivize honest and accurate feedback by rewarding those who predict correctly and penalizing those who predict wrongly.</p><p>Decentralized communities, DAOs, information markets, and other decentralized concepts can be used for RLHF and crowdsourcing purposes by providing a scalable, efficient, and reliable way of collecting and aggregating human feedback for AI agents. Decentralization can also foster a more collaborative and participatory approach to AI development, where humans and machines can learn from each other and co-create value.</p><h1>Verifiable Models and ZKML</h1><p>How can we ensure that the data used for ML is not leaked or tampered with? How can we verify that the ML models are trained and executed correctly and honestly? How can we protect the intellectual property and ownership rights of ML developers and users?</p><p>One possible solution is to combine ML with zero-knowledge proofs (ZKPs) and blockchain. ZKPs are a cryptographic technique that allows one party to prove to another that a statement is true without revealing any information beyond the validity of the statement. Blockchain is a distributed ledger that records transactions in a secure and transparent way without relying on a central authority.</p><p>By using ZKPs and blockchain, verifiable machine learning, a paradigm that enables the verification of the correctness and privacy of ML processes and outcomes, can be created. Verifiable machine learning has several benefits, such as:</p><ul><li><p>ZKPs can be used to prove that the ML models are trained and executed correctly without revealing the data or the model parameters. For example,&nbsp;<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://arxiv.org/abs/2304.05590">Zero-Knowledge Proof-based Federated Learning (ZKP-FL)</a>&nbsp;is a scheme that leverages ZKPs for both the computation of local data and the aggregation of local model parameters, aiming to verify the computation process without requiring the plaintext of the local data. Blockchain can be used to store the proofs and the results of the ML models, creating an immutable and auditable record that can be checked by anyone.</p></li></ul><img src="https://storage.googleapis.com/papyrus_images/531b6909992e7c9b6a71aca28772ad5a.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAASCAIAAAC1qksFAAAACXBIWXMAABYlAAAWJQFJUiTwAAAFUElEQVR4nH2UfUwTZxzHb8n+W5aYbP/tLXFxiW7Zlsw/ZrI//GfRvbjXxJmoGzrnfEkQcGCok1ZeWpCiHBTQQqzYRaCAQKkXyglX4Spcax/GcVCkB7jWlpej9Up7R3u39lngCKjL9v3ncs9z93x+z+/7fR4E/rdS6VQqnYYQlnbYL9v6IISPuUgqnY48WTY3WZs7cMO1W8LKSkKWw0tLqsY2A+aIxWN3MFxbVoNWmx75QxBC5H8AEMKEJK+Iy8jbnyAffA4hXOCXf7N0mzrtJ09k5ubmfP3todvt7ebWtsH+PuTjPVsPn/LPsBeLyjCsj3Q+YJiHzwBCoZDZfBPDsPb2dgzDCIIYGBzwz8ycUGkQ5GXko88SkszOPvq+QHeuqOzkibP5qguHj5zBbLbMQp1tcAgMO/vs9kWeByPjYGQcv0t6J6c3AaIochxH06MDJKlWq3Ecr6ioLCzSSZL0JwB5OVk3TCZJkkeYccxmGxp24f1DPbizy9oXDM31Ew6Xx8PH4kIimZBl0ukGI+N3CecmQJJkaVUyhJDneQAAxy18t//Ui1t2LHLhNTNWnUil01FxJSKI4WgsKopxSRISSY6PRQTxiSAmJJmPxVPptNtN08ykwzG0DhBFcfWBvLV1++5YLOJwODiOk6Tk9vd2IciWzq5uCOHPNTdP1TcrJkuSFI7wBuOt6+aOotJaMZFIpdPLkchPl43qZluUjzSaWyrRmvp6MzvtX105kUyuAVaVSgkU5QIAdHV1ZWdn5efna7XaAdyOvLYT2fGpYvKPDU1Vf7QfzTiel5v91TcHb5hMNaabPRiG7N734bEz/pmZRrPll+OnCov1gcDCeosycy4oAIqieJ7fsF1pWjN+D0FeRXbuFRJJ75TvoKowr6AoK0ul0RQfPpLlGR4+V9vQ7XT5mFHnAPF4fp6dDpTr0dKyyr/8c+sAr3fi9OnjBkP1IrdEkiRN05Ik/S3LgiBACCcYxlBRZm1vFUXROz0LgIednfWMTgLaN+SmA8GQzY4P3B/eMNk7OVNWdllbemUdIEkShBDH73IcpxTOsixBEMqrtDYbecLHhVWrJEkecrmsNpvtzh1LW5u9t7fJYmlta7/d0dXS2ubyeCCEPjag0ZSoNaWbLVpbdMpQ24ggiLG+SckSRVHKViCEAIBQKLSBd7mGK/Tl51XnLmrUeXm5OTnZGIaNM4zyDRgZJxz3rdbe9RQpAZWSsRde2oYgyCuv74QQKvWyLEuSpN/vjwuikpaoKMqplJxKDblpS0evvY9cjgtRUUzIspBI8jEhlU5TrtFnYqrUSFGUWq3es3efXq9XWq8oHOHHnIPI7h+QL45CCIMRfiI4fx/Q+QWleP9Ankq7yIUf+nyMm0I+PfDm2eJoZPGCuqS72+50ehhmahMAAPB6J0KhIMv6lJGN8zUVWkC2vIu8v3eVF429c748s6j8WMavRzMOfblvf8PVqypDfTOGF2gv/a67ND07g+P3PIBx3HM9swMAwEbhSjohTCtTLZ3dyBvbdh3IWI4Lw2DEUFfX1GSpqjVX195QF1fTNK2pqu3E+4OBgN/vj4qi20339DhaLLbnAU+fgKclCPFQ4FFkaVGSpKi4osQxuBAOzHPBhfDcUvihj2VnZ9fujxUhkRRFsbKq4brJsn5HPLcDgiBwHKdpmiRJAABJkl6vd8I7OTbGKG2kaZqiqHGGcTpJZozm+ei/a/J6vRupWwcovwEAjMZrOp2upqYGRVG9Xo+iaF1tHYqiVyor9Xq90Wg0GAwl2hIURXW6UhRFCYJ44PG4PQ/ApkasVuvc3LwC+AcrLkIpROCRWAAAAABJRU5ErkJggg==" nextheight="970" nextwidth="1714" class="image-node embed"><ul><li><p>ZKPs can be used to protect the privacy of the data and the model parameters by hiding them from the verifier and the public. ZKPs can enable shielded transactions, viewable only to those who possess the decryption key, for ML models. Blockchain can be used to enforce access control and encryption policies, ensuring that only authorized parties can access the data and the model parameters.&nbsp;(<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://a16zcrypto.com/posts/article/checks-and-balances-machine-learning-and-zero-knowledge-proofs/">ref</a>)</p></li><li><p>ZKPs can be used to prove the ownership and provenance of the data and the model parameters by linking them to digital signatures or identities. Blockchain can be used to create smart contracts and tokens that represent the rights and rewards of the data and the model owners, enabling a fair and decentralized marketplace for ML.</p></li></ul><figure float="none" data-type="figure" class="img-center" style="max-width: null;"><img src="https://storage.googleapis.com/papyrus_images/eb8877f5a9aec4cd9c4ea08cd05e6dab.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><h2>Giza</h2><p>An example project working on model verifiability is Giza Tech. Giza is a company that leverages blockchain technology and zero-knowledge cryptography to enable verifiable AI model inference on-chain. They allow anyone to deploy their AI models in a serverless and secure manner and to use them in smart contracts without revealing the model details or the input data. This way, users can benefit from the power and utility of AI models while also having the assurance that the models are running correctly and producing valid outputs.</p><img src="https://storage.googleapis.com/papyrus_images/134cb8856d1d38c23b7b1b074d5aacaf.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAWCAIAAAAuOwkTAAAACXBIWXMAAAsTAAALEwEAmpwYAAADyklEQVR4nKWVz2seRRjH95Ye/BdEYYU3l1xetO9lm8N78L3k9FJ4MRA87OkVeW8vyJ4WhD1tvezBpYJjiZNAJmmdFjMVmzm0I9buxYHagSATxU4VR2rcUhjbw0j2sdv0bdIG/RyWd4ed5zvzfX68gT8Bzjnv/fra50HDZwilT8iyLM9zhFBd10fuDU4usLz8Lgi0Ky1CCGPM/xXY3r4aBEH46huUUkII55xSWhSF1loIobX+7wKAEMJ7L6VkDdPpVAhBCBENSin/AgH3LO2KP4QQoq4fQCyE0NzcHELIWqu1ppS+ROAkMMa895zzLMuqqoInQggsklIeK+CcS9O02+0uLS2NRqMoiiaTSRzH4/G4vZ/3vixLxljVQCntdOallO3rS3IAbmKMOecYYykl/HCHXHLOGWNsQ13XnPPRaJTn+fNF9SKLfr1375e75u5vB84aY2SDeoJ+grWWEDIcDvM8N8ZordtvlFKwBfY+FYBTXPziyocfvPf+cnRn98dbt76Dj6SUEPfwZkppXdfW2jZo61hVVVALlNJnbvDob/fHw4dtNz1+9LjVVkoJIZxzVVVhjIUQGGPGWFmW4L5zDmNMKWWMaa0554wxY8ysRddv3Pj+5s2r61sz6Q3DsNOZX1lZiaLIOTcej+M4LssSTu2cG41Gnc78wsLCYDCIoiiOY0KIUmpWgBDy5/6+vX9fKWWtBX+ttc45xpgQotvtGmPGDXAJaOk0TZMkOYgYBAsLC5RSKJZZAYzxD7dvv3P27KlTrwwGbw8awjDs9XpJkkA3IIQYY7QhyzKlFMYY0g7uQxEihI4QWF9b27hIgyAQ33y1t/cz5FZKaYxJkoQxdmRFlmXpnIMKPvxBVVWzAlub5GOEgyC4/vWn+389sPZ30+C9hxNJKcuyJISACYwxa21ZlhDOWltVFW+A0grAYqUUDHSEkBBie/vLb6uqrb+iKJRSEHQ6nRpjOOdFUdR1DUZBh0dR1Ov1ptNpHMeTyeRfiyaTyXg87vV6aZoqpRbPnHnzrdOXNjaWFhejKEqSJI5jrXW/3w/DsNvthmEI5+j3+957rXUURUEQ9BqyLIMqz7KsLMuDnDvn6rqGp/f+o3Pn7uzunn7t9SAILm1tXdvZuXL5spQyjuPhcAjjpLUIIQSB0jSVUiKEYBQWRcE5z/P86DIVQlzb2Vm9cGFrc3N1dRVjTAiBNFRVVT9H2zHwCnMe1o9IMudca/3T3h7kFgoDhr5S6rj/xeOGnRBiVkAIgRDCGH9y/jxCqGxAh4CV4ygaEEKUUmiFfwDuS4ARrDQ0FgAAAABJRU5ErkJggg==" nextheight="1403" nextwidth="2069" class="image-node embed"><p>Giza Tech’s solution is based on two key components: a transpiler and a verifier. The transpiler converts any AI model in the common ONNX format into a verifiable model that can be executed on-chain. The verifier uses zero-knowledge proofs to check that the model inference has been done correctly off-chain and to provide a proof of validity that can be verified by anyone on-chain. The verifier also ensures that the model and the input data are kept private and encrypted and that only the output and the proof are revealed. The network is designed to be scalable, interoperable, and compatible with various blockchain platforms and protocols.</p><h2>Modulus</h2><p>Modulus Labs is another company that aims to make machine learning (ML) models verifiable. They address this issue by using zero-knowledge proofs. Modulus Labs constructs and trains ML models utilizing their own data and code, converting them into zero-knowledge (ZK) circuits. These circuits are mathematical depictions of the model's logic and parameters.</p><p>They publish these ZK circuits on a public blockchain, accompanied by the model's metadata, such as its name, description, and input/output format. Modulus Labs also offers an API that allows dApp developers to access their ML models. They do this by sending queries and receiving responses via smart contracts.</p><img src="https://storage.googleapis.com/papyrus_images/c956929ad5fad0e72bdea2eb884d9537.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAJCAIAAADcu7ldAAAACXBIWXMAAAsTAAALEwEAmpwYAAACbElEQVR4nF2RzUrrQBSAR8QqWVRwI7juRhGlbgQXUtGV6K7G5+gzdONeQUgX11cQd24ElyqkKmljQmiStulkMmNMmhJi0pxLE8yV+zEcDgPf+ZlBcRz3+33jB9MwZzEjDEP4xXQ6tW07juPflwDg+76u66Y5Ewt0XccYT6dTpGmaqqjUoTa2i+MQR9f119dX3/eDIPA8L03T+/v7lZWVu7s7APA8L8j4/PwURRFjTAjBGNsZGGPGWLvdJoQgWZYnk8l/QwFAHMeiKFqW5bouY8zzvKenp9PT02azads2IYRS6rquYRidTgcA0jQt3Dw3TdOyLKSqqmEYiqJ0Oh1Jkrrd7sfHx/v7u67rsiwDQJIkURTlDsZYVVXGWFElDENJkvKk1+spiiLL8tfXFwD0er1ZA1mWgyDICyVJkmYAQBRFualpmiRJrVbr6urq4uKi1WoJgnB5efn8/AwA4/G42CCKoslkMh6Pv7+//23Q7XY1TWOMEUIcxyEZlNJ+vy+KIgAMh8PDw0OEUKlU4jiulIEQqtVq+RwvLy+EENu2HcehGYQQ13Xb7fasAaWU5/l6vc7/cH5+zvP8ycnJ29sbADDGlpeXy+VytVqt1Wq7u7s7Ozvz8/N7e3u+7wPA7e3t0dFR7p5l8Dx/fHzcaDSSJEGj0Whtba1SqVSr1a2trfX19c3NzY2NjXK5/Pj4CACu666urs7NzS0tLS0sLCwuLnIchxDa39+nlIZheH19jRCqVCrb29uFznHcwcFBHMcoSZLRaDQYDCzLGg6Hg4w8Kf724eFBEISbm5s/PwiCkD8gAARBUCi/o+M4APAXljGLickZO/IAAAAASUVORK5CYII=" nextheight="407" nextwidth="1400" class="image-node embed"><p>They employ a specialized ZK prover, named Remainder, to create zero-knowledge proofs for each query. These proofs are also published on the blockchain. They confirm the model's output was correctly calculated from the input, based on the model's logic and parameters, without disclosing any information about the model itself.</p><p>Modulus Labs allows anyone to verify the ZK-proofs on the blockchain using a ZK verifier, which can be implemented in any smart contract language. The ZK verifier checks the validity of the ZK proof and returns true or false. This enables a new paradigm of decentralized and verifiable AI, where ML models can be verified and shared across a distributed network of nodes without compromising their security or performance.</p><h1>Conclusion</h1><p>Decentralization can address challenges in AI model building, including the need for large data and compute resources, privacy and security concerns, and model verification. Many of the problems that we covered haven’t been able to find sufficient solutions yet, which shows both the great potential of the possible solutions and how tricky the situation is in the first place. As a core part of the AI process, model building is a step that deserves attention in terms of the decentralized AI landscape and AI discussions in general.</p><p>In the next article, we will explore how retrieval augmented generation can complement model building and the decentralized solutions that can further increase RAG’s potential impact.</p><h1>References</h1><ul><li><p><a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://arxiv.org/pdf/2304.05590.pdf">https://arxiv.org/pdf/2304.05590.pdf</a></p></li><li><p><a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://arxiv.org/pdf/2310.14848.pdf">https://arxiv.org/pdf/2310.14848.pdf</a></p></li><li><p>A decade of exponential growth in raining FLOPs. Sevilla et al (2022).</p></li><li><p><a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://www.datawallet.com/crypto/what-is-bittensor">https://www.datawallet.com/crypto/what-is-bittensor</a></p></li></ul><div><div class="callout-base callout-info" data-node-view-wrapper="" style="white-space:normal"><img src="https://paragraph.xyz/editor/callout/information-icon.png" class="callout-button"><div class="callout-content"><div><p>This article was originally posted on the <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.firstbatch.xyz/blog/">FirstBatch blog</a>. It was reviewed by <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/batuhan-aktas-38692b148/">Batuhan</a> and <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/kerim-kaya-552878129/">Kerim</a>, edited by <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/ilkyazyesilserit/">İlkyaz</a>.</p></div></div></div></div><p></p>]]></content:encoded>
            <author>lab@newsletter.paragraph.com (omer)</author>
            <category>ai</category>
            <category>blockchain</category>
            <category>web3</category>
            <category>bittensor</category>
            <enclosure url="https://storage.googleapis.com/papyrus_images/09a6f18d88f0a18020bcb44f612f39d1.png" length="0" type="image/png"/>
        </item>
        <item>
            <title><![CDATA[TOWARD DECENTRALIZED AI, PART 1: DATA COLLECTION]]></title>
            <link>https://paragraph.com/@lab/deai-data-collection</link>
            <guid>FftxEhHfPQ2aVT9EiHsY</guid>
            <pubDate>Tue, 30 Jan 2024 00:00:00 GMT</pubDate>
            <description><![CDATA[This is the first installment of a five-article series on where and why decentralization is needed in the AI pipelines. We are starting with data collection.]]></description>
            <content:encoded><![CDATA[<p>This is the first installment of a five-article series on where and why decentralization is needed in the AI pipelines. We will not only cover one specific AI task but include a variety of AI / ML-related areas, from language processing and vision to robotics. Instead of starting with decentralization and looking for where we can fit it into the AI processes, we will be looking at where AI practice needs decentralization and focus on how that integration can be designed.</p><p style="text-align: start">When talking about AI, especially the decentralization of it, people tend to think solely about the model development stage. This leads to an unbalanced focus on the training itself and the compute supply problems that arise from it while ignoring other areas whose importance may be greater or equal to that of model development. </p><figure float="none" width="483px" data-type="figure" class="img-center" style="max-width: 483px;"><img src="https://storage.googleapis.com/papyrus_images/e98a7dc841bbc22e8db122d1f667b0ad.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="">Overview of LLM Challenges. Designing LLMs relates to decisions taken before deployment. Behavioral challenges occur during deployment. Science challenges hinder academic progress.</figcaption></figure><p>In this series, we will have five articles covering different stages of the AI pipeline and the safety concerns that relate to all of those stages:</p><ol><li><p>Data Collection: Quality, Copyrights &amp; Ownership (you’re here!)</p></li><li><p>Model Building: Compute, Open Sourcing, Fine Tuning</p></li><li><p>RAG: Bringing in World Knowledge, Collective AI Memory</p></li><li><p>Safety: Manipulation, Deepfakes</p></li><li><p>Applications: AI - Human Interaction</p></li></ol><figure float="none" data-type="figure" class="img-center" style="max-width: null;"><img src="https://storage.googleapis.com/papyrus_images/0b58fd3c36ff81bc7d3e5e8e270f9f60.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><h1 style="text-align: start"><strong>Data Collection Overview</strong></h1><p style="text-align: start">In the AI development process, data collection is the step in which relevant data is gathered and organized according to the identified goals and objectives. Of course, the type of data and how it is collected heavily depends on the specific use cases (you probably won’t be collecting a lot of images for a text-only model), but some general patterns emerge across all domains.</p><p style="text-align: start">Data collection can include getting existing data from open or proprietary sources, generating new data specifically for the task at hand, augmenting low-quality/volume data with various techniques, and data labeling, among other techniques.</p><img src="https://storage.googleapis.com/papyrus_images/da05d7ae1e852686fe412e5a3ff00d74.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAJCAIAAADcu7ldAAAACXBIWXMAABYlAAAWJQFJUiTwAAACEElEQVR4nG2RIWzjMBSGM1ZWMjhQ6QoCSgpKDAJCBkoi7QpGJt0i3QyKQ0JCSkJChkoCvAMlIUEmJiZGISYPGb0DJkZGIzmdvVW60z70bD3/z/peorUuimKxWNzc3KzX677vGWPDMDw+PiaBH8/P4zhWVZUkye3tbdd1hJDlcnl3d7def4v9jDHOOSEk5pRlOY4jY+xyuSTe+2maOOcygAHnnDFGSimEMMY45wBASqmUMsZM06SUklJyzhERALTW8U8P3x/yPKeUTp8k87+8v7/P86yUEkJora8p8zxDILZZa4UQsfbeG2OqAA28vPzs+x7xNyJ+PUBK2bbt5XKpqqppmpilA7ENEYdhiLVzDhEppU9PT0VRHALn8xkRjTGJc+7t7RfnPKrXWkchzjlrrQtELSYQxwAAIk7TFI9SyuPxmGVZnue73S7LsrqulVJ/FXnvX19fh2E4n89d13HOowGt9TXOOee9j+vxAWstAFzrcRz3+/3hcNjv94SQWAshAOBrRc45SmnXdVVVlWUZ1QshxnGMbQDQdd1VlzHmdDoRQqKfNE0ZYwCglPp/QEyPcowx1tr4PopGxOuN/QQRtdZ1XZdlSSltmqYoCkrpxw4AgHMuApxzpVTTNKvV6v7+PsuyNE232+1mszmdTnVdU0qllIwxQshqtWrbVggxDMN2u93tdnmebzabLBBFcc7/AEiFUfgTfk90AAAAAElFTkSuQmCC" nextheight="292" nextwidth="1020" class="image-node embed"><p>Data collection and quality challenges cannot be solved in a single phase, but they require continuous attention throughout the entire process (ref). For this reason, we will keep coming back to the data issues in the following articles of the Towards Decentralized AI series. Still, most of the foundation of how data collection and decentralization relate to each other will be laid out here.</p><h1 style="text-align: start"><strong>How Decentralization Can Solve Data Collection Problems</strong></h1><p>In order to explore the intersections of decentralization and data collection, first, we will look at the current challenges that AI teams/developers face during the data collection process. One of the most well-known ones today has become data volume, as the training data needed for generative models recently started, leading to questions about limited data supply, but the quality of that data is also a major factor for obvious reasons. Storage issues, privacy controls, and copyright are other challenges that have key roles in AI pipelines, and all have possible solutions that lie within decentralization.</p><h2 style="text-align: start"><strong>Data Quality</strong></h2><p style="text-align: start">With the ever-growing need for larger datasets for training, AI developers face a major challenge in both finding high-quality data to use and checking the quality of the data they collect. At current rates, it’s getting impossible to reliably check the quality of their training datasets (<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://www.springboard.com/blog/data-science/machine-learning-gpt-3-open-ai/"><strong><u>even GPT-3 had &gt;45TB training data</u></strong></a>), and relying on automated methods can lead to inaccuracies and poor data quality.</p><figure float="none" data-type="figure" class="img-center" style="max-width: null;"><img src="https://storage.googleapis.com/papyrus_images/a711cb206fa0308e4dfc5c4fb48088c7.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="">Refinitiv Artificial Intelligence / Machine Learning Global Study</figcaption></figure><p>Some major problems in terms of data quality stem from the repeated, duplicate or outdated data that is collected. While these problems can be overcome with model editing and retrieval augmentation methods, the edited versions and augmentation datasets still require robust data collection and quality check processes. Here’s an example from the 2023 paper “Challenges and Applications of Large Language Models” (Jean Kaddour et al.) going over an example of outdated training data:</p><img src="https://storage.googleapis.com/papyrus_images/adf688b9bebc323059d8d7c5f88cc14a.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAHCAIAAADmsdgtAAAACXBIWXMAAAsTAAALEwEAmpwYAAACTUlEQVR4nE3OQWjTUBgH8HcQPOhhshUmg0JhJzPBh5ps8DrMFCOFqLOb0Kn0lFPRw5CZXYJzq5AOWhq69jnLUuhhUFIPAdki2ubS5rbTytj2BGluOThy6Xp7sgSGP/68w8f38X8AY2xZlhnAGCOERFFMp9OCIEiSxPO8IAg8z0uSFK5ZAdM0dV3HGNd0vabrzWZT12sIIU3THKdj2y07oOs62NzcTKfTGxsbiqLIsiwIgiiKqVQKIcQHRFEMWzVNUxQFY+y6LqXUNE2GYThu+s499s2rJcMwMMa5nNrpdn/8ajWbhuM4rusCnucBACzLxmIxlmWFAMuyEEJRFCGE8XhclmVN02zbPjg4IIR4nkcpxRcqjuMkU68BAP1+X83lksnkp/X1W/A+AFcPDw//9PtAUZRYLPZwbk4QBBi4PTWFEJrmuNn4LEIIQri4sPA5m5Vl2TAMSulwOKSUGoahaRrGlfpug0u8pJR2ut2w+4tej07NUEoJIaDRaNi23Wq3W+323v7+z1br+97e0fGx7/t/z87CN4zneeHE930v8H7lw41r18dHRxNPHjMMAyFkGGZ5eZmQ3+EnLgqmOe7m2NhEJDI+MjIZjU5GoxORyPNEoloqXaZSKNS2t6ul0rdGo9Pthscnp6cAXFl6sYjg3cy7t5IkZTIZSZJWV1cTz+YfJOYppb1eD6x9XNspl2uVyk65fJlqqVRU1aKqbuXzRVUtZLNb+Xwhm92t1y3LIoS4rksIOTo59X1/+J/BYEApLX+tzTx6en4+6PV6/wDWqHYSyCRIHgAAAABJRU5ErkJggg==" nextheight="536" nextwidth="2628" class="image-node embed"><h3 style="text-align: start"><strong>Community Curation &amp; Review via Open Data Hubs</strong></h3><p style="text-align: start">Over the years, open data hubs and data-sharing platforms proved their usefulness in the data science and machine learning world. Hugging Face and Kaggle are among the prime examples, with Hugging Face being more focused on training data and open models and Kaggle being home to competition-style predictive competitions and related datasets/notebooks. Not limiting it to ML-focused platforms, even human-curated knowledge sources like Wikipedia play a crucial role in the open knowledge movement. These play a huge part in creating available sources for training, fine-tuning, and retrieval.</p><p style="text-align: start">By creating decentralized open data hubs, a large amount of people can curate, review, and rate the data that will be used for AI tasks. This shifts the burden of ensuring the quality of datasets from relatively small teams and masses of people while the contribution and rating processes run transparently without any single party’s manipulation or intervention. With the right incentive structures, these decentralized data hubs and data curation/review activities can ensure the quality of data coming in with high velocity and volume.</p><p style="text-align: start">These decentralization solutions can apply to newly emerging data hubs as well as existing ones.</p><h3 style="text-align: start"><strong>Information Markets</strong></h3><p style="text-align: start">Another type of structure is information markets, where a dynamic and competitive environment ensures the most up-to-date and high-quality data sources in critical areas. These types of markets can be set up using smart contracts to ensure fairness and provide the necessary incentives to the participating parties. People from a wide range of backgrounds, from domain experts to everyday internet users, can participate in these markets in different capacities, contributing, reviewing, and monetizing data.</p><p style="text-align: start">Ocean Protocol has been working on this for a long time, creating markets where users share data not for web3 use cases only but also for AI and data science applications.</p><img src="https://storage.googleapis.com/papyrus_images/975f1903dbaab1a7a52db71243e3c5d2.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAVCAIAAACor3u9AAAACXBIWXMAABYlAAAWJQFJUiTwAAADZElEQVR4nJ2UTW/jRBjHmwTFLzhjezIznhfbGXv8vglseugCF0ppuXbV/QB72pWQ6AE+AHtE3BAfoZf0AOrHoFz7LXqIVKk9VgqynaQLbBN1fxrZ84ys5/+8jXfMBsMwIIScc8/zMMaO4wRBoJRijFmWZT6OVS+rXY31r3PTNHdqw7I0TZNSnpycvDp5dXh4eHBwUJXl0dHR3t6e67q6rrdBWCsGg4FlWbppaIbW1/uarmm6rumavkLTl+e1gGmauq5jjMuyTBqyLEuSJI7j1qyqKs/zoiiUUpRSLgQXPOSCDQizWUhGkRdxyillHvIwRBgi36UZij5jxUrAMAAACCFCiG3bruMIIXADdJcghAAAhm5AC4QDL7GDnGc5VSkaZVgWQ1UMkwnNn7GsHKYxHDGHYkhrAcM0bQuwIQ6DgDG2O51+8eJFWRSRlL4QcjRKk2S9z9K0TPM4kFLKLE9VEgdhk9AozLNUypEQPAh8YIO+1td1fccwDQxs+unQh0KpOPD9/f39746OJuNxnmWRlGmSTJ8/b/USpZ5V1WQ8jiIZR7Isit3ptJXP0nR3Oi3yPJJSxTF0oWkYdZMJsKkFkYEQwh71aggRnDNKPUJoAyHk/T3BePVhTbOtDYzx2qyLuRQwa+/1OaUYYwShL0dJmUPPcxBEEDqO47ouhMte1IFgjBrw4zwIRFT5foCb3hJMHAR3PfVafPW9/PpLlvFR8PlkUlWVUmo8rvI8o7SN/oEtArbjDAYDo5kiSmjXsX6JXy7kbwv/91/Dl6bA46Ke0TRN22fdXCnjOF5vPqjxIGDo9Wsl4GkueKu++Uv9eKl++kF9a3ouo6ytPqWUMbYu/ZotGax/FQCAtksDG3SMfueTbmen0+12e71ed0Wn0wEAPOZ0uwClVNO0t2/e/H15efHHn+ez89lsdnZ2dt4wm80uLi6Oj48ty1qX5WkZMMZ6vd7P794tFov5fH53d3d9fX11dTWfz29vb29ubu7v709PT/v9PqUUIcSFn+Yl5+L/E/VhAUIIQojxmrboXsO6B+09eN/R0zIghEAIpZRhGGK8vDKkUV2ztfpbmlxfosY3qxMRXAjGOHkK20sUBEGWZUKIJ4X8HzYJ+L6fJIkQYvOk441sLBFC9c/nY2PfLrBhNj5C4B+5qdqD3yvqAAAAAABJRU5ErkJggg==" nextheight="1396" nextwidth="2176" class="image-node embed"><h2 style="text-align: start"><strong>Data Volume</strong></h2><p style="text-align: start">Despite the major advancements in LLMs, the amount of data used for training these AI models is still tiny compared to the amount of input a human receives:</p><div data-type="twitter" tweetid="1750614681209983231"> 
  <div class="twitter-embed embed">
    <div class="twitter-header">
        <div style="display:flex">
          <a target="_blank" href="https://twitter.com/ylecun">
              <img alt="User Avatar" class="twitter-avatar" src="https://pbs.twimg.com/profile_images/1483577865056702469/rWA-3_T7_normal.jpg">
            </a>
            <div style="margin-left:4px;margin-right:auto;line-height:1.2;">
              <a target="_blank" href="https://twitter.com/ylecun" class="twitter-displayname">Yann LeCun</a>
              <p><a target="_blank" href="https://twitter.com/ylecun" class="twitter-username">@ylecun</a></p>
    
            </div>
            <a href="https://twitter.com/ylecun/status/1750614681209983231" target="_blank">
              <img alt="Twitter Logo" class="twitter-logo" src="https://paragraph.xyz/editor/twitter/logo.png">
            </a>
          </div>
        </div>
      
    <div class="twitter-body">
      I've made that point before:<br>- LLM: 1E13 tokens x 0.75 word/token x 2 bytes/token = 1E13 bytes.<br>- 4 year old child: 16k wake hours x 3600 s/hour x 1E6 optical nerve fibers x 2 eyes x 10 bytes/s = 1E15 bytes.<br><br>In 4 years, a child has seen 50 times more data than the biggest LLMs.…
      
      
      <div class="twitter-quoted">
        
  <div class="twitter-quoted twitter-embed">
    <div class="twitter-header">
        <div style="display:flex">
          <a target="_blank" href="https://twitter.com/tomosman">
              <img alt="User Avatar" class="twitter-avatar" src="https://pbs.twimg.com/profile_images/1592517044578226176/19SSiHyW_normal.jpg">
            </a>
            <div style="margin-left:4px;margin-right:auto;line-height:1.2;">
              <a target="_blank" href="https://twitter.com/tomosman" class="twitter-displayname">Tom Osman</a>
              <p><a target="_blank" href="https://twitter.com/tomosman" class="twitter-username">@tomosman</a></p>
    
            </div>
            <a href="https://twitter.com/tomosman/status/1750461836078764367" target="_blank">
              <img alt="Twitter Logo" class="twitter-logo" src="https://paragraph.xyz/editor/twitter/logo.png">
            </a>
          </div>
        </div>
      
    <div class="twitter-body">
      "A 4-year-old child has seen 50x more information than the biggest LLMs that we have." - <a class="twitter-content-link" href="https://twitter.com/ylecun" target="_blank">@ylecun</a> <br><br>20mb per second through the optical nerve for 16k wake hours <img class="twitter-emoji" draggable="false" alt="🤯" src="https://twemoji.maxcdn.com/v/14.0.2/72x72/1f92f.png"><br><br>LLMs may have consumed all available text, but when it comes to other sensory inputs...they haven't even started. 
      <div class="twitter-media">
      <img class="twitter-image" src="https://pbs.twimg.com/ext_tw_video_thumb/1750460045819600896/pu/img/OEDqQGjGPVZHtC7D.jpg"> 
    </div>
      
       
    </div>
    
  </div> 
   
    </div> 
    </div>
    
     <div class="twitter-footer">
          <a target="_blank" href="https://twitter.com/ylecun/status/1750614681209983231" style="margin-right:16px; display:flex;">
            <img alt="Like Icon" class="twitter-heart" src="https://paragraph.xyz/editor/twitter/heart.png">
            4,362
          </a>
          <a target="_blank" href="https://twitter.com/ylecun/status/1750614681209983231"><p>11:20 PM • Jan 25, 2024</p></a>
        </div>
    
  </div> 
  </div><p>When it comes to the sheer volume of data being collected, providing social and economic incentives to contributing teams and individuals can speed up the process by bringing in new types of data that are not widely available yet or pushing more data sources to open up with the new monetization frameworks. Right now, even though there are TBs of data sources publicly available on the internet, this is concentrated in certain data formats (mostly text-based) and certain sectors/verticals.</p><p style="text-align: start">Of course, the volume itself is not meaningful if the data quality is lacking, but when coupled with the curation and review processes mentioned in the previous section, it can be accelerated much more confidently, knowing the newly added volume will always be going through the necessary controls.</p><p style="text-align: start">This also applies to the synthetic data. While the overcrowding of the training datasets with machine-generated data and differentiating it from human-generated ones pose a big challenge, even relying on these iterated processes of collective reviews and feedback can make a great difference in the effect of the synthetic data, essentially making it a much more useful part of the data collection process instead of a bottleneck.</p><h2 style="text-align: start"><strong>Data Storage and Maintenance</strong></h2><p style="text-align: start">When we talk about these massive datasets in different shapes and sizes, another key aspect that teams/developers need to deal with is the storage and maintenance of this data. When you are collecting large datasets, there should be a reliable storage method that ensures you will not lose, damage, or change the data unintentionally. The simplest risk involved is the monthly subscription fees for data storage, where the initial timeline of a project can limit how long these storage units will be maintained, although the collected data can be needed for much longer for different purposes. Similar issues could arise with on-site storage units with their dedicated hardware, as the subsequent projects will also need storage.</p><p style="text-align: start">Another issue is the cost involved in the data storage. As the collected data grows, monthly payments for storage grow cumulatively, meaning that if you collect 5TB of data every month, what you will pay is not a flat fee, but instead, it will always be more than what you paid last month. Of course, paying for more than you need upfront and getting a discounted deal is always an option, but that means there will always be a portion of storage units that you are not utilizing, which leads to inefficiency both technically and financially.</p><p style="text-align: start">Since many of the general use models have large overlaps in the data sources they use, such as Wikipedia pages, books, and crawled web page content, they collect and store mostly the same data in different storage solutions over and over, paying unnecessary fees that add up to millions of dollars for something that’s already paid for. The lack of coordination and collaboration is a serious setback against the efficiency of the system.</p><img src="https://storage.googleapis.com/papyrus_images/e3959f04e306d25bf72eaa09457fec07.png" blurdataurl="data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAACAAAAAPCAIAAAAK4lpAAAAACXBIWXMAABYlAAAWJQFJUiTwAAACoUlEQVR4nI1Ur28bMRg9ULJ1f8OaSZUqVVpJedlaKdJKAiqtBZkUkIIDpZHWAzc1AyetoEdMDgSkAwcspUcMQtxJJpZmcmBGBkdMLHUmHvim5WuT/lYfOCn+Pn/P7/nFEcwQXgC4aYOn8bAaPdP96P5wM+Ie8VNbrgmyLIvjOEmSNE0JIcPhsCiKfv+QUtrr9ZIk6Xa7QggA8N7PJ1JK+/3DwWCQzNDr9eI4ds49QiCEYIxNp1MxA2NMCME5r+t6PB5PZ9BaI4Fzzlr7S6nfWltrpZR5nmutjTGcc2ut9x6/C4J72p+xyDkXQiCEtNttQghjLE2/tlrvsixjjEkpv5+e3tZ6TbC8vLy1tQUAg8EgiqKmaQBgfX19d/cjABBCoihSSiEH53x1dXVtbe3V6zdpmr7f2Njc3Ox0Om9XVlqt1t7e3oft7fPzHwuLvPdRFC0tLQHAzs7OfFY0AwB87najKCqKApu11pTST/v7eZ6fnJx0Op0vx8dnZ2dHR0ftdnv4H98mFxfoxLWC0WjEGAOAuq7zPMcaug8AxpiiKHAR78A5p5TyM+BPY4yUUgiBK1haEPT7h0hAKY3jGBfjOKaUAsB0Oj04OJhfj9Y6SRLcL6XE2FhrtdYh/A0hIPedFKkZAMA5xxjDQUIIKSWemlI6D0Jd1/M4KqXmBDgBG+ZhXcQUU2iMEULgIKVUXdfIitpRhPd+Mplc/blCccYYlIXhCSEwxqy1C4KmaQgho9HIe88Yw1Djqcuy9N4rpcbjMZ4AkyqEwLxKKW8TID3nHHO4SNHPy0t0QylVVZW1NoRQVdVkMnHOSSmLokA1AMA5z7IMzamqinMeQhBClGVprW2apixLzvkdi14IvAa0G014+A7efqMe+Sc/fM6eqr4Q/wCSJbzAeXLkgwAAAABJRU5ErkJggg==" nextheight="1020" nextwidth="2250" class="image-node embed"><p>What decentralization can provide in the first issue, the data loss is clear. Just like you never lose your financial transaction history on a blockchain, you can also make sure you never lose the data you have collected. Using a decentralized solution dedicated to permanent storage, such as Arweave, you essentially mitigate the risk of losing data by opting out of time-limited storage options that only store as long as you keep paying a monthly fee. There is also an opportunity to create a shared data lake that different projects can use without going through the process of collecting and storing the same data in different places. This reduces the total cost and risk involved with data storage significantly while also making it much easier to build platforms where people collectively collaborate on data collection and review, as previously mentioned in the data quality and volume sections.</p><h2 style="text-align: start"><strong>Privacy and Personally Identifiable Information</strong></h2><p style="text-align: start">The privacy and the security of the users are always serious concerns for any type of data collection task, more so with the type of large volume of meaningful data that we have been talking about. A key term here is PII, or personally identifiable information, which refers to any type of data that can be used to identify an individual, such as their name, ID or phone number, address, biometric credentials, health history, etc. PII is all over the internet, and using data that originally has some name attached to it such as social media posts or blog entries is fine in principle, these should be stripped away up from any type of context that can lead to use of PII. Regulatory measures such as the recent announcement of The EU’s AI Act put a significant focus on privacy concerns like this.</p><p style="text-align: start">Just like the data quality control problem, privacy and PII issues also get significantly harder to solve with the high volume of data that AI developers need to manage. An important difference here, though, is the lack of viability that some quality assurance solutions have in terms of privacy protection. You can create collaborative environments where thousands of people go over the collected datasets to rate their quality and usefulness, but outsourcing the privacy and PII controls of a dataset to a collaborative platform or a market is definitely not an option.</p><p style="text-align: start">Though cleaning the collected datasets from PII and ensuring privacy standards are met are still possible, a complementary approach preventing these sensitive data from entering the datasets upfront can be much more effective. Blockchains enable various use cases for technologies like zero-knowledge proofs (ZKP) and decentralized identities (DID) that make the privacy-preserving approach the default mode of interaction for online activity.</p><h2 style="text-align: start"><strong>Intellectual Property and Copyright</strong></h2><p style="text-align: start">Another widely debated topic around data collection for AI is intellectual property rights and copyright infringement. There are many online platforms and media outlets that challenge AI companies on the use of material that is subject to copyright, saying that the training and use of these AI models <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://www.nytimes.com/2023/12/27/business/media/new-york-times-open-ai-microsoft-lawsuit.html"><strong><u>create a competitive advantage</u></strong></a> over the rightful owners of the original content.</p><p style="text-align: start">Though copyright and IP issues have always been a tricky topic, first, the mass adoption of social media and now the growing usage of AI chatbots have made it even tougher to resolve conflicts related to these matters. Luckily, however, one of the most well-known technologies related to blockchains and NFTs has been focused on this problem ever since its inception, though the crazy prices for cartoon profile pictures overshadowed the discourse for a while.</p><p style="text-align: start">NFTs make the ownership and use rights of creative/intellectual material very clear and transparent thanks to the transparency and immutability of on-chain actions. These tokens can be used to verify and recognize which material is subject to what type of procedures, making both the data cleaning process and dealing with lawsuits easier. One example of how copyrights and licenses can be handled well is <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://monkedao.io/"><strong><u>MonkeDAO</u></strong></a>’s <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://monkedao.tensor.trade/trade/smb_gen3"><strong><u>SMB Gen3</u></strong></a> collection. With each NFT, there is an NFT license living in the token’s metadata that provides information on how the token holder can use their NFT and create derivatives and what rights they have over it.</p><figure float="none" width="227px" data-type="figure" class="img-center" style="max-width: 227px;"><img src="https://storage.googleapis.com/papyrus_images/0c36514ee95d37e8c676f804150eb3cf.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><figure float="none" width="422px" data-type="figure" class="img-center" style="max-width: 422px;"><img src="https://storage.googleapis.com/papyrus_images/51475437f19980fae9ca9d2d8f672eb4.png" class="image-node embed"><figcaption htmlattributes="[object Object]" class="hide-figcaption"></figcaption></figure><p>Proving the origin of an IP online has also become easier and more reliable through decentralization. With centralized social media and content platforms, there’s always a significant risk of losing access to your content or account or even the platform shutting down completely. This makes it very challenging to prove you were the first person to create/post that piece of content and risks your right to distribute and monetize it in the future. Decentralized alternatives like <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://www.notion.so/2e5daf5e49114cb883d9ea939661550c?pvs=21"><strong><u>Zora</u></strong></a> enable creators to share their works in a way that makes the origin very clear, hence protecting their intellectual (or, <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://zora.co/writings/mintellectual-property"><strong><u>mintellectual</u></strong></a>) property rights over it.</p><h1 style="text-align: start"><strong>Risks and Challenges in Decentralization of Data Collection</strong></h1><p style="text-align: start">While we went over many use cases where decentralization can help in solving data collection problems, there are still some risks and open problems involved with these methods, as is the case for any other solution.</p><p style="text-align: start">One of the problems that decentralized solutions often introduce is the default anonymity of the users. As previously mentioned, there are definitely cases where this anonymity comes as a strength in terms of privacy, but this doesn’t fully cancel out the potential problems that can arise from it. When there is a regulatory issue, for example, related to copyrighted or harmful content, the offenders being anonymous can create an even bigger issue, putting the platform at risk. A now famous example of anonymity rights being criminalized is <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://techcrunch.com/2022/08/08/treasury-tornado-cash-laundering-stolen-crypto/"><strong><u>Tornado Cash</u></strong></a>, where one of the developers got arrested.</p><p style="text-align: start">A similar concern can arise from permanently storing data on a decentralized network as there can be harmful content in the data uploaded from crowdsourced platforms, though these are already taken into account in the decentralized <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://www.arweave.org/yellow-paper.pdf"><strong><u>storage solutions</u></strong></a> through <a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://jonniesparkles.medium.com/censorshop-and-content-moderation-e9bbb1293c43"><strong><em><u>self-censoring</u></em></strong></a> models.</p><p style="text-align: start">The last issue is the eventual tradeoff between volume and quality incentives since no matter how the platform is structured, there will always be people who will upload more data with lower quality or high-quality data with low volume. Collaborative platforms and information markets need to balance out these initially contradicting incentives to create an environment where quality checks help accelerate the higher volume contributions. These can be done through various points mechanisms where scoring weights are adjusted based on the priorities of the platform, ensuring the overall data quality in the ecosystem.</p><h1 style="text-align: start"><strong>Conclusion</strong></h1><p style="text-align: start">While there are still open problems remaining when it comes to data collection, rapid developments in AI and increasing adoption of its apps will require new solutions that will take the existing approaches to the next level. The power of decentralization mainly comes from the collective contribution possibilities where transparency and accountability ensure the best quality data flows in high volumes while preserving the users’ privacy as well as creator rights.</p><p style="text-align: start">As we progress further in the intersecting platforms that bring decentralization and AI data collection together, there will be more opportunities to empower better coordination paradigms toward smoother data collection pipelines.</p><p style="text-align: start">In the next article, we will talk about how decentralization can help further down the pipeline, improving the model-building and fine-tuning processes.</p><h1 style="text-align: start"><strong>References</strong></h1><ul><li><p>“Challenges and Applications of Large Language Models” (<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://arxiv.org/pdf/2307.10169.pdf"><strong><u>Jean Kaddour et al., 2023</u></strong></a>)</p></li><li><p>“Data Collection and Quality Challenges in Deep Learning: A Data-Centric AI Perspective” (<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://arxiv.org/pdf/2112.06409v3.pdf"><strong><u>Steven Euijong Whang et al., 2022</u></strong></a>)</p></li><li><p><a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://techcrunch.com/2022/08/08/treasury-tornado-cash-laundering-stolen-crypto/"><strong><u>https://www.springboard.com/blog/data-science/machine-learning-gpt-3-open-ai/</u></strong></a></p></li><li><p>“Insights from the Refinitiv 2019 Artificial Intelligence / Machine Learning Global Study” (<a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://www.lseg.com/content/dam/lseg/en_us/documents/media-centre/press-releases/refinitiv/refinitiv-ai-ml-survey-report.pdf"><strong><u>link</u></strong></a>)</p></li><li><p><a target="_blank" rel="nofollow noopener noreferrer" class="dont-break-out" href="https://5xwvoh2nm2ugdyaahqfk5mb2vrfa2dlvrs7vn5vbpvqlhvuymrcq.arweave.net/7e1XH01mqGHgADwKrrA6rEoNDXWMv1b2oX1gs9aYZEU"><strong><u>SMB Gen3 NFT License</u></strong></a></p></li><li><p><a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://techcrunch.com/2022/08/08/treasury-tornado-cash-laundering-stolen-crypto/"><strong><u>https://zora.co/writings/mintellectual-property</u></strong></a></p><div><div class="callout-base callout-info" data-node-view-wrapper="" style="white-space:normal"><img src="https://paragraph.xyz/editor/callout/information-icon.png" class="callout-button"><div class="callout-content"><div><p>This article was originally posted on the <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.firstbatch.xyz/blog/towards-decentralized-ai-part-1-data-collection">FirstBatch blog</a>. It was reviewed by <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/batuhan-aktas-38692b148/">Batuhan</a> and <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/kerim-kaya-552878129/">Kerim</a>, edited by <a target="_blank" rel="noopener noreferrer nofollow ugc" class="dont-break-out" href="https://www.linkedin.com/in/ilkyazyesilserit/">İlkyaz</a>.</p></div></div></div></div></li></ul><p></p>]]></content:encoded>
            <author>lab@newsletter.paragraph.com (omer)</author>
            <category>ai</category>
            <category>blockchain</category>
            <category>web3</category>
            <enclosure url="https://storage.googleapis.com/papyrus_images/5ed585e40260019abdd4cca71832ab7b.png" length="0" type="image/png"/>
        </item>
    </channel>
</rss>