Cover photo

The Web3 + AI Daily #70

Daily insights into the fascinating convergence of crypto and AI.

What's Hot in Web3 + AI?

Story Protocol Becomes DATA Foundation to Tackle AI Copyright

In June, Story Protocol. rebranded to The Data Foundation and shifted focus to exclusively solve the copyright issues existing in the AI training infrastructure. Here's why this is important.

One of the most serious problems the current wave of generative AI is creating is copyright infringement - enormous amounts of artistic creations, be it literature, music, or images, are used to train AI models without in any way recompensing the authors. And once again, the crypto space is coming to the rescue.

The DATA Foundation will operate the DATA Network, an onchain registry designed to verify the origins, licensing and consent history of datasets used to train artificial intelligence models.

The startup’s shift comes as AI developers and Big Tech face mounting copyright lawsuits over the data used to train their models and increasing pressure to prove that datasets were collected with proper consent. DATA Foundation is betting blockchain can provide a transparent record of ownership, licensing and provenance for AI training data.

For context, before the rebrand, Story Protocol's mission was to turn every piece of intellectual property into a programmable and permissionlessly licensable asset. Now, it pivots to use its technology to tackle one of today's most urgent issues.

DATA Foundation plans to address this by launching Trace, a public audit and search platform that creates unalterable cryptographic receipts for individual data contributions. The receipts are designed to help AI developers verify the provenance, licensing and consent history of datasets before using them.

I want to highlight this news because it serves as an important reminder of something we should never lose sight of:

The shameless scraping of the public internet and the brazen extraction of value from artists and creators the major AI labs are performing is not an inherent feature of the AI technology. Training LLMs undoubtedly requires vast amounts of data, but the choice not to compensate those who provide that data is one made solely by OpenAI, Anthropic, xAI, and their peers.

Now that the DATA Foundation provides a straightforward way to trace and license training data, let's see what excuse Sam Altman, Dario Amodei, and Elon Musk come up with next for not paying artists.


Web3 + AI Readings & Conversations

How Can Humans Recognize Each Other Online?

In a stimulating essay for Noema Magazine, Renee DiResta, an associate research professor at the Georgetown University McCourt School of Public Policy, raises a number of urgent questions about how we are proving (or not) our personhood online.

I've often written about digital identity and proof of personhood because I believe they're among blockchain's highest-value use cases - DiResta also quotes some of the Web3 companies I've covered here, like World and Humanity Protocol.

However, no existing solution, blockchain-based or not, checks every box a robust digital credential system should. As a result, proving our humanness online, while preserving privacy, security, and usability, remains an open problem. And a problem with crucial implications.

‘Proof of personhood’ technology is … constitutional infrastructure that will shape who can act, speak, transact, delegate authority — and be trusted online.

Current Solutions Are No Longer Working

As DiResta writes, for over 30 years we have been taking it for granted that a human being is on the other end of our online interactions. But that wasn't just because there were far fewer bots than there are today. It was also because intermediaries, like email providers, telecom operators, banks, and others, effectively served as validators of our humanity.

Yet, now, as our shared digital space becomes increasingly populated by AI agents, these validators can no longer manage this task:

When generative and agentic AI began to improve and democratize rapidly, many common tests for establishing personhood broke down. LLMs could suddenly solve many CAPTCHAs and generate photorealistic ID documents to beat tests that asked for uploads. Video and voice filters made it possible to chat in real-time while appearing as someone else entirely.

As a result, scams and fraud are quickly proliferating and becoming more and more sophisticated. Here's a telling example: crypto journalist, author, and podcaster Laura Shin, a person with a decades-long experience in reporting on various social engineering schemes, recently admitted that she had almost gotten hacked.

We are in a golden age of scams. Con artists prey on the fact that it’s increasingly difficult to know who or what is real.

Play Video

AI is Here, and We Need to Act Fast

In an attempt to cover the existing landscape of private digital credential solutions, DiResta cites a number of fascinating blog posts, research papers, and opinion pieces, and I urge you to read them. The author also breaks down the actual challenges these systems face:

The private market for proof-of-personhood credentials has exploded as companies see an opportunity to meet emerging regulatory and anti-fraud needs while letting users avoid civil identity disclosure. The approaches taken within this burgeoning ecosystem vary along two dimensions. The first is where trust lives: what root mechanism verifies that a person is real and eligible for a credential? Does it ultimately trace back to the state, through an underlying government ID? Is it reliant on a physical body — as when a system relies on biometric uniqueness — or the community, through other users vouching for you? Does it leverage a specific physical electronic device? Each choice puts a different limitation or chokepoint around who can appear as a person. The second dimension is the presentation layer: once a person is proven, what credential does she receive? How is it stored, carried and what is disclosed when she uses it?

I would add that the challenges become even more significant when we take into account that, on top of authenticating humans, we'll have to be able to authenticate AI agents, too. We need to be able to authoritatively distinguish between the agent I've empowered and financed to act on my behalf and the one that pretends to be my delegate after stealing my funds.

(Reminder that the Ethereum Foundation's dAI team is working on the latter - learn more here.)

The truth is that, in a machine-dominated internet, no single solution will be able to serve all our needs. The answer, rather, as DiResta writes, lies in building the identity layer of the internet that consists of different protocols operating in unison. Amid the growing abundance of AI agents, the sooner the world decides what this idenity stack looks like, the better.


Thank you for reading! My name is Albena, and every day I share insights into the groundbreaking convergence of blockchain and AI. If you’re enjoying them, hit the subscribe button and never miss a key Crypto × AI update.

Subscribe

The Web3 + AI Book Club is live on Fable! Join us in exploring our July title - 'Empire of AI' by Karen Hao. Follow the link below to read with us.

If you want to support the publication financially, you can either purchase my writer token $WEB3AI, or buy my creator token $ALBENA on ZORA.

I'm looking forward to connecting with fellow Crypto x AI enthusiasts, so don't hesitate to reach out on social media.


Disclaimer: None of this should or could be considered financial advice. You should not take my words for granted; rather, do your own research (DYOR) and share your thoughts to encourage a fruitful discussion.