Cover photo

The Web3 + AI Daily: On The Commoditization of Inference On Chain

Daily insights into the fascinating convergence of crypto and AI.

Hello and welcome!

This is the 74th edition of The Web3 + AI Daily - your definitive guide to the intersection of blockchain and AI. Today's agenda revolves around AI inference, as it has overtaken training as the dominant share of global GPU demand. As a result, not only that Web3 companies like Boundless are redirecting their GPU clusters to AI inference, but on-chain inference capital markets started to emerge.

Thank you for being here! Let's dive in.


What's Hot in Web3 + AI?

Boundless Shifts Focus from ZK to AI

The decentralized compute firm Boundless has announced a pivot to AI, involving an expansion of its GPU network into AI inference.

Boundless initially focused its 4,000 GPU cluster on settling Ethereum- and Base-based zero-knowledge proofs on Bitcoin.

The move is in line with the wider trend of compute networks and mining companies directing GPU capacity towards AI operations, to address the unsatiable thirst for machine intelligence.

Of note, Boundless said it will continue running its zero-knowledge proving network in parallel, and plans to “give its native token, ZKC, a role in its AI network,” by requiring AI operators to stake ZKC to join the network, according to the announcement.


Web3 + AI Readings & Conversations

Centralized vs. Decentralized Inference Providers

Maybe I should have started by explaining what AI inference is. That's the process of an already trained AI model running a prompt, an agent loop, an image or a video generation. Every time you ask your chatbot a question, it processes your inquiry to deliver a response through the act of inference.

As you can imagine, there are centralized and decentralized inference providers. In the centralized category, besides hyperscalers like Microsoft, Google, and Amazon, one can also find platforms like Fireworks AI, Together AI, Replicate, Baseten, and Groq.

At the other end of the spectrum are companies building the permissionless and confidential alternative, such as Chutes AI, Dolphin AI, Pearl Research Labs, Venice.ai, and others.

These two categories of inference providers often come with diametrically different value propositions. As 0xSammy writes:

The mistake is to compare all of these providers as if they are competing in the same market; they are not.

Traditional providers sell reliability, developer experience and enterprise procurement.

Crypto AI networks sell cheaper supply, open access, privacy, verifiability and new incentive loops.

Decentralized Inference: Confidentiality, Privacy, Affordability

I’ve often discussed Web3 inference providers, but I believe it would be useful to compile a list of the most prominent ones and highlight what differentiates each of them. Here it is:

  • Chutes AI combines an offering of open-source models with using decentralized GPU supply underneath. I've recently covered a breakthrough of theirs here.

  • Targon AI / Manifold Labs focuses on delivering confidential compute to serve data-sensitive industries like finance and healthcare:

The problem it addresses is obvious: many users will not run sensitive prompts, models or data on infrastructure operated by unknown third parties. Targon’s answer is protected execution through trusted execution environments, encrypted virtual machines, remote attestation and confidential GPU infrastructure. In plain English, the aim is to prove the workload is running in a secure environment and reduce what operators can see.


Intelligence Is Becoming Tradeable, Both On- and Off-Chain

Commoditizing AI Inference: Who, What, Why?

Naturally, with the number of people and enterprises using AI, and AI agents actively operating both on the rise, the need for computation resources continues to grow exponentially.

The demand for intelligence is swallowing the industry whole. AI providers don’t have enough compute to service the demand, leaving the gap for Inference providers to step in and service those demand.

Just as enterprises are currently shifting away from frontier AI labs toward more affordable open-source models, they are also starting to switch to more cost-efficient inference providers. Moreover, as inference is becoming an ever more valuable resourse, subscribers of AI models now sell their unused credits, while owners of idle GPU clusters monetize their capacity.

As 0xSammy noted, the inference market will soon resemble the electricity one: many suppliers with similar offerings, competing more on reliability and resilience than on price. Yet, the Web3 alternatives still boast much lower prices than hyperscalers.


On-Chain Inference Markets

As tradition goes, the Web3 space supercharges the trend of AI inference commoditization with experimentation, incentive design, and innovative tokenomics. The inference is being tokenized, not to mention verified, and it's becoming yield- or reward-bearing.

Last week Galaxy published an in-depth analysis into the emerging AI inference markets, and how the Web3 space is pricing compute, tokenizing access, and financing hardware. I'll share a summary below, but I still recommend you read the full article.

The term onchain inference capital markets describes the set of networks, protocols, supporting infrastructure, and applications coordinating AI model inference outside the centralized API surface controlled by frontier labs and hyperscalers, together with the financial layer now forming on top of that activity. Rather than routing every API call through frontier model providers like OpenAI, Anthropic, or the underlying cloud providers that service them, users can send prompts to networks of GPU operators coordinated by crypto token incentives and onchain settlement, and in some configurations receive cryptographic or economic guarantees about output correctness and privacy.

In recent years, decentralized GPU marketplaces, inference protocols, payment rails, tokenization, capital formation vehicles, and onchain liquidity each had its moment. What is new is that these primitives are converging into a single integrated system, an inference capital market, which is projected to find growing demand as inference is increasingly used for all work.

At first glance, the main difference with regular inference providers is that the crypto-native ones can "source capacity from decentralized GPU networks, accept stablecoins or tokens as payment, include privacy guarantees, or attach tokenized access rights to usage."

However, the uniqueness of what Web3 has to offer comes to light on the financialization side, where crypto changes how inference is owned, priced, and financed. As Galaxy outlines:

Financializing inference has attracted a range of onchain projects that use blockchain payment rails and tokenization to turn inference activity into tradable assets. This takes three forms.

Inference service providers like Venice and Morpheus tokenize inference access, turning a claim on future inference into something that can be held, priced, and resold.

Proof of useful work projects like Pearl and Ambient tokenize inference production, paying out a token for the work of serving it.

Credit providers like USD.AI do something different. Rather than tokenizing inference, they finance the hardware it runs on, using stablecoin deposits to fund the GPUs and data centers underneath.

Collectively, these components come together to form onchain inference capital markets.


Venice AI's Model

All of these models are quite interesting, but I'm sure the tokenomics nerds among you would appreciate the Venice.ai one the most. It uses a two-token system to turn inference into an ownable and easily transferable asset. Here's how it works in short:

DIEM is an experiment in how to tokenize and deliver inference access. What makes it distinct is ownership. It lets users own the inference they consume rather than rent it. A buyer paying per request gets nothing back once the inference is spent, while a holder of tokenized access owns an asset they can keep, transfer, or sell.

DIEM wraps a claim on future inference into something a holder can mint, own, and resell.

Since $DIEM represents a $1/day inference credit, a user can decide not to use it and sell it instead. Accordingly, people have already started speculating on its future value and secondary markets have appeared:

Secondary markets enable users to stake DIEM or sell the unused credit at a discount. Thus, earning yields 10-20% APR on DIEM. Morpho has this wstDIEM/DIEM where wstDIEM (LSD version of DIEM) can be used as a collateral to borrow DIEM which essentially allow users to leverage farm DIEM (up to ~56% APR according to Liquid protocol, the team behind wstDIEM).


Decentralized inference is a fascinating subject that I'll keep covering here. Let me know which aspects or companies you find compelling, I'm looking forward to hearing your thoughts.


Thank you for reading! My name is Albena, and every day I share insights into the ground-breaking convergence of blockchain and AI. If you’re enjoying them, hit the subscribe button and never miss a key Crypto × AI update.

Subscribe

The Web3 + AI Newsletter is an independent, ad-free publication that I have been building on my own since 2023. If you find value in my work, please consider supporting it at the link below. I greatly appreciate it.

The Web3 + AI Book Club is live on Fable! Join us in exploring our July title - 'Empire of AI' by Karen Hao. Follow the link below to read with us.

I'm looking forward to connecting with fellow Crypto x AI enthusiasts, so don't hesitate to reach out on social media.


Disclaimer: None of this should or could be considered financial advice. You should not take my words for granted; rather, do your own research (DYOR) and share your thoughts to encourage a fruitful discussion.