# The Web3 + AI Daily: On The Commoditization of Inference On Chain

*Daily insights into the fascinating convergence of crypto and AI.*

By [The Web3 + AI Newsletter](https://web3plusai.xyz) · 2026-07-23

web3plusai, web3ai, albena

---

Hello and welcome!
------------------

> **This is the 74th edition of The Web3 + AI Daily - your definitive guide to the intersection of blockchain and AI.** Today's agenda revolves around AI inference, as it has overtaken training as the dominant share of global GPU demand. As a result, not only that Web3 companies like [**Boundless**](https://www.linkedin.com/company/boundless-networks/) are redirecting their GPU clusters to AI inference, but **on-chain inference capital markets** started to emerge.

**Thank you for being here! Let's dive in.**

* * *

**What's Hot in Web3 + AI?**
----------------------------

### **Boundless Shifts Focus from ZK to AI**

The decentralized compute firm [**Boundless**](https://www.linkedin.com/company/boundless-networks/) has announced a pivot to AI, involving an expansion of its GPU network into AI inference.

> Boundless initially focused its 4,000 GPU cluster on settling Ethereum- and Base-based zero-knowledge proofs on Bitcoin.

The move is in line with the wider trend of compute networks and mining companies directing GPU capacity towards AI operations, to address the unsatiable thirst for machine intelligence.

> Of note, Boundless said it will continue running its zero-knowledge proving network in parallel, and plans to “give its native token, ZKC, a role in its AI network,” by requiring AI operators to stake ZKC to join the network, according to the announcement.

[

Distributed compute startup Boundless expands 4,000-GPU network from ZK to AI
-----------------------------------------------------------------------------

Boundless, which built a 4,000 GPU cluster to settle Ethereum and Base ZK proofs on Bitcoin, is expanding to AI compute.

https://www.theblock.co

![Distributed compute startup Boundless expands 4,000-GPU network from ZK to AI](https://storage.googleapis.com/papyrus_images/5171f4d796a6527a4afd602cef5d709b66fa258737afd4b2190ef60e034ee233.jpg)

](https://www.theblock.co/post/408213/distributed-compute-startup-boundless-gpus-zk-to-ai)

* * *

**Web3 + AI Readings & Conversations**
--------------------------------------

### **Centralized vs. Decentralized Inference Providers**

Maybe I should have started by explaining what AI inference is. That's the process of an already trained AI model running a prompt, an agent loop, an image or a video generation. Every time you ask your chatbot a question, it processes your inquiry to deliver a response through the act of inference.

As you can imagine, there are centralized and decentralized inference providers. In the centralized category, besides hyperscalers like [**Microsoft**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), [**Google**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), and [**Amazon**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), one can also find platforms like [**Fireworks AI**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), [**Together AI**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), [**Replicate**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), [**Baseten**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#), and [**Groq**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#).

At the other end of the spectrum are companies building the permissionless and confidential alternative, such as [**Chutes AI**](https://www.linkedin.com/company/chutesai/), [**Dolphin AI**](https://dphn.ai/), [**Pearl Research Labs**](https://www.linkedin.com/company/pearlresearch/), [**Venice.ai**](http://Venice.ai), and others.

These two categories of inference providers often come with diametrically different value propositions. As [**0xSammy**](https://x.com/0xSammy) writes:

> The mistake is to compare all of these providers as if they are competing in the same market; they are not.

> Traditional providers sell reliability, developer experience and enterprise procurement.

> Crypto AI networks sell cheaper supply, open access, privacy, verifiability and new incentive loops.

[

The Agentic Future (06.23.26): AI Inference Wars
------------------------------------------------

Every prompt now has a cost; as trillions of agents deploy live products, inference is turning into a lucrative market: routed, priced & fought over by cloud giants, GPU networks & crypto AI protocols

https://www.0xsammy.com

![The Agentic Future (06.23.26): AI Inference Wars](https://storage.googleapis.com/papyrus_images/1ad90dd2f815d9b9c1362d35146e92c1cf0b5e516c9456808e09625af4fbb71c.jpg)

](https://www.0xsammy.com/p/the-agentic-future-062326-ai-inference)

### **Decentralized Inference: Confidentiality, Privacy, Affordability**

I’ve often discussed Web3 inference providers, but I believe it would be useful to compile a list of the most prominent ones and highlight what differentiates each of them. Here it is:

*   [**Chutes AI**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#) combines an offering of open-source models with using decentralized GPU supply underneath. I've recently covered a breakthrough of theirs [**here**](https://www.linkedin.com/pulse/web3-ai-daily-71-albena-kostova-nikolova-2tlaf).
    
*   Targon AI / [**Manifold Labs**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#) focuses on delivering confidential compute to serve data-sensitive industries like finance and healthcare:
    

> [**The problem it addresses is obvious: many users will not run sensitive prompts, models or data on infrastructure operated by unknown third parties. Targon’s answer is protected execution through trusted execution environments, encrypted virtual machines, remote attestation and confidential GPU infrastructure. In plain English, the aim is to prove the workload is running in a secure environment and reduce what operators can see.**](https://www.0xsammy.com/p/the-agentic-future-062326-ai-inference)

*   [**Venice.ai**](http://Venice.ai) is best described as an encrypted and private AI chatbot, although, unlike [**ChatGPT**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#) for instance, it gives access to various closed- and open-source models. Venice boasts that it collects no user data and that prompts are never stored.
    
*   [**Dolphin AI**](https://dphn.ai/) began by offering uncensored open models, and then built the inference network around it. [**"Its architecture is often described as peer-to-pool, which means \[that\] GPU owners contribute capacity into model-specific pools, rather than each buyer renting a specific node directly."**](https://www.0xsammy.com/p/the-agentic-future-062326-ai-inference)
    
*   [**Pearl Research Labs**](https://www.linkedin.com/article/edit/7482799282073563136/?author=urn%3Ali%3Afsd_profile%3AACoAAAzEEOMBXZf5gMJpTucLPk-a9SXgLXQI5TI#) is a blockchain network that replaces Bitcoin's [**"SHA-256 hashing algorithm with matrix multiplication, the core operation in AI inference and training. Pearl’s claim is that the same matrix multiplication that serves a customer's inference can double as a mining attempt."**](https://www.galaxy.com/insights/research/inference-capital-markets-ai-compute-gpu-futures-onchain-crypto) What does that mean in practice? Instead of spending huge amounts of computation and energy on math solutions that nobody uses, the network validators compete to win the right to mine the next block by performing actual AI operations.
    
*   Similarly to Pearl, [**Ambient**](https://ambient.xyz/) also let miners run an AI model. However, unlike Pearl, it [**"standardizes the entire network on one large open-weight model and builds its consensus around verifying that model's output."**](https://www.galaxy.com/insights/research/inference-capital-markets-ai-compute-gpu-futures-onchain-crypto)
    

* * *

**Intelligence Is Becoming Tradeable, Both On- and Off-Chain**
--------------------------------------------------------------

### **Commoditizing AI Inference: Who, What, Why?**

Naturally, with the number of people and enterprises using AI, and AI agents actively operating both on the rise, the need for computation resources continues to grow exponentially.

> [**The demand for intelligence is swallowing the industry whole. AI providers don’t have enough compute to service the demand, leaving the gap for Inference providers to step in and service those demand.**](https://defi0xjeff.substack.com/p/the-rise-of-inference-capital-markets)

Just as enterprises are currently shifting away from frontier AI labs toward more affordable open-source models, they are also starting to switch to more cost-efficient inference providers. Moreover, as inference is becoming an ever more valuable resourse, subscribers of AI models now sell their unused credits, while owners of idle GPU clusters monetize their capacity.

As [**0xSammy**](https://x.com/0xSammy) noted, the inference market will soon resemble the electricity one: many suppliers with similar offerings, competing more on reliability and resilience than on price. Yet, the Web3 alternatives still boast much lower prices than hyperscalers.

* * *

### **On-Chain Inference Markets**

As tradition goes, the Web3 space supercharges the trend of AI inference commoditization with experimentation, incentive design, and innovative tokenomics. **The inference is being tokenized, not to mention verified, and it's becoming yield- or reward-bearing.**

**Last week** [**Galaxy**](https://www.linkedin.com/company/galaxyhq/) **published an in-depth analysis into the emerging AI inference markets, and how the Web3 space is pricing compute, tokenizing access, and financing hardware.** I'll share a summary below, but I still recommend you read the full article.

> The term _onchain inference capital markets_ describes the set of networks, protocols, supporting infrastructure, and applications coordinating AI model [**inference**](https://hai.stanford.edu/ai-definitions/what-is-inference) outside the centralized API surface controlled by [**frontier labs**](https://www.longtermwiki.com/wiki/E820) and [**hyperscalers**](https://www.redhat.com/en/topics/cloud-computing/what-is-a-hyperscaler), together with the financial layer now forming on top of that activity. Rather than routing every API call through [**frontier model**](https://www.nvidia.com/en-us/glossary/frontier-models/) providers like OpenAI, Anthropic, or the underlying cloud providers that service them, users can send prompts to networks of GPU operators coordinated by crypto token incentives and onchain settlement, and in some configurations receive cryptographic or economic guarantees about output correctness and privacy.

> In recent years, [**decentralized GPU marketplaces**](https://www.galaxy.com/insights/research/decentralized-ai-training), inference protocols, payment rails, tokenization, capital formation vehicles, and onchain liquidity each had its moment. What is new is that these primitives are converging into a single integrated system, an _inference capital market_, which is projected to find growing demand as inference is increasingly used for all work.

[

Inference Capital Markets
-------------------------

GPU futures, tokenized inference access, useful proof-of-work networks, and onchain GPU lending are converging into a new financial layer for AI compute. Galaxy Research maps the emerging inference capital market.

https://www.galaxy.com

![Inference Capital Markets](https://storage.googleapis.com/papyrus_images/9e7a5f07eafdcd51a75817f8fd5ea1c5df4bd9d662a324a2f91926232e412575.jpg)

](https://www.galaxy.com/insights/research/inference-capital-markets-ai-compute-gpu-futures-onchain-crypto)

At first glance, the main difference with regular inference providers is that the crypto-native ones can [**"source capacity from decentralized GPU networks, accept stablecoins or tokens as payment, include privacy guarantees, or attach tokenized access rights to usage."**](https://www.galaxy.com/insights/research/inference-capital-markets-ai-compute-gpu-futures-onchain-crypto)

However, the uniqueness of what Web3 has to offer comes to light on the financialization side, **where crypto changes how inference is owned, priced, and financed**. As Galaxy outlines:

> Financializing inference has attracted a range of onchain projects that use blockchain payment rails and tokenization to turn inference activity into tradable assets. This takes three forms.

> _Inference service providers_ like Venice and Morpheus tokenize inference access, turning a claim on future inference into something that can be held, priced, and resold.

> _Proof of useful work_ projects like Pearl and Ambient tokenize inference production, paying out a token for the work of serving it.

> _Credit providers_ like [**USD.AI**](http://USD.AI) do something different. Rather than tokenizing inference, they finance the hardware it runs on, using stablecoin deposits to fund the GPUs and data centers underneath.

> Collectively, these components come together to form onchain inference capital markets.

* * *

### **Venice AI's Model**

All of these models are quite interesting, but I'm sure the tokenomics nerds among you would appreciate the [**Venice.ai**](http://Venice.ai) one the most. It uses a two-token system to turn inference into an ownable and easily transferable asset. Here's how it works in short:

*   $VVV is the main Venice token. Holders can stake it to earn yield or to mint the $DIEM token.
    
*   Venice runs a buy-and-burn mechanism on $VVV - [**a discretionary burns funded from general revenue, and a programmatic burn that routes a fixed portion of every new subscription into buying and burning the token. To date, 42% of VVV has been burned.**](https://www.galaxy.com/insights/research/inference-capital-markets-ai-compute-gpu-futures-onchain-crypto)
    
*   **Each $DIEM token gives holders $1 of daily credits to spend on Venice's AI tools, in perpetuity.** [**"**](https://defi0xjeff.substack.com/p/the-rise-of-inference-capital-markets)[**It's a tokenized platform credit that expires and replenishes itself daily."**](https://defi0xjeff.substack.com/p/the-rise-of-inference-capital-markets)
    

> DIEM is an experiment in how to tokenize and deliver inference access. What makes it distinct is ownership. It lets users own the inference they consume rather than rent it. A buyer paying per request gets nothing back once the inference is spent, while a holder of tokenized access owns an asset they can keep, transfer, or sell.

> DIEM wraps a claim on future inference into something a holder can mint, own, and resell.

Since $DIEM represents a $1/day inference credit, a user can decide not to use it and sell it instead. Accordingly, people have already started speculating on its future value and secondary markets have appeared:

> [**Secondary markets enable users to stake DIEM or sell the unused credit at a discount. Thus, earning yields 10-20% APR on DIEM. Morpho has this wstDIEM/DIEM where wstDIEM (LSD version of DIEM) can be used as a collateral to borrow DIEM which essentially allow users to leverage farm DIEM (up to ~56% APR according to Liquid protocol, the team behind wstDIEM).**](https://defi0xjeff.substack.com/p/the-rise-of-inference-capital-markets)

* * *

**Decentralized inference is a fascinating subject that I'll keep covering here. Let me know which aspects or companies you find compelling, I'm looking forward to hearing your thoughts.**

* * *

> Thank you for reading! My name is Albena, and every day I share insights into the ground-breaking convergence of blockchain and AI. If you’re enjoying them, hit the subscribe button and never miss a key Crypto × AI update.

[Subscribe](https://web3plusai.xyz/subscribe)

> The Web3 + AI Newsletter is an independent, ad-free publication that I have been building on my own since 2023. If you find value in my work, please consider supporting it at the link below. I greatly appreciate it.

[Support](https://buy.stripe.com/14A4gs0gs14wgY1b873gk01)

> The Web3 + AI Book Club is live on Fable! Join us in exploring our July title - 'Empire of AI' by Karen Hao. Follow the link below to read with us.

[The Web3 + AI Book Club](https://fable.co/club/the-web3-ai-book-club-with-albena-417116729458?club_type=free)

> I'm looking forward to connecting with fellow Crypto x AI enthusiasts, so don't hesitate to reach out on social media.

[LinkedIn](https://www.linkedin.com/in/albena-kostova-nikolova/)

[Firefly.Social](https://firefly.social/profile/farcaster/albena)

[Zora](https://zora.co/@albena)

* * *

**Disclaimer:** None of this should or could be considered financial advice. You should not take my words for granted; rather, do your own research (DYOR) and share your thoughts to encourage a fruitful discussion.

---

*Originally published on [The Web3 + AI Newsletter](https://web3plusai.xyz/inference_commoditization)*
