28 Sep, 2026
Today, Grass supplies most of the world’s leading AI labs with data to train frontier AI models. Because most of the internet is inaccessible to modern AI crawlers, Grass uses millions of residential internet connections to access it. What a lot of people might not know is that training data is our first product. The long-term goal is to give AI a live connection to the world’s information.
Historically, access to information has expanded human capacity. It allowed people to learn from work they didn’t personally perform, build on the discoveries of others, and solve problems that would otherwise be beyond the limits of their individual experience. The internet pushed this to an extreme, and humanity now has access to more information than any human could ever consume.
While the public internet made information abundant, artificial intelligence has made our ability to use it abundant as well. For the first time, machines can consume and reason over information on a scale that is approaching our species’s ability to produce it.
This is why, back in 2022, we predicted that access to the public internet would quickly deteriorate. Over the last four years, internet data has become one of the most important inputs to scaling intelligence, while access to it has become increasingly restricted.
Training data and inference data An AI model can do two things with data: it can learn from it, or it can use it to inform its decisions. These processes are called “training” and “inference”, and in today’s model architectures, they are distinct processes.
Today, the demand for inference is growing exponentially. This is easy to track using a variety of public reports, and has led some people to wonder why we even bother with training data when the market for inference (and by transitivity, inference data) is exploding.
It’s a fair question, and the simple answer is that training data is a much quicker path to growth today. At the frontier, companies are training 12 trillion parameter models. Under Chinchilla optimal scaling laws, this means frontier AI models require 240 trillion tokens of training data after deduplication, parsing, and processing. To put this size into perspective, it is more than 100 times more information than every single published book in human history combined. This scale of data only exists on the internet, and the internet is becoming exponentially more difficult to access.
This has created an enormous market for anyone capable of accessing parts of the internet that are blocking AI crawlers and scrapers. It has also given a way for Grass to generate revenue today while continuing to build the infrastructure required for inference data. Obviously, funding long-term goals with existing demand allows Grass to grow with less dependency on outside capital. But a more nuanced point is that fragmentation of training and inference will eventually be solved, and under a future continuous learning paradigm, they will be the same thing.
Live data The top AI labs in the world use Grass to absorb trillions of tokens of data as discrete snapshots across training cycles that can last months to years. Once a model with fixed weights is trained and released into the wild, it is relied upon to operate in a world that is constantly changing. This creates an obvious limitation, which is that a model’s knowledge starts becoming stale the moment training completes.
Today, our infrastructure is optimized for collecting enormous amounts of data from the public web ahead of a training run. The next step is to take the same exabyte-scale infrastructure, and allow AI models to tap into it directly at inference time.
Evolving Grass from delivering snapshots of the internet to becoming AI’s live connection to the real world will be materialized over multiple products.
First, we will take the access layer that we already use to collect the internet at scale, and make it available programmatically. Our Contents API will allow a model to open and render any public webpage through any of the Grass network’s millions of residential internet connections.
Of course, the ability to access anything on the internet is only useful if a model knows where to look. Our training data business has already made Grass excellent at crawling any domain in depth, and now it’s time to scale our crawling efforts horizontally, across every domain possible. Soon, Grass will have indexed enough of the public web to release its own Search API and allow models to not just access the web but also discover where to look for information.
The internet is much more than text. An enormous quantity of information that humans produce exists in other modalities as well. Videos, images, spreadsheets, PDFs, and other formats are difficult for traditional search engines to process efficiently. Multimodal grounding will allow AI to retrieve and reason over any form of data.
Though they will be released as separate products, these tools are all expressions of the same underlying goal: AI models should be able to find information, access it, and reason over it regardless of its modality.
Why Grass will be the information layer for intelligence We have several unique advantages that we believe will continue to compound. Grass has access to millions of residential connections that have opted in and are being compensated for being part of the network.
We are already crawling the web at extraordinary scale in order to serve training data demand that exists today. Instead of financing entirely against hypothetical future inference data revenue, every new training dataset gives us a reason to expand our coverage, drive down costs, and improve our infrastructure.
This creates a flywheel. Training data demand incentivizes Grass to crawl more of the internet. Crawling more of the internet improves coverage and freshness, which improves the Grass web index. A better web index makes search and inference products more useful. And as demand for all of these products grows, it gives us more reason to both expand the network and expand crawling activity.
The plan In summary, this is our plan:
Build the world’s largest network for public web data collection. Use it to supply the data required to train the world’s most capable AI models. Use that money to expand our coverage and build search and retrieval. Give AI a live connection to the world’s information. Let’s free the internet.
Written by
Theo Brandt
Contributor Education Writer
Theo Brandt writes Grass's getting-started and trust-and-safety guides, making the network easy to understand for anyone, what Grass is, how sharing your unused internet bandwidth works, and what to expect, in plain language, without the jargon.
Follow on X
™