Open Weight Thoughts
All articles

When Generated Content Becomes the Internet, Trust Will Move Behind Doors

By E. Yao

  • artificial intelligence
  • internet
  • online communities
  • platforms
  • media
When Generated Content Becomes the Internet, Trust Will Move Behind Doors

ChatGPT was released on November 30, 2022. Four years later, its most important effect is not that it can draft an email or make an illustration. It has lowered the cost of producing a plausible public artifact, whether that artifact is a product review, a tutorial, a comment, a search-optimized article, a song, a short video, or a fake personal recommendation. The cost of publishing was already close to zero. The cost of generating enough material to occupy attention is now falling toward zero as well.

By December 31, 2030, the open web will be dominated by machine-generated material by volume, while a growing share of valuable human interaction will move into closed, identity-bearing spaces: paid communities, private group chats, workplace networks, local clubs, invite-only forums, and small services with real moderation. This is a prediction about where people will place trust and conduct conversation, not a claim that the public web will disappear, that humans will stop publishing, or that every closed community will be healthier than an open one.

For this essay, “internet-generated content” means material whose first publishable version is produced by a generative system rather than written, filmed, recorded, or drawn by a person. A human can prompt, revise, fact-check, and approve it and it still counts as internet-generated under this definition. A person using spell-check, transcription, or a camera filter does not. The distinction matters because the coming shift concerns the supply of finished-seeming material, not every use of software in creative work.

Volume will become a bad proxy for value

The majority threshold will arrive first in countable formats, not necessarily in the material people remember. A million generated pages can be produced faster than one investigative report; a hundred thousand synthetic replies can be distributed faster than a useful answer from someone who has done the work. That makes content volume a leading indicator of the change, but a weak measure of cultural importance. Attention, repeat visits, purchasing decisions, and private sharing will lag behind it.

Public platforms will initially treat the extra supply as a benefit. More posts mean more opportunities to place ads, more pages to index, more search queries answered, and more apparent activity in a feed. But abundance changes the economics of attention. When nearly every query returns fluent text and every product category acquires infinite reviews, the scarce input becomes evidence: who made a claim, what they have done, whether they can be contacted, and whether a group will impose a cost for deception.

That is why the likely destination is not a return to the old internet in a literal sense. Independent websites, message boards, and mailing lists will persist, but the larger movement will be toward bounded systems where membership, history, and moderation can be observed. A neighborhood Slack group, a paid trade community, a Discord server run by known experts, a union forum, or a hobby club with an application process can make a member’s reputation portable within the group. A public comment box generally cannot.

The mechanism is a collapsing cost curve for plausible speech

The driver is the cost curve. Generative models reduce the marginal cost of producing credible-looking language, images, audio, and video more quickly than open platforms can reduce the cost of verifying origin and intent. This imbalance creates the trend. It does not require a sudden collapse in human judgment, a single malicious actor, or a universal preference for private spaces.

The causal chain has several steps. First, a publisher, marketer, political operator, or hobbyist can create far more variants of a message than a human team could make by hand. Second, distribution systems reward a portion of those variants because ranking models can observe clicks, watch time, keywords, and shares more cheaply than they can observe truth, authorship, or good faith. Third, users encounter a rising proportion of material that is coherent but ungrounded. Fourth, the cost of deciding what deserves belief rises for each user. Finally, users route consequential questions toward places where context is denser and the sender has something to lose.

This is why better model output can intensify the move away from the open web rather than prevent it. Obvious junk is cheap to dismiss. Competent imitation is expensive to audit. Once generated material is good enough to pass a casual glance, public feeds become less useful for deciding whom to trust, even if they remain useful for discovering what exists.

Search will feel this pressure first. Search engines were built around a premise that links, citations, and repeated publication created signals of authority. Generated pages can reproduce the outward forms of those signals at industrial scale. Search companies will respond with answer boxes, source selection, provenance labels, tighter crawling rules, and more reliance on sites they already know. Those changes will make search more efficient for routine questions, but they will also concentrate traffic among a smaller set of recognized publishers and platforms. The open web will remain visible, yet less evenly reachable.

Social feeds have a related problem. Their historical advantage was access to strangers: an unexpected person could post something useful, funny, revealing, or newsworthy and reach a large audience. When generated accounts can imitate that output continuously, the feed’s promise weakens. Platforms can label synthetic media and remove automated accounts, but enforcement must identify both the content and the coordinated behavior around it. The more public reach is worth, the more effort will go into evading those controls.

Closed spaces solve a specific problem, and create others

Closed forums will grow because they can impose friction. A membership fee, an invitation, a verified workplace address, a local event, a moderator who knows the regulars, or a public history inside a group all make mass impersonation more expensive. None proves that a person is honest. Each gives a community more information with which to judge them.

The product that people will buy is not privacy in the abstract. It is lower verification cost. A parent asking about a school, a software engineer debugging a production system, a collector judging an expensive item, or a patient comparing a treatment experience wants an answer from somebody whose experience can be assessed. In a bounded group, members can notice when someone has never attended, never contributed, or never been accountable for a bad recommendation. That social memory is difficult to manufacture at scale.

There is a cost. Closed groups fragment public knowledge. Useful answers become trapped behind subscriptions, invitations, and workplace boundaries. Moderators can become arbitrary gatekeepers. Small groups can also amplify local myths because members see fewer challenges from outside. The next internet will therefore not divide neatly into a false public web and a true private web. It will divide between low-cost discovery and higher-cost trust, with many failures on both sides.

The email analogy explains the pressure, but not the whole outcome

Email offers the strongest historical analogy. It made communication inexpensive and open, then spam made indiscriminate outreach cheap enough to degrade the inbox. Filters, reputation systems, authentication standards, and controlled sending infrastructure followed because users needed a way to separate expected messages from unwanted ones. The important lesson is economic: when sending becomes nearly free, receiving and evaluating become the expensive side of the exchange.

The analogy breaks down in a material way. Spam is usually trying to reach a recipient. Generated content can earn money, influence a purchase, fill a search result, train another system, or create the appearance of consensus even when no one reads it closely. It can also be entertaining and genuinely useful. The problem is not that machine-made material exists. The problem is that open distribution rewards volume before it rewards provenance.

There is also a warning in the history of digital publishing. Adoption and profit do not necessarily travel together. The web made publishing available to far more people, while much of the advertising profit accumulated in the companies that controlled search, feeds, and ad targeting. Generative tools will make production available to more people again, but the largest financial gains may go to model providers, cloud operators, identity services, payment systems, and platforms that own the distribution bottleneck. More creators will be able to make things. That does not mean more creators will earn a living from them.

The best opposing case is that abundance expands culture

The strongest opposing view is not that synthetic material will stay rare. It is that cheap production will create more useful knowledge than confusion, and that ranking systems will improve fast enough to preserve open discovery. Marc Andreessen has made a broad version of this argument about artificial intelligence: more capable software can expand human capacity, lower prices, and produce more goods and services rather than merely displace existing work.

That case gets an important point right. A large supply of generated material will be valuable in domains with clear feedback loops. Code can be tested. Product descriptions can be checked against inventory. Language translation can be reviewed by a fluent speaker. Educational exercises can be tailored to a student. In these settings, low production cost is an advantage because errors are relatively cheap to find.

The disagreement is about public, low-context claims. A generated guide to repairing a bicycle can be tested in a garage. A generated account of a local scandal, a restaurant recommendation, a medical anecdote, or a claim about an obscure historical event often cannot. Where truth is hard to observe, the public internet will accumulate text faster than it can accumulate confidence. Better ranking will help, but it cannot eliminate the value of knowing who is speaking and what they risk by being wrong.

What would falsify this forecast

By December 31, 2028, I expect a reproducible sample of 10,000 newly published English-language public pages from major search indexes to show that at least 70 percent were first drafted by generative systems, whether or not human editors later altered them. I would consider that prediction wrong if a credible, transparent sampling effort finds less than 50 percent, or if major publishers and platforms can demonstrate that automated publication remains a small share of new public material.

By December 31, 2031, I expect at least 25 percent of United States adults who use the internet to pay for, or receive through work or membership, access to one bounded online community they regard as more trustworthy than public search or social feeds for a consequential question. I would consider that prediction wrong if representative surveys put the figure below 15 percent and show that users instead rely more heavily on open, provenance-rich public systems. The concrete test is where they ask for advice when the answer can cost them money, time, or health.