The old search box offered a bargain. You supplied a few words. It supplied ten blue links, several advertisements wearing civilian clothes, and the faint suspicion that the useful answer lived on page two.
This was not a golden age. Much of the web had already been reorganized around pleasing a ranking system whose inner life was private and whose appetites were expensive. We learned to open six tabs, avoid the recipe writer’s childhood, and recognize a page built by someone who had met the keyword but not the subject.
Then the answer moved to the top.
Large language models can search, compare sources, translate jargon, explain disagreements, and synthesize a landscape into a paragraph. This is often wonderful. A difficult question that once required an afternoon can now receive a useful first answer before the coffee has accepted its responsibilities.
But the paragraph changes the bargain. A ranked list exposes some of the territory between the question and the conclusion. A synthesized answer collapses that territory. The user sees less of the route, fewer boundaries between sources, and fewer signs that reasonable people may have reached different conclusions.
The danger is not merely that machines will answer badly. Bad answers eventually attract attention, if only because they recommend glue or gasoline as a pizza ingredient. The deeper danger is that machines are answering well enough that we stop noticing what was omitted.
Search Was Always an Editor
Search engines never presented the whole web. They crawled some pages, indexed some of what they found, ranked the index, removed material, promoted other material, and displayed the result through a business model. Libraries, newspapers, schools, publishers, and friends have also mediated discovery for a very long time. Humanity did not wander an unfiltered garden of knowledge until a chatbot installed a gate.
we built a global information system and sorted it by popularity, timestamp, and ability to drive ad revenue
So the argument cannot be that old search was neutral and AI search is not. Old search was an editor too.
What changes is the visibility and finality of the editing.
A list of links says, more or less: Here are some places the answer may be. A synthesized response says: Here is the answer. The second form is easier to use. It also makes a long sequence of editorial decisions feel like a property of the world.
Google describes its AI Mode as issuing many searches across related subtopics before producing a response, a process it calls query fan-out (Google, 2025). That may let the system explore more of the web than a person would have patience to visit. It may also leave the person with less knowledge of what was explored, what was rejected, and why.
The machine can travel farther while showing us less of the trip.
Five Ways to Call Search Successful
We usually evaluate search near the beginning of the problem.
Retrieval: Did the system find the requested information?
Relevance: Was the information pertinent to the person’s intent?
Understanding: Did it improve the person’s mental model?
Discovery: Did it reveal something the person did not know to request?
Serendipity: Did something unexpected prove useful?
The first two are easier to measure. A result was clicked. A task was completed. The answer contained the expected fact. The latter three concern what happened to a person’s understanding, which is slower, stranger, and less cooperative with a quarterly dashboard.
Information-science researchers have long distinguished simple lookup from exploratory search for learning and investigation. In exploratory search, the knowledge acquired along the route matters alongside the destination. People compare results, reformulate questions, follow clues, and discover that the problem they began with was not quite the problem they had (White et al., 2007).
The wandering was not always waste. Sometimes it was the work.
LLM search complicates this picture in both directions. A 2026 preprint analyzing more than 200,000 human–ChatGPT interactions found that almost 80 percent of the prompts were not readily answerable through conventional search, suggesting that conversational systems expand what people can ask. Yet for comparable searchable questions, its authors found less diverse information in AI responses than in Google results across most topics. They also observed that diversity in the response predicted changes in the diversity of what users asked next (Yu et al., 2026).
One study reveals this tension without settling it. The machine may help us ask a question we could not have searched, then answer it with a world smaller than the one it found.
The Answer Is a Compression Artifact
Before a sentence reaches the screen, a search system has already narrowed reality several times. Some sources were available to be indexed or retrieved. Some were found. Some ranked highly enough to enter the model’s context. Some influenced the response. A smaller number appeared as citations. The person saw whatever survived.
- 01AvailableWhat exists
- 02RetrievedWhat the system found
- 03RankedWhat it preferred
- 04CitedWhat it showed its work with
- 05SynthesizedWhat survived the prose
- 06SeenWhat reached a person
Every narrowing is a decision, even when the decision arrives wearing the costume of an answer.
Compression is unavoidable and often useful. Thinking itself requires selection. Nobody can read the entire internet before deciding where to eat lunch, and anyone who tries will starve beside several persuasive reviews.
The problem is that compression can conceal disagreement and provenance while retaining the tone of authority. A fluent paragraph can merge distinct claims into one smooth surface. Sources that disagree may be cited together as though they formed a committee and issued a joint statement. A minority view may disappear because it ranked poorly. A primary source may be replaced by a summary of a summary whose author has excellent metadata.
Citation helps, but citation is not the same as provenance. In an audit of four early generative search engines, Nelson Liu, Tianyi Zhang, and Percy Liang found frequent gaps between generated claims and the sources attached to them. Their specific results describe systems from 2023, not the current state of every product, but their distinction remains useful: a system must cite comprehensively, and each citation must actually support the sentence wearing it (Liu, Zhang, and Liang, 2023).
A useful footnote provides an escape hatch from the answer.
Discovery Is a Form of Agency
Search appears to be about information. It is also about attention.
What we encounter influences what we can consider. What we never encounter cannot be compared, challenged, chosen, or refused. When a system decides which sources deserve our finite time, it exercises a form of delegated judgment over the material from which our judgment will later be made.
This is the second delegation.
The first delegates retrieval: Find information for me. The second delegates discovery: Decide what is worth my noticing.
Retrieval serves an articulated need. Discovery helps form the need itself. It introduces the unfamiliar concept, the neglected source, the disconfirming result, the person outside the usual network, or the fact that the original question contains a bad assumption wearing a confident hat.
When AI remains a conversational tool, a person can still ask for alternatives, inspect sources, or reformulate the question. Agentic systems add another turn. An agent may search, synthesize, decide, and act without presenting the informational territory to a person at all. Source selection then travels directly into action.
The report is not merely read. The supplier is hired. The treatment is recommended. The grant is rejected. The meeting is never scheduled because the system concluded, somewhere upstream and invisibly, that the idea was not relevant.
At that point discovery is no longer a pleasant side effect of browsing. It is part of the control system.
Should the Machine Surprise Us?
A common answer is to build serendipity into the system. Recommender research often defines a serendipitous item as relevant, novel, and unexpected. Researchers have shown that systems can be optimized for more of it, though gains in serendipity and diversity may trade against narrow measures of accuracy (Kotkov, Veijalainen, and Wang, 2020).
This sounds charming until the machine begins deciding what surprise will be good for us.
If an AI knows what we seek, should it occasionally show us something it believes we were not seeking? How often? How far away from the request? Should a medical search introduce a frightening possibility because it is statistically adjacent? Should a political question include an opposing view even when that view is poorly supported? Should a student receive the efficient explanation or the difficult source that might teach them how the explanation was made?
There is no neutral setting. Maximum relevance can become intellectual enclosure. Maximum surprise can become paternalism with confetti.
Randomness can produce surprise; useful discovery changes the map. It may provide a counterexample, expose an assumption, connect two domains, or reveal that the obvious consensus is only obvious inside one community.
Good discovery therefore requires judgment about relationships between ideas. The system must know not only what resembles the request, but what productively complicates it. Then we must decide whether we trust the system to make that judgment, which returns us to the problem wearing a different badge.
Who Pays for Friction?
It would be easy to romanticize the old path. We could praise the six tabs, the dead links, the badly scanned PDF, and the afternoon lost inside a bibliography. Suffering builds character, particularly in people writing essays about how other people should search.
The costs of friction are unevenly distributed. Fast synthesis can make specialized knowledge more accessible to people with limited time, unfamiliar vocabulary, disabilities, language barriers, or no institutional subscription. A system that explains a field before asking someone to navigate it can expand participation.
The goal is to remove avoidable inconvenience while preserving agency.
People should be able to see where an answer came from, understand where serious disagreement remains, widen the search when the stakes require it, and leave the path the system prepared. A better interface can leave the old web’s obstacles behind while preserving the exits.
Pew Research Center’s analysis of 68,879 Google searches in March 2025 found that users clicked a conventional result on 8 percent of pages containing an AI summary, compared with 15 percent of pages without one. Links inside AI summaries received clicks on only 1 percent of visits (Pew Research Center, 2025). This does not prove that people learned less. It does show that the answer changes whether they enter the underlying web.
The destination has begun replacing the road.
Design for Discovery
If discovery matters, systems should be evaluated for more than whether the answer was correct and the user stopped searching.
They should preserve a visible source trail. They should distinguish broad agreement from unresolved dispute. They should let people choose between precision and breadth, and explain why an unexpected source was introduced. They should preserve direct access to original work rather than treating citations as ceremonial trim.
They should also be audited for exposure. Which sources repeatedly appear? Which communities remain outside the retrieval boundary? Does the system favor institutions already rich in links, authority, and machine-readable prose? Does personalization introduce useful context, or does it turn a person’s past into the wall around their future?
For consequential agentic work, the system should pause when discovery changes the decision. A new source that alters the evidence, introduces a material disagreement, or exposes a hidden assumption creates a judgment point.
These requirements will make some systems slower. They may also make them better. Speed measures performance along one axis; trustworthy judgment requires inspection along several others.
After Search
The future of search will probably contain both lists and answers, exploration and synthesis, machines that retrieve and people who decide when the retrieval has become a worldview. The useful question is not which interface wins. It is which forms of human judgment survive inside the winning interface.
We should not demand that every answer become a seminar. Most questions need an answer. Some need a source. A few need the machine to say that the question is contested, the evidence is thin, or the most relevant thing lies just outside what was asked.
A useful system answers the question.
A humane one helps us notice when the question has made the world too small.
Sources
- Google. “AI Mode in Google Search: Updates from Google I/O 2025.” May 20, 2025.
- Kotkov, D., Veijalainen, J., and Wang, S. “How Does Serendipity Affect Diversity in Recommender Systems? A Serendipity-Oriented Greedy Algorithm.” Computing 102, 393–411 (2020).
- Liu, N. F., Zhang, T., and Liang, P. “Evaluating Verifiability in Generative Search Engines.” Findings of EMNLP 2023, 7001–7025 (2023).
- Pew Research Center. “Do People Click on Links in Google AI Summaries?” July 22, 2025.
- White, R. W., Drucker, S. M., Marchionini, G., Hearst, M., and schraefel, m. c. “Exploratory Search and HCI: Designing and Evaluating Interfaces to Support Exploratory Search Interaction.” CHI 2007 Workshop (2007).
- Yu, Y., Li, Y., Suri, S., and Counts, S. “From Searchable to Non-Searchable: Generative AI and Information Diversity in Online Information Seeking.” arXiv:2604.10258 (2026). Preprint.