What happens when a thousand independent sources turn out to have one parent?
A while back I went looking for a specific piece of television. A 2012 late-night interview with a sitting president, forty-five minutes long, broadcast on a major network to several million people.
Finding out about it was trivial. An episode database has the record: season, episode number, air date, runtime, and a summary of what was discussed. Wire coverage exists. Clips exist. A national newspaper posted the complete video the following morning, and the URL for that page is still indexed. Contemporary articles quote it. Later articles quote those articles.
Finding the thing itself was considerably harder.
A version of this essay treats that as sinister. This is not that essay. Broadcast rights change hands. Video platforms get retired. Formats go obsolete. Nobody has to intend anything for a forty-five-minute artifact to become difficult to inspect while everything written about it remains a search away.
What interests me is the shape that leaves behind, because I think it is becoming the normal shape of our information environment:
What happens when the source disappears but everything derived from it remains?
The memory hole doesn't have to be empty
We tend to imagine information loss as absence. A document disappears, a database is deleted, a recording is destroyed. Something that existed no longer does.
Modern information systems produce a stranger failure mode. The original can vanish while its descendants multiply.
Picture a primary artifact that generates ten contemporary news stories. Another hundred articles cite those stories. Wikipedia summarizes several of them. Blog posts cite Wikipedia. Podcasts discuss the blog posts. Social posts quote the podcasts. Years later, AI systems ingest some combination of it all.
The ecosystem now contains thousands of references to an artifact almost nobody can examine. Retrieval works. Search works. There may be enormous agreement about what the original contained. But something has quietly changed underneath all that agreement.
The system hasn't forgotten the story. It has forgotten how to prove the story.
That leads to the claim this whole essay rests on, so I will state it once, plainly, before dressing it up in examples:
Document count is a terrible proxy for evidentiary independence.
417 sources can't all be wrong, right?
Suppose an AI system is answering a question about a disputed event and finds 458 relevant documents. Of those, 417 support one interpretation and 41 support another. The tempting conclusion writes itself.
417 > 41
But documents aren't votes.
Suppose 290 of those 417 ultimately trace back to the same wire-service report. Another 92 descend from the same organizational statement. The remaining 35 cite one another. Meanwhile, the 41 documents supporting the competing interpretation include several independent primary sources.
The interesting question was never how many documents agree. It is how many independent provenance chains support the claim.
The web is exceptionally good at copying information, which is precisely why counting copies tells you so little. Generative AI sharpens the problem, because the final answer collapses hundreds of derivative sources into one confident paragraph. The reader sees consensus without seeing the genealogy that produced it.
This is not a thought experiment, and you can check it yourself in about a minute. Ask an answer engine a general knowledge question and look at what it cites. A crowd-edited encyclopedia will turn up more often than you might expect. That encyclopedia is a tertiary source: a summary of secondary reporting about primary artifacts. When it appears in a citation list, nothing in the interface mentions that the chain already runs three deep before it reaches anything anyone actually witnessed.
Enter the Golden Country Tire Company
Real disputes carry emotional freight, so let's use tires.
Imagine the Golden Country Tire Company is the world's largest tire manufacturer. Golden Country has just released its flagship product, the Super-Duper Road Tire. It's fine—perfectly adequate tire. Golden Country would nevertheless very much like the world's humans, search engines, and AI systems to regard it as one of the finest achievements in the history of vulcanized rubber. The company and every site named below are invented. The structure is not exotic.
So Golden Country does what any competent marketing organization does. It creates genuinely good content: technical documentation, comparison pages, FAQs, buying guides, structured data, product specifications, expert commentary, and articles answering every question a person might plausibly ask about road tires.
From a GEO and AEO standpoint, Golden Country is doing its job well.
Then Golden Country goes further and funds or controls a collection of apparently independent sites:
RoadTireExperts.example
UltimateDrivingGuide.example
TirePerformanceLab.example
BestRoadTiresToday.example
DefinitelyNotGoldenCountry.example
Each publishes high-quality, well-structured, machine-readable content. Each concludes that the Super-Duper Road Tire is fantastic.
Now ask an answer engine which tires are best for highway driving.
Retrieval surfaces dozens of sources praising the Super-Duper Road Tire. The model isn't hallucinating. The documents exist. The recommendations exist. The citations exist.
Five sources are not five independent sources if Golden Country is standing behind all five of them.
Golden Country may not have fabricated a single claim. Every individual statement might be technically defensible. What Golden Country manufactured is not a falsehood. It is the appearance of consensus.
An answer engine that understands URLs sees five sources. An answer engine that understands provenance sees one organization speaking through five domain names. Those are very different information environments, and nothing in the retrieval layer distinguishes them.
When the copies start citing one another
The problem gets more interesting once Golden Country's ecosystem develops internal links.
RoadTireExperts.example publishes a review calling the tire exceptional, citing a braking-distance comparison from TirePerformanceLab.example. The lab article points to a roundup at UltimateDrivingGuide.example. That roundup cites customer-satisfaction figures summarized by BestRoadTiresToday.example, which links back to the original Road Tire Experts review.
From the outside, the provenance graph looks rich:
Multiple domains. Multiple articles. Multiple authors. Multiple citations. Apparent corroboration throughout.
The graph is a circle. No independent evidence ever entered the system. The sources don't corroborate one another. They are recursively laundering the same claim.
Here is the same graph with one more fact restored:
The dashed edges are the only thing that changed, and they are the only thing that matters. They are also the only part of this picture that no retrieval system draws, because nothing in a URL, a byline, a schema block, or a citation announces who funded the page.
This is where counting citations becomes as misleading as counting documents. A densely connected graph can look authoritative while having almost no independent roots. If every path eventually terminates at Golden Country, the graph contains repetition, not corroboration.
None of this is new. Human information ecosystems have always contained circular citation, press-release recycling, unattributed copying, and claims that gain acceptance through sheer repetition.
What changes with generative AI is the economics. Another plausible article is cheap. Another plausible site is nearly as cheap. Rephrasing a claim so it reads as linguistically independent is cheap. Producing structured, answer-friendly content at volume is cheap.
Apparent consensus can now grow much faster than independent evidence.
I could give you a number here. A widely circulated figure estimates how much of the newly published web is now AI-generated, and I have seen it quoted in a dozen places this month. I went looking for where it came from. The first article cited a second article. The second cited a marketing blog. The marketing blog cited a crawl study, described but not linked. I gave up at the fourth hop, which is either a failure of diligence on my part or the entire thesis of this essay demonstrating itself at my expense. Possibly both.
So take the number as read, and notice instead that I cannot show you its parents.
And synthetic content no longer needs to copy the original wording. Fifty pages can express the same unsupported claim fifty different ways. Textual similarity becomes a weaker signal of shared ancestry, even though the underlying provenance hasn't changed at all.
The result is an information environment optimized beautifully for retrieval and architecturally terrible for verification.
But doesn't somebody catch this?
The reasonable objection is that platforms already police this. They do, and they do it reasonably well. Search engines have spent years developing policies against scaled content abuse and coordinated networks built to manipulate rankings, and enforcement actions have removed entire sites from indexes.
Notice what those policies target: low-quality content produced at volume, and thin content built to game a ranking. That is the crude version of Golden Country, and the crude version does get caught.
Our Golden Country doesn't do that. Its technical documentation is accurate. Its comparison pages are useful. Its specifications are correct. Its structured data is well-formed. Every site in the network would survive a quality review on its own merits, because every site deserves to.
Golden Country is not violating the spam policy. It is following the content marketing playbook competently, five times, from five domains it happens to own. The enforcement regime was built to detect garbage, and Golden Country isn't producing garbage. It is producing a well-made monoculture.
That deserves a name, because it will keep happening. Call it a provenance monoculture: an information environment that is diverse in sources, formats, and domains, and uniform in origin. Nothing in it is false. Nothing in it is thin. Everything in it grew from the same root.
That is the gap. Quality enforcement and independence verification are different problems, and we currently have infrastructure for one.
Nobody needs a pneumatic tube to the furnace
In 1984, controlling history requires destroying evidence. Winston Smith rewrites the record and the original goes down the memory hole.
Our systems don't require anything that dramatic. A primary source becomes gradually inaccessible. Links rot. Licensing changes. Platforms retire. Archives migrate. Formats go obsolete. Meanwhile the derivative material stays exactly where it is, and new material keeps accumulating around one interpretation of the missing source.
Nothing has to be deleted on purpose. Nothing has to be centrally coordinated. The environment simply becomes asymmetric, and the systems grounded on that environment inherit the asymmetry.
A modern memory hole is surrounded by more information than ever.
The economics of the hole
GEO and AEO are usually discussed as marketing disciplines: make your organization, product, expertise, or terminology retrievable and comprehensible to answer engines. That is legitimate work, and good technical content should be understandable by humans, search engines, and answer engines alike.
The problem starts when information availability gets confused with independent corroboration.
When a primary source is missing, something determines which secondary representation becomes its machine-readable substitute. An organization with sufficient resources can produce a large body of coherent, optimized material around its preferred representation of reality. It doesn't need to falsify anything. It only needs to become disproportionately represented in the environment from which answers get assembled.
The Super-Duper Road Tire doesn't become better. It becomes better represented.
If answer engines treat frequency as confidence, domain count as independence, citation density as authority, or repetition as corroboration, then the organizations best equipped to populate the environment gain an advantage with no relationship whatsoever to the quality of their evidence.
There is a further wrinkle, and it's also easy to test. Put the same question to two different answer engines and compare the lists of sources underneath. The answers will often agree. The evidence behind them frequently does not overlap much at all. Whatever consensus a reader perceives is partly an artifact of which pipe they happened to ask.
When the source is missing, say so
This is where provenance stops being an archival concern and becomes part of memory architecture.
A trustworthy system should distinguish among primary evidence, independent corroboration, derivative reporting, organizational claims, unknown provenance, and unavailable primary sources. Those are not equivalent categories of knowledge, and collapsing them is a design decision, not a technical necessity.
If 417 documents descend from three sources, the system should know that. If five apparently independent tire sites belong to Golden Country, the system should know that too. If a citation graph contains no independent evidentiary root, the number of edges in the graph should not manufacture authority.
And if a source once existed but can no longer be examined, the system should preserve that fact rather than silently filling the gap with the statistical weight of everything surrounding it. Absence is itself a provenance category. It is a thing worth recording, not a hole to be smoothed over.
A provenance-aware system might answer like this:
Multiple secondary sources report this claim, but the primary artifact they reference is unavailable. Several of those sources also derive from the same upstream reporting, so they should not be treated as independent corroboration.
That is not a weaker answer. It is a more honest one, and honest answers are the only kind worth building infrastructure for.
Information without provenance is just gossip
Memory is not simply the ability to preserve information. Trustworthy memory preserves the relationship between information and its origins.
Who created this? What evidence supported it? Was the source primary or derivative? Was it independent? Can that authority still be verified? Does this source depend on another that no longer exists? Are apparently independent sources controlled by the same organization? Does the citation graph lead outward to evidence, or eventually curl back onto itself?
Without those relationships, a system can accumulate an extraordinary volume of knowledge while gradually losing the ability to explain why any of it should be believed. That isn't memory. That's a very well-indexed rumor mill.
Orwell imagined that controlling history required destroying the evidence. Our problem is subtler and considerably cheaper. We can preserve enormous quantities of information while losing the provenance required to evaluate it, and we can surround a missing source with so many summaries, restatements, and synthetic corroborations that the absence itself becomes invisible.
Increasingly, machines stand between that environment and the person asking the question. So the question is no longer whether a system can find an answer. It is whether the system can tell a thousand independent witnesses from one witness repeated a thousand times.
Because when a source falls into a memory hole, something always fills the space around it. Information without provenance is just gossip, and gossip scales beautifully.
An open question, and I mean it as one.
I have described a problem and stopped short of a fix, because I am not sure what the first move is. Disclosure obligations for funded networks? An independence signal carried alongside citations? Answer engines surfacing shared upstream sources when they detect them? Something else entirely, or nothing, because the incentives point the other way?
If you build retrieval systems, work in trust and safety, or just have a view: what would a first step actually look like, and who is positioned to take it?


Top comments (11)
Until AI architectures combine statistical fluency with deterministic, rule-based logic and explicit provenance, human judgment remains the only reliable filter against well-packaged nonsense.
Frequency is not confidence.
History isn't truth. Same watermark mistake.
It's a mistake cycle in life, the loop cycle can never be eradicated. Nuance will always exist.
I think the "frequency is not confidence" distinction is the important one. I'd add that human judgment isn't entirely immune to the same problem either. We see ten apparently independent sources making the same claim and naturally give that more weight than one source, unless we know those ten all descend from the same origin.
That's why I keep coming back to provenance as infrastructure, not just another quality signal. Neither a human nor a model can properly evaluate independence if the relationships between sources are lost.
We design architectures around our cognitive defaults: we mistake repetition for truth, confuse volume with consensus, and build storage systems that prioritize rapid recall over tracing origins. When a thousand derivative nodes echo a single unverified parent, the system isn't failing—it is executing our default human heuristic at machine scale.
Building durable systems requires moving past statistical retrieval to treat provenance as a first-class execution boundary like you rightly said 👍
Great article! Honestly, some of these problems existed long before AI. For example, the number of articles saying the same thing doesn't equal independent evidence.
And it doesn't even have to be deliberate. Someone writes about something, their article spreads, then 100 other people write about the same thing. Who can prove they didn't all come up with it independently?
I recently had exactly this situation with my article about the Claude watermark. A few days after I published it, a very popular Polish tech website covered the same topic, with a structure surprisingly similar to mine. Do I have any proof that they were “inspired” by my article? Of course not! 😄
But here's the funny part: if they'd actually done the research independently, they would have noticed that Claude had updated its documentation in the meantime. Instead, they still wrote that “we don't know anything yet” 😂
Because why check the latest sources when you can just copy? 😄
That's a great example, particularly because of the outdated documentation detail. You may never be able to establish whether the later article derived from yours, but the fact that it preserved the same outdated assumption is at least an interesting provenance signal.
And I completely agree that none of this started with AI. Humans have been copying, summarizing, syndicating, and independently rediscovering the same ideas for as long as we've published things. What AI changes is the scale and economics of producing and synthesizing those derivatives.
I think your example also gets at why provenance is harder than simply building a citation graph. Two articles can look independent while sharing an unrecorded ancestor, and two genuinely independent articles can arrive at nearly identical conclusions. The graph can only tell us about relationships we actually know.
That's why I think the system sometimes has to preserve uncertainty too. "These appear to be independent sources" is a different claim from "these are independently derived sources." We shouldn't manufacture provenance certainty any more than we should manufacture factual certainty.
And yes, checking the current documentation apparently remains optional. 😄
The graph-is-a-circle point is the one that should stick. A dense citation graph looks like corroboration and codes as repetition the moment you trace the edges, but nothing in the retrieval layer draws those edges, and nothing incentives the answer engine to. What I keep coming back to is that this is a scoring problem, not a knowledge problem. You could compute effective sample size for a claim the way you compute it for a portfolio of correlated assets, weight each source by how much of its variance is shared with its parents, and a 417-to-41 "consensus" collapses to maybe four independent chains. The hard part isn't the math; it's that provenance has to be recoverable at query time, and right now the interface hands you a flat list and calls it evidence.
I really like the effective sample size analogy. That's a much more precise way of describing what I was getting at with the 417:41 example. The raw document count says 417 independent observations, while the provenance graph might reveal an effective sample size of four.
I think your last point is the architectural catch, though. You can't calculate that at query time if the system discarded source ancestry at ingestion. A flat retrieval result has already collapsed "417 documents from four provenance chains" into "417 documents."
That's increasingly where I land on memory architecture in general: provenance has to survive the write boundary if you expect to reason about authority later. Retrieval can't reconstruct relationships the system never preserved.
There's probably an interesting scoring model hiding in your portfolio analogy too: relevance score on one axis, provenance independence on another, rather than treating more matching documents as inherently stronger evidence.
This has a practical engineering lesson: validate the system at the real boundary, not only where the code looks clean. Integration inputs, permissions, retries, and state transitions are where the expensive surprises tend to hide.
This exact problem showed up on a supplier risk project: we had what looked like a dozen independent sources on one manufacturer, but tracing the chain they'd all started from the same press release, just reformatted by a couple of data vendors along the way. Each passed quality checks fine on its own so there was nothing to flag at ingest, and we only found out when we had to explain a call to a client and realized we couldn't actually justify the confidence. If there'd been something above the citation layer that knew those sources were all owned by the same vendor, it'd have collapsed the false independence before any retrieval happened. Honestly surprised no one's built that into a standard system yet, feels like the obvious next thing.
That's a fantastic real-world example of the failure mode. Twelve individually credible sources can pass every source-quality check and still collectively represent one piece of evidence.
What really stands out in your example is that the problem only became visible when someone asked you to explain the decision. Retrieval had worked. The quality checks had worked. The system had plenty of evidence. The provenance behind the confidence couldn't survive examination.
I'm increasingly convinced that's the missing layer. We spend a lot of effort scoring documents individually, but much less asking whether a collection of documents actually represents independent evidence. Ownership is one signal, but derivation matters too. Twelve separately owned publications can still trace back to the same press release.
That's essentially what I was trying to capture with "provenance monoculture": diversity at the document/domain layer hiding uniformity at the origin layer.
Maybe we can manage provenance. Not by solving provenance globally, but by making provenance loss impossible inside the system's own state-transition boundary.
And that shouldn't only apply when the root disappears. Provenance can be unknown in either direction. We may know what produced a claim but no longer have access to it, or we may have the source but have no reliable record of what produced it.
In either case, the system should preserve the boundary honestly:
source unavailable
origin unknown
rather than inventing a relationship to make the chain look complete.
The goal isn't perfect provenance. It's never pretending that missing provenance is known provenance.