The Concentration of Knowledge
Knowledge is concentrated in the hands of a few global players based on the biggest theft in human history where they took knowledge from millions of decentralized websites and redirected the users to their own websites. Thus, there is a serious risk for other knowledge providers on the internet to become unprofitable and thereby unsustainable.
Ever since Chat GPT, Claude and friends were made publicly available a few years back, critics have been concerned that our ability to learn and evaluate may decline once we outsource more and more of our thinking and knowing to the use of LLMs. Here on Dataethics.EU1, Olivia Heslinga has presented some of the first studies on the cognitive impact of outsourcing and offloading critical thinking skills. In short, she outlines how extended AI use over time leads to cognitive debt. Elisa Caeli discusses how the spread of AI content makes even more difficult for young people to assess the reliability of information2. Even though it is still early days both in terms of using LLMs and research on the effects thereof, the first indications are not promising for the human mind.
My impression is, however, that another potential effect of the increasing reliance of individuals and companies on chatbots for answers to all sorts of questions, is not discussed just as frequently. I am talking about the concentration of knowledge in the hands of a few global players. The companies behind the large LLMs, predominantly from the United States and China.
What do I mean about the concentration of knowledge?
Well, in short, LLM companies have trained their models on knowledge obtained from the biggest theft in human history, i.e. whatever sources were available on the internet. They have sourced data predominantly without consent or reimbursement of the copyright holders or authors of the knowledge that they have fed into their LLM’s. That is why the training process of LLM’s amounts to the biggest theft in human history. OpenAI, Anthropic etc have grabbed knowledge that had been decentralised to millions of websites, publishing platforms, digital libraries and so on to absorb the knowledge and provide knowledge centralised through their chatbots.
That is, by the way, what makes chatbots so useful for their users. They are a one stop for knowledge. They remove the friction of having to research something on different webpages and the friction of having to search for something useful through a couple of pages of a google search or an academic database. Just type one line of command. Voilà.
But is that not great?
Well.
First of all, the AI-companies have made the compound human knowledge their product. They sell us something that has never been their own. They are basically the people that re-sell stolen bikes on online platforms. But that is not the core of the knowledge-concentration problem.
Since LLM’s got popular, the Danish health platform sundhed.dk, a platform that provides knowledge on illnesses and health more broadly written by authorised doctors and researchers has lost one third of its visitors. The number dropped from 1.5 million to 1 million visitors per month. For sundhed.dk that is not an economic problem, because it is publicly funded.
But sundhed.dk is by far not the only website experiencing this problem. And many websites do earn money through their visitors, either through advertising or purchases.
Chatbots redirect traffic from websites to, well, chatbots and thereby strip companies and individuals behind many websites of their reward.
You might argue that this is just the force of economic competition doing its deed.
But that is untrue. It is rigged competition, because LLMs have stolen the knowledge from these providers, redirecting the users to themselves, stripping the knowledge providers of visitors and income.
The long-term consequence of this trend might be that it becomes economically unattractive if not downright impossible to be a knowledge provider on the internet. Greater-good initiatives like Wikipedia or a publicly funded websites like sundhed.dk might prevail, but for others it might not be attractive enough any longer to run websites. Even public websites might have to close at some point, because it could become hard to justify to pour tax payer money into websites that are visited by a steadily declining user base.
Why is that a problem? I mean we still have access to knowledge through chatbots, and it is easier than googling!
Probably. But at that point, we have monopolised (well, oligopolised) human knowledge primarily in the hands of companies that are originating from China and the US. Two countries, that at least from a Western European perspective might not be well-suited as the guardian of human knowledge for obvious reasons.
Philosophers like Michael Foucault knew that knowledge is power. Nowadays one might say: Knowledge is data, and data is both money and power.
If we give away knowledge to LLM-companies, (or rather: if we let them get away with their theft), while simultaneously giving up other knowledge structures, we give these corporations massive power. Some might argue that this has already been the case with Google’s search engine. That is true to some extent, but Google “only” has a monopoly on presenting knowledge. OpenAI, Anthropic, DeepSeek and the other LLM companies now (illegally) own the knowledge. We are one decisive step further.
Recently, I have frequently heard the argument that my worries about knowledge concentration are ill-founded. The argument goes like this: First of all, LLM companies have an interest in providing the knowledge to us, because that is what they sell to us, i.e. they need to provide knowledge to earn money. Furthermore, we will always have alternative sources of knowledge, for example books or of course, the internet. Nobody can take that away from us.
It is true that LLM companies must keep selling their knowledge to earn money. It would be stupid for them to shut down access to their knowledge, since that (for now) is the product they sell. But since the large corporations are based in China and the US, we must not ignore the risk of sanctions, for example in the realistic scenario of a NATO-internal conflict due to a US-takeover of Greenland. You cannot know for sure, as long as Trump and his allies are in power. The same goes for China, for different potential conflicts.
Furthermore, we must be very much aware of the risk that the knowledge presented to us is skewed. There has been much talk about the inherent bias of LLM’s, for instance racist and sexist biases. But if global powers like the US and China are controlling how we access knowledge, they can also control the knowledge that we get.
Imagine, again, the scenario of a NATO-internal conflict due to a US-takeover of Greenland. We know how intertwined US-tech companies like OpenAI are with the Trump government. I can easily imagine how a user searching for the history of Greenland on ChatGPT will no longer be presented to historic facts but to some kind of fake news that justifies the US takeover of Greenland.
In short: LLM’s open up for a new level of disinformation campaigns.
We might still have alternative sources of knowledge, for example books. I do not believe that the book-shredding by AI-companies will be enough to erase human knowledge that way around (I hope I will not be proven wrong here). But as I outlined above, there is a serious risk for other knowledge providers on the internet to become unprofitable and thereby unsustainable.
At that point, LLM-use will have damaged and reduced available alternatives to finding knowledge.
If LLM companies control what we know, they can also easily influence what we believe to be facts, what we like and do not like, what we buy, how we act.
We have to resist by insisting on our own human agency, by not letting ourselves become dependent on ChatGPT, Claude and friends for every little search. We have to insist on political solutions, and we could start by insisting that the biggest theft in human history must not go unpunished.
If someone steals your bike, or breaks into your house, you would demand a punishment for the thief. Just because the thief now provides a nice and seamless service, we must not forget the crime, and we must not let the crime go unpunished.
As radical as it might sound: In the beginning of this century, platforms where copyright-protected music and videos were distributed illegally were massively chased by law-enforcement internationally, and many platforms were ultimately shut down.
To be honest, I cannot see, how LLM’s are any different.
To borrow a little Foucault once again: Knowledge must be defendeed!
Photo: Pixabay, royalty-free
1 https://dataethics.eu/the-price-you-pay-for-ai-convenience-is-cognitive-debt/
2 https://dataethics.eu/young-peoples-ability-to-assess-reliability-of-information-is-decreasing/