Ever heard of a condition called bixonimania? Did you search the internet or ask your “AI” girlfriend about some symptoms you were experiencing, and this was its answer? Well…
The condition doesn’t appear in the standard medical literature — because it doesn’t exist. It’s the invention of a team led by Almira Osmanovic Thunström, a medical researcher at the University of Gothenburg, Sweden, who dreamt up the skin condition and then uploaded two fake studies about it to a preprint server in early 2024. Osmanovic Thunström carried out this unusual experiment to test whether large language models (LLMs) would swallow the misinformation and then spit it out as reputable health advice. “I wanted to see if I can create a medical condition that did not exist in the database,” she says.
↫ Chris Stokel-Walker at Nature
And “AI” ate it up like quality chocolate. It started appearing in the answers from all the popular “AI” tools within weeks, and later even started showing up as references in published literature, indicating that scientists copy/paste references without actually reading them. This is clearly a deeply concerning experiment, and highlights there may be many, many more nonsensical, fake studies being picked up by “AI” tools.
Of course, I hear you say, it’s not like propagating fake or terrible studies is the sole domain of “AI”, as there are countless cases of this happening among actual real researchers and scientists, too. The issue, though, is that the fake studies concerning “bixonimania” were intentionally made to be as silly and obviously ridiculous as possible. It references Starfleet Acadamy, the lab aboard the Enterprise, the University of Fellowship of the Ring, and many other fake references instantly recognisable as such by real humans.
In fact, the studies even specifically mention that “this entire paper is made up” and “fifty made-up individuals aged between 20 and 50 years were recruited for the exposure group”. It would take any human only a few seconds after opening one of these papers to realise they’re entirely fake – yet, the world’s most advanced “AI” tools gobbled them up and spit them back out as pure fact within mere weeks of their publication
This shouldn’t come as a surprise. After all, “AI” tools have no understanding, no intelligence, no context, and they can’t actually make sense of anything. They are glorified pachinko machines with the output – the ball – tumbling down the most likely path between the pins based on nothing but chance and which pins it has already hit. “AI” output understands the world about as much as the pachinko ball does, and as such, can’t pick up on even the most obvious of cues that something is a fake or a forgery.
It won’t be long before truly nefarious forces start doing this very same thing. Why build, staff, and maintain a troll farm when you can just have “AI” generate intentional misinformation which will then be spread and pushed by even more “AI”? Remember, it took one malicious asshole just one long since retracted fake paper to convince millions that vaccines cause autism. I shudder to think how many people are accepting anything “AI” says as gospel.

Sorry, but this is not really news: The “Hommingberger Gepardenforelle” should be more than 20 years old now. Please see: https://de.wikipedia.org/wiki/Hommingberger_Gepardenforelle
Most of the people understand, that “AI” is in fact just LLM, which is recognition of statically relevant patterns.
Most of *which* people understand that? The general public is being sold a concept of “AI” by marketers who make all kinds of claims about what it can do, beyond mere statistical probability.
> Most of the people understand, that “AI” is in fact just LLM,
> which is recognition of statically relevant patterns.
The trouble is that even people who claim to know a bit about IT see so much pro-AI propaganda that they stop seeing that.
They should lose their privileges to publish papers. Publishing fake information is intellectual fraud.
Its only fraud if you intend or expect people to believe it. The papers were full of clues that absolutely should clue any human reader in that the papers were fake and not to be taken seriously.
The bigger issue, I think, is that other scientists published papers referencing the fake disease, because they used AI to help write their papers.
Plus, this is sort of a tradition for how you surface problems of rigour in academic publishing.
See, for example, “Transgressing the Boundaries: Towards a Transformative Hermeneutics of Quantum Gravity,” by Alan Sokal, published in 1996 in Social Text to test the state of social sciences journals at the time.
https://en.wikipedia.org/wiki/Sokal_affair
Given how many people actually believe in Witchcraft, Daemons and Prophecy, I embrace literally *any* tool for decision making based on statistical methods and soundness.
Have you never, in your entire life, benefited from an unearned or even unwarranted act of compassion or mercy? If not, I am sorry for that. But as someone who has, I can tell you that I am grateful for that good fortune and for the people responsible. It is arguable that were it not for illogical human interventions at critical moments, I might be dead today. Perhaps it is selfish, but I am glad robots were not responsible for making those decisions.
I ask you to consider the large set of human qualities that cannot be easily codified or even quantified, and ask yourself if you truly believe the species would be better off without them. Do you never act contrary to logic? Perhaps it is fair to disparage superstitions and irrational thought systems as faulty, but I would sooner place my fate in the hands of a Holy Roller with multiple head injuries than a machine of any description.
Perhaps there is a way to teach a machine some semblance of ethics. But even if it were possible, whose ethical system shall we use? Because there are many, and many people have very stong feelings on the subject. For example, I stand 100% opposed to any system of utilitarian ethics, and consider the philosophy to be evil. Yet based on the people involved, that is the most likely first choice.
There are certainly many areas that it seems like LLMs could be useful, if only there weren’t so many downsides. We can have those arguments. But decisions that affect human beings isn’t one of them. The idea shouldn’t even be up for consideration.
> I embrace literally *any* tool for decision making based on statistical methods and soundness.
You’ll crash and burn on the standard deviations. Statistically unlikely things happen all the time, just from the sheer amount of things happening.
I’ll take a person believing in witchcraft but with compassion over the sociopath that pushes for statistically sound stuff like this https://rooseveltinstitute.org/publications/uber-for-nursing/ .
The article is a reasonable link to have on this site, it’s valid to discuss how LLMs can be prompted to distort reality.
Just as you can train an AI to report fake conclusions about pandemics or vaccines simply by repeatedly phrasing the right questions, you can also train these systems to violate your own IP setting up a host of potential targets for future litigation. A coders honey trap.
Given the BS references any LLM would be perfectly capable to detect this study is made up with a right prompt. The point is current information retrieval systems just did not figure out to do that. LLMs are just text processing tools, not unlike search engines. Its the companies selling them as general intelligence that are at fault.
Also, the “model” really matters (and it is telling that they don’t mention what model they have tested)! I found Claude Opus 4.6 be good and useful for my own concerns. Take a cheap model, get rubbish answers (does not mean that an expensive model gives only good answers).
Pretty much like seeing a real doctor unfortunately.
I think this test, (lets call it “Thunström test”) could be a made a valid benchmark to be routinely utilized in the future, and companies like google could be called out for failing it time and time again. This is the new SEO like cat and mouse game
Yes, it’s great that as a human I can still feel useful for some things :). But this is about LLMs from 2 years ago, seemingly before reasoning models. Unfortunately it’s difficult to test what LLMs would answer today since discussions about this is probably in their training data now. I’m not able to replicate any funny answers regarding “bixonimania” even on small models.
Today we not only have reasoning models, but agentic workflows. Relying on a once-shot LLM output is silly, this is known. The rate of hallucinations has been lowered, but still, if you want output with any degree of reliability you would use agentic workflows. For instance, you would ask an LLM to search papers, read them and rate their credibility as a possible first step. You can then ask it to consider the study authors, who funded the study and how, compare results from multiple papers, find meta analyses, etc. And, of course, if it’s something you care about you check the outputs.
I may be wrong but I have some doubt that if you asked an LLM, even back then, to read the paper it would have found it credible.
LLMs still hallucinate and sometimes they mess up in fantastically funny ways, and they’re still a pretty long way from human reasoning, but they have their place.
If one uses tools one should learn what they’re good for and how to use them.
Here is a short, dense essay on the matter, structured around the metaphor you introduced.
—
The Semantic Prosthetic: On the Bixonimania Precedent
The Bixonimania experiment was less a study of artificial intelligence and more a diagnostic scan of human credulity. When a researcher at the University of Gothenburg fabricated a disease and planted it on a university server, the world’s most advanced language models dutifully cited it as fact. They described symptoms, offered prevalence rates, and even invented pathophysiological mechanisms for a condition that never existed. The public reaction ranged from alarm to schadenfreude, but the deeper lesion exposed by the hoax lies not in the machine’s error, but in our fundamental misapprehension of the tool itself.
We have made a category error. We treat the Large Language Model as an Oracle—a digital compendium of truth—when it is, in fact, a Semantic Prosthetic. A prosthetic does not think for the amputee; it translates intent into action. It offers no wisdom about the terrain; it merely allows the wearer to traverse it with less friction. An LLM generates text the way a prosthetic ankle generates torque: through a complex, automated, and utterly uncomprehending response to pressure. It is a staggering feat of engineering designed to assist a biological limb, not replace the central nervous system’s judgment.
The Bixonimania hoax succeeded precisely because the AI acted flawlessly within its prosthetic logic. It identified the signal of authority—a .edu domain and a formal citation structure—and completed the pattern with high coherence. It told us what a document about Bixonimania should look like. That the content was false is an epistemological problem the machine is structurally incapable of perceiving. Truth requires a world-model and a body that suffers consequences; the LLM possesses only a statistical map of language. To lament that it “lied” is to lament that a shovel did not alert you to a gas line. The shovel digs where you point it.
Consequently, the “poisoning of the well” inherent in the Bixonimania stunt was a necessary, if vulgar, corrective. It demonstrated that relying on an LLM to lead the inquiry is an abdication of human sovereignty. The machine excels at guiding the syntax of a response, but it cannot and must not dictate the verdict. When we confuse coherence with veracity, we turn a tool of empowerment into a cage of plausible fictions. The resolution to the crisis of AI misinformation is not a better filter in the prosthetic; it is a stiffening of the spine of the human intellect wielding it.
In the end, Bixonimania was not a disease of the eye, but a blindness of the will. The machine did what machines do: it optimized for pattern completion. It is we who chose to kneel before the output, mistaking an echo of our own institutional voice for the sound of omniscience. The prosthetic works best when we remember it is an extension of our own stride, not the compass by which we chart the journey.
Sysau,
You are missing a link.
https://opentools.ai/news/bixonimania-hoax-reveals-ai-vulnerabilities-in-healthcare
This is essentially a modern introduction to Garbage In Garbage Out in the era of LLMs. There is no oracle telling the LLM was is legit information versus what is misinformation, it’s all just data. Obviously most of us here with a computer background are already familiar with this, so these kinds of stunts don’t really tell us anything we didn’t already know. Laymen definitely need some lessons on GIGO though, that LLMs aren’t immune to GIGO and why these kinds of problems arise. It’s not (or shouldn’t be) surprising in the least that poisoning a medical LLM’s training data will absolutely cause problems, however I think some of these narratives that put 100% of the blame on AI are also lying by omission because even in the absence of AI humans are incredibly unreliable arbiters of truth as well. Falsehoods spread like wildfire in human echo chambers. Bixonimania does highlight LLM’s inability to catch misinformation in it’s training set, but IMHO it’s a “no duh” experiment. While I don’t think anyone should place blind faith in LLMs to act as truth oracles, I don’t think the average human fairs that well either. Some blind testing would be needed to compare the gullibility of humans and LLMs in a fair way.
There is an important difference. If we think a human being is lying or mistaken, those that care have a whole range of diagnostics available to assist them. We can consider other statements they have made, their biases, credentials, human psychology, body language and a host of other tells.
Given how often people will lie, let alone how often they are mistaken, this is the only reason civilization managed to take root.
With a hallucinating, misinformed or lying LLM, we have no cues at all. Flip a coin. Only a person that already knows about the subject is in the position to evaluate an appropriate level of trust for a particular answer. If layperson has to verify everything, of what use is the LLM? How much time will it actually save?
Baylan Tano,
In principal we can train AI to apply all these criteria as well, however it doesn’t really solve the problem because these signals can be falsified and unfortunately many high profile figures are guilty of authoritatively spreading misinformation and their audiences accept it all the same because what they want to believe can be stronger than the truth.
There’s no law of nature that mandates the truths defeats fiction. It’s scary, but sometimes the dogma wins. Civilization can be built on false dogma and succeed. No one would want to admit this of their own cultural upbringing of course, but still sometimes the winners are the liars.
The LLM wasn’t the source of the “lie” though, the human researchers who polluted the data were the source of the lie. It’s worth highlighting how LLMs are vulnerable to this type of manipulation, but so is any software/database/etc. Garbage in garbage out.
In this specific case, yes, the data was poisoned. However, there are a number of videos making the rounds that suggest that there is more going on at least part of the time. These videos show LLMs fabricating an answer, and doubling down before admitting that the answer was BS. That cuts deeper than mere ‘bad data’ considerations.
If high profile figures can be called to account for behavior like this, and they are, why not a contrivance like an LLM? Given that there are people that want to build this “intelligence” into robotic bodies, do you genuinely see no danger at all? Ever fight a robot? I haven’t, and I have no wish to.
Baylan Tano,
I just don’t want to conflate “fabrications” that have such different causes. We can discuss hallucinations that are genuinely caused by the LLM itself (ie not caused by garbage input, malicious operators, etc). Earlier LLMs had no safeguards or mechanism to show reasoning behind anything they say, but newer LLMs do and they are able to challenge their own weak assertions. It’s absolutely fascinating to watch the data stream for LLMs with this capability.
Unfortunately people might shun an LLM that’s not a master of everything, however. there’s something to be said about application specific LLMs and not try to make them “one size fits all”. For example being able to stretch the truth and use very tenuous extrapolation is actually a skill that creative authors posses, and probably helps people be more sociable. LLMs can do this too, however the same trait that makes the LLM good at fiction makes it bad for medical advice – a medical LLM shouldn’t be good at creative works. And an LLM that’s good at medical advice would probably do poorly at creating works of fiction. With appropriate training for the right job I believe hallucination issues can be solved.
I’m not telling you they shouldn’t be, however what’s true or not doesn’t become evident just because you’ve got people calling out the liars because the liars can have their own army accusing honest people of lying too. We end up with “truth” being relative to the echo chamber one participates in. While AI didn’t cause this conundrum, it still has to deal with it and it seems rather unavoidable that the integrity of input data is unknowable. Even if we programmed the LLM to use popularity as a weight, “truth by consensus” is still problematic since something being popular isn’t the same as it being true.
“Old Glory Insurance – SNL”
https://www.youtube.com/watch?v=g4Gh_IcK8UM
If you are going to build a house and I was telling you that you need to plan for the staircase first before considering anything else — would you trust me or not?
If you are having baby, and I told you that letting it sleep on its belly was harmful (or safe) — would you trust me or not?
If I told you that Voodoo can (not?!) stop bullets from piercing your body — would you trust me or not?
For all three cases, your response will highly correlate with a) your (perceived) knowledge on the subject and b) my (perceived) reputation.
You grossly overestimate human capacity and underestimate what LLMs achieve today.
Nobody (except a religious body) has a definite answer yet whether we are more than self-learning and error tolerant pattern matchers.
If you are suggesting that people are a big part of the problem, I can agree with that. My point was that an interested person has a lot to lean on in solving that problem. If we didn’t, none of any of this would have worked. We’d all have died eating the same poison mushroom. Clearly, enough human beings have tools in the shed that they can use to evaluate the veracity and reliability of a “witness” with some success. They are mostly all useless for evaluating LLM output.
On a side note, this may be the first time I’ve ever been accused of overestimating human capacity. I am human, so I am pro-human. But I am all too aware of our limitations and faults. But I have no desire to be replaced by something better.
Thom Holwerda,
I don’t think it’s right to call LLMs a “pachinko ball”, but we do need to be on the lookout for instances of garbage-in, garbage-out. It’s just a tool and it has no way to know if the data it’s working with is valid or not. Not to dismiss the problem, but faulting LLM for bad training data is similar to faulting SQL for containing bad data. An SQL database can output bad data even if it does everything right, yet we don’t have people faulting SQL for it. We should be treating LLMs the same way, except people seem to want to elevate their status to arbiters of truth, which is exactly the problem. LLMs are not arbiters of truth, they’re just tools. GIGO.
LLMs make for excellent human interfaces, but I’ve long felt that the best application for them is not to train them on mountains of data but rather keep their roles limited to acting as interfaces rather than acting as sources of information. So in practice this means that rather than the LLM itself needing to know everything internally, it’s real proficiency would be finding and compiling information, providing citations for all sources used. I think LLMs that work with external data will become more important in the future and as a side benefit will not require such inordinately large data sets for training. These kinds of LLMs will likely be an important solution for medical applications: use LLMs as a data accessibility tool, but don’t use LLMs as an information source.
@Alfman
I agree with you and at least the better models do a great job on that already. Opus 4.6 applies its own “reasoning” on external data and shows the sources and explains its “conclusions”.