Fairy tales.
"If a capable agentic AI tool were available".
If wishes were horses ...
AI hype is colliding with reality yet again. Wiley's global survey of researchers finds more of them using the tech than ever, and fewer convinced it's up to the job. Academic publisher Wiley noted a link between researchers' use of AI and their confidence in its ability to perform as well as a human in a preview of its second …
The problem is that not all knowledge has been codified as text. Or can be.
Example 1: weather forecasts. They are often wrong even without AI. (But still useful.) Weather is also ill-defined because specific location is granular ("rain at a specific address"), but forecasts are region-based ("rain in a city").
Example 2: configuring a router. Documentation is missing online. Experimentation with a concrete device is necessary to figure out what works. The router mechanism itself is codified knowledge (but not as text).
Example 3: repairing a washing machine. Similar to weather forecasts there could be hundreds granular reasons for a break. Until you open the machine and investigate, the specific reason cannot be known.
Example 4: driving a car. An "impossible scenario" could be "a highly unlikely scenario". Like a meteorite falling on the road. Not all possible conditions are in the database.
Example 5: making a good pizza. Manual skills and physical strength are necessary to spread the dough. A perfect text description is not enough.
Example 6: refueling a car in year 1736. Actual petrol to refuel a car is necessary. You cannot refuel a car without a petrol station in the desert. Or specific petrol type.
Example 7: making a CPU. Good luck without a machine made by ASML and a knowledgeable technologist operating the machine. The machine itself is knowledge codified in its parts.
Books are compressed descriptions of reality. Many are subjective, too generalizing, or plain wrong. Scientific articles are relatively new in the history of humanity.
So expecting the AI to "know it all" is unrealistic. It is application-specific. Knowledge without physical or technical prerequisites is useless.
I'd add:
There's often conflicting information. Especially when it comes to history. It's not just the winner that writes the history: It's any of the survivors who does that. So a 'resounding victory due to a well planned out campaign' could also be reported as 'a complete shambles due to shoddy planning and poor communication that we were damn lucky to survive, let alone win'. Both are 'true' in the view of the reporter, but would AI understand that and be able to accurately report on it? Well, not if it's taught to speak with authority and only view and not multiple.
More, that it is in our nature to 'forget' bits of information because 'it's common knowledge' so 'everyone knows it'.
Classic Example: How to make a cup of tea. How many people would think to include: What temperature the water needs to reach. What exact measure of tea needs to be used. How long should it be left to steep. Milk first or last.
And then think about instant coffee: I'm looking at the 'preparation' instructions on Kenco instant: Add 1 or 2 teaspoons of coffee to your cup & add hot water just off the boil.
Makes sense to us, right? But no mention of milk or sugar/sweetener, nor how much hot water, or what 'off the boil' means. It also doesn't indicate if it's level teaspoons or heaped. It doesn't mention stirring. It doesn't indicate the size of cup and how that might affect the measure of coffee powder. So much missed, but WE know what is meant because we already know how to make instant coffee.
But to an AI that lacks context? Lacks the 'common knowledge'? That's going to be one interesting 'cup' of coffee... and potentially quite the mess to clean up, too.
"Classic Example: How to make a cup of tea.
And then think about instant coffee"
Both examples that can't even rely on common knowledge but also have to be cross-referenced to the individual drinker's personal preferences and the environment.
Nowhere in the common knowledge of tea or coffee making does it say that one of my colleagues is lactose intolerant, so I have to use pea milk in that tea instead of cow milk, or my wife only ever wants her coffee cup half filled, and kills it with too many sweeteners, or that I might only have decaf in the evening if I hope to sleep, and that while I have milk in a lot of tea, I prefer Earl Grey with lemon.
And am I at work throwing a tea-bag in a mug, or going all posh and using tea leaves in a pot, or having chimarrão with my wife, who is Brazilian, which is a whole different way of making and drinking tea.
I've read that standard. I'd say that NO one aside from tea leaf testers makes it that way. There's no mention of "straining the leaves out of the tea", and the measurement of tea leaves is within "+-2% accuracy".
Much like the current ISO standard for "Absorbency of adult absorbent undergarments" is not measuring the practical absorbency.
I gave a little in-class assignment prior to talking about algorithms. Each team was to write instructions for making a peanut butter and jelly sandwich, assuming a new loaf of bread, PB and jelly, a plate, and a knife were on the table in front of them. OK, everybody done? Great, now hand your instructions to another team.
Then out came the real loaves of bread, jars of PB and jelly, paper plates, knives. Follow the instructions you've been given, but feel free to interpret anything ambiguous as you like so long as you don't disobey the instructions.
Oh my.
Some groups had to stop immediately: they hadn't been told to open the bag holding the bread (I told them to open it -- and the jars -- and keep going). One particularly creative team applied the peanut butter to the crust around the edge of each piece. General hilarity ensued, and everyone got a bit of free lunch (though some were a bit messy to try to eat).
And then there's false info. Like the AI which picked up the joke on Reddit about putting glue on pizza and didn't realise it was a joke, since current AIs can't. Or on the history side reading propaganda (ancient or modern) as fact.
For another thing there's picking up the wrong info and applying it. Maybe it'll give you some info about configuring a different router if it can't find anything for the one you're asking about and claim it as how to configure this one, which might work or might not. I've seen an AI do much the same sort of thing when asked about an obscure game.
A bigger problem is information is updated so quickly these days the information given by AI is often months or even years out of date. It often assumes something is true because some evidence exists stating the thing is true, without considering if newer documentation supersedes that information.
If you talk to AI about cloud configuration, or (to use one of your examples) router configuration, it will give you instructions for menu trees that existed 3 years ago, but have had 20 updates since and now the menus and available options look nothing like the instructions provided.
This is getting better (creating a project and giving it explicit instructions to never trust assumptions and to always go online and get the latest information helps a little), but I've frequently found this to be more problematic than missing information or hallucinations.
If you talk to AI about cloud configuration, or (to use one of your examples) router configuration, it will give you instructions for menu trees that existed 3 years ago, but have had 20 updates since and now the menus and available options look nothing like the instructions provided.
Not that different to human instructions. I'm constantly advising people on how to do xy & z in some software and having to caveat it with "but.. I've not had to do that since 2022 so it may have changed a bit since". And if you just Google for the answer any human-written result is likely to be the same: some random blog purposely published without a date because *assholes exist* but that's probably a number of years old, and if you've had to do this more than a couple of times you're immediately sceptical.
But agreed, at least I have the self-awareness that my instructions might be out-dated. Current LLMs lack this context are far too confident in how they explain everything, fact or lie. The one piece of advice I give anyone insisting on using them is just to try not to be over-influenced by their confident style.
This is what many of us have been reminding people for several years now. There is no understanding of truth vs false, no 'theory of mind' to work out what the questioner might have really meant to ask, little if any ability for the LLM to interrogate it's own capabilities (ie, "does this LLM know what it knows').
LLMs don't even know their own 'names' without a system prompt saying "You are ChatFoo, trained by Bar" somewhere.
.....which really underscores the difficulty of making really good documentation. Technical writing is a skill which is seriously underappreciated, it can make writing software code seem almost trivial.
(BTW -- I'm not a technical writer, I just write some design documentation and some code. But I've worked with enough of them to realize just how difficult it is to describe something succinctly, clearly and with no forward references.)
> Example 1: weather forecasts.
There are AI weather models and MetCheck is using the output of one. It seems to do OK but even they have admitted that it just seems to take the average (or middle ground) of all the other models which sounds a bit like the ensembles they already use!
I use AI to get a viewpoint that I might not have considered. Much of the time it is simply wrong but it does prompt new ideas. So, it has its uses. But much of it is dangerous nonsense.
I can't stand AI written articles on websites. Most of them are wrong, contradictory and misleading. AI summaries can be useful but are also often wrong\hallucinations. The problem is the hard-of-thinking taking them as fact.
The problem is that the 'AI' is good at being plausible BUT does not 'know' what is 'Right' ... this is the problem that hits the headlines everyday.
Some person thinks that the 'AI' is all it has been sold as and squeezes out a report or two ... this looks OK as the user is not a knowledge expert and it is passed on as ;All Good'.
Sometime later the customer or the knowledge expert reads the report and finds it is a contrivance of plausible 'facts' that match the pattern for the report type BUT is not actually valid/useful.
If the output of 'AI' was meant to be 'Gold Metal' the actual output would be 'Fools Gold' ... looks like the real thing but is of no value at all and cost much more to produce than it could ever be worth.
The 'AI' Bubble is getting bigger & bigger ... soon it will collapse due to the huge size it has achieved being unable to support its own weight !!!
There are more useful things to spend the money on !!!
:)
That's more true than it seems. The main point of divination (unless you actually believe in faeries) is to introduce some randomness to speculation. This is useful as inner thoughts, if left alone, tend to run in circles. Some random input can break the thought pattern, so when it settles again you might come up with a new idea, or a new angle on the same idea. Of course, a couple books and a tarot deck, or a home telescope, or even a farm animal and a sharp knife, are all a lot cheaper than a data center...
John Pertwee's (4th) Doctor used to toss a coin to make decisions, and then do what he wanted to do anyway.
Julius Caesar got himself qualified to take the auguries. While other generals had to wait for favourable omens to be able to fight the battle they wanted, Caesar got to decide what the Gods were saying himself - and was then qualified so the priests couldn't argue with him.
Whereas poor Publius Claudius Pulcher hadn't taken that precuation. So when the sacred chickens refused to eat before the naval battle of Drepana - he wan't able to stop the augurs from declaring that he shouldn't fight the next day. His response was to throw the chickens overboard reportedly saying, "if they will not eat - then let them drink!", then fight anyway. The battle was a disaster, the Romans lost most of their fleet and he only just avoided being tried for treason and was done for sacrilage instead.
David Hicklin,
He would have been executed for treason. Was merely banished for sacrilage. There's a story that his wife remained in Rome. Points for solidarity! But massively blotted her copy book when she was held up in a massive crowd while out in the streets of Rome. And was supposedly overheard to say, "if only my husband could lose another battle, this place would be less crowded."
She's here all week. Don't forget to tip your waitress.
Surely Jon Pertwee was the third Doctor, with Tom Baker being the fourth.
Oh bugger! I've annoyed the nerds now, I'm doomed! The worst thing is, I even counted on my fingers. Doubted myself that he was third, maybe there was another after Troughton? Then thought no, because Baker was 4 and Davison was 5. Then left 4th in the post anyway and merrily went upon my way, having wasted all that thinking time and not then changed the bloody text!
Made a note in my diary today. It simply says, "Bugger!"
Divination is a matter of perceiving patterns to predict outcomes.
I can predict that throwing a stone into the canal will not change the flow of water. I can also predict that next month (November) will be colder on average than this month. I can predict the canal will ice up over winter. I can predict that some birds will fly south soon. I can predict a lot of things because there's a repeated pattern that's been fairly consistent for decades.
And I can also get some of those predictions wrong because it might not be cold enough for the canal to ice up. And we might get a surprise heat wave next month. And the birds have already flown south that intended to do so, unless it gets warmer.
Divination is just that with a bit of extra spin to make it sound like you can see the future :p
So that would say AI can divine the future: It can read past data and calculate the pattern to predict what happens next.
Unless you're talking about the other divination, such as finding water, which is sometimes just a con trick, and other times just blind chance.
" I can also predict that next month (November) will be colder on average than this month."
But people use divination to decide whom to marry, when to fight battles, and where to build their offices.
We hear from every study that the outcome of AI is just as useful as Astrology and Feng Shuifot such decisions..
Ooh, another example of writing around that time was this writing about tulpas and egregores, basically mental projections that practitioners say can take their own appearance and volition:
I don’t believe in AGI (yet) and I don’t think ChatGPT or other LLMs are sentient. But I am struck by the similarity here to reports of weird chat LLM behavior, which go way back now—and continue to appear, along with incantations like repeating the letter “a” one hundred times and watching them spew craziness. Weird behavior seems particularly common when people try to jail break them.
I honestly do feel like LLMs tap that kind of tendency for us to assign meaning and volition towards what is essentially a stochastic process, and unconscious biases can drive some pretty out there and apparently uncontrollable behavior.
Often as not, asking an LLM for a more detailed description / explanation of an issue results in a word salad.
I've played around with one or two over the past year and Copilot for example seems to be getting worse, presumably due to ingesting the BS generated by itself and other LLM's.
That's the problem the AI hype train has. Getting people to pay $200+ a month requires some service they can't get at home or the office so they push for bigger and bigger >100 billion parameter models, but most people only need the capabilities that <10 billion parameter models can provide, which can usually run locally.
There's already two (announced) 1-trillion parameter models. Qwen3-Max and Kimi K2. Since ChatGPT doesn't seem to publish parameter numbers, the current belief is that their largest model is more than 1 trillion parameter, and possibly more than 10 trillion.
The 'parameter arms race' is real.
I have been working on some IP networking recently. Chatgpt was useful - it replaced much reading of command line "man" pages. It suggested ways ahead and how to diagnose when I was stuck. However, it eventually got stuck in a loop. ie.
chatgpt:Remove X replace with Y
me: Ok you fixed problem 1 but now problem 2 shows up.
chatgpt: Remove Y replace with X.
me: hmm...
But that gave me the hint of where to go looking for the problem.
For me, chatgpt was like an assistant that had read all the books and knows a lot, but without the logic to spot inconsistencies.
As long as it is treated as a helper, and not a solver, and you don't swallow all the marketeers hype, it is useful.
The problem is that unless you then go and read the man pages yourself you have no way of knowing if it just made some of it up. Even if its suggestions work, they may not work the way they're supposed to, which can lead to misconfigurations that are harder to diagnose.
LLMs can be a substitute for menial tasks under close supervision but they should never be a substitute for your own knowledge, because without that knowledge you've no way of knowing when or how they're screwing up.
Basically, if you use LLM ‘advice’ to fix surface-level problems, you set up yourself for being absolutely lost later – when tasked with solving the subtle and non-local problems the previous fix created. Or you may be lucky and doing something simple for which the LLM-regurgitated stuff is sufficient. It can happen. Usually, could just do a web search and copy'n'paste from some guide/tutorial in such cases, but whatever.
See this type of problem a lot when it comes to doing stuff with “Windows”.
The problem is the LLMs/AI assistants can’t tell the difference between an XP tutorial and one for W10. Hence given the number of “How to do xyz on Windows nn” tutorials the AI will present an amalgamation … which will always be wrong. The safest and quickest approach is to ignore the AI results and dive straight into the articles, giving greater credibility to those from recognised tech sites.
I see that too. "Try looking in the settings for <whatever> and change it there". It's a good place to look, I've looked there already myself, and if the software had implemented it that's where it would logically be. It seems its memory of settings pages is as bad as mine. If you ask an ai with an agent like comet to then show you the page, it gets all confused and accuses me of having heard a rumour.
Yes LLMs can be useful to experts and non-experts alike, just like a Search Engine occasionally turns up an interesting link discussing the problem at hand from a new angle or with new information.
However, that is not the idea that makes LLMs seem worth billions of dollars...
The billion-dollar use case, replacing white collar labor with AI, is currently about as viable as replacing those seats with Google Search licenses.
"Most Scientists Use AI But Question Its Usefulness".
Yeah right. Like "Nobody goes there anymore because it's too crowded".
These scientists are just exhibiting symptoms of Upton Sinclair Syndrome: "It's difficult to get a man to understand something when his salary depends on his not understanding it".
> "Most Scientists Use AI But Question Its Usefulness".
Much depends on the question.
In recent surveys, I know Google et al use “AI” but I don’t really know how. I also use Grammarly - another tool which uses “AI” in some form. So depending on how the “use AI” question is asked, the honest answer is yes, with the usefulness question getting a no or maybe response; but those probably isn’t the intended answers, as what they were really wanting to know is whether you interact with a LLM chatbot and cut-and-paste its output.
There are problems with the application of AI in medicine:
"ChatGPT Salt Advice Triggers Psychosis, Bromide Poisoning in 60-Year-Old"
https://www.medscape.com/viewarticle/chatgpt-salt-advice-triggers-psychosis-bromide-poisoning-60-2025a1000qab
(I've stripped the tracker out of the URL)
Sodium is an obviously vital part of the sodium-potassium pump, which drives action potentials down the axon, to the synapse, an electrochemical event. There is such a thing as too little, and too much.
I'm using AI to help me compile a forensic report, from text with all data anonymised. I'll reverse the process when it comes to rewriting it again.
My experience with GitHub CoPilot is pretty much the same as suggested by this article. Once you've adjusted you IDE so that CoPilot doesn't splash a half-dozen lines of code in your editor everytime you type a enw character you can get down to work. At that point you, you are just asking questions of a chatbot. It if very familiar to most programers who have grown accustomed to "Just Google"ing it over the last decade or so. The results are very much the same. You need to actually read source of the results and gather context for the answers. The summaries provided by the chatbot makes this much harder to do. The GitHub chatbox seems to just summarize all the code found in the GitHub repositories, regardless of code quality or purpose. Quite often things are presented like they are "best practices" when they are simply a survey of the latest things a bunch of junior developers have pushed up to their private repos.
In the end, you just need to understand stuff. AI, at best, is a "spicy search" that can help you find connections to things. The human still needs to understand those "things". Chatbots need to make it easier to find the sources of the infomation it is presenting so that the chatbot user can look at what the human did or said.
In real life when someone's faced with a task who only has access to incomplete or contradictory information they'll come back to you and say "I don't understand this" or "That doesn't make any sense to me". At least they should do -- we've all worked with colleagues who claim to be always right and plow on regardless, treating every question about what they're doing and why as a direct assault on their very being. They might produce a work of genius but more often than not what they usually make is a mess.
Current versions of AI seem to be 'that' colleague on steroids --- it will always give you an answer. Whether or not its correct or logical is "TBD". So using it in situations where its not tightly constrained or well supervised is asking for trouble.