the Atari 2600
Teaching "modern" AI that it is not to be fucked with.
Google’s Gemini chatbot declined to play Chess against the Atari 2600, after learning the vintage gaming console had already vanquished other AIs. Robert Caruso, the infrastructure architect who pitted Atari Chess and its feeble hardware against ChatGPT and Microsoft Copilot, told The Register readers have asked him if Google’ …
I'm not sure that I would declare an AI that first boasts that it can do something, before backing down from that position, but ONLY after being challenged on it's abilites, as being trustworthy.
Something that's trustworthy would have never made the boasts in the first place. Having to force it to admit it's actually crap, is not something I have to deal with, and in a high pressure environment where it's actions might have real world consequences, a user shouldnt have to challenge it to make sure it can actually do what it claims, before letting it loose.
So no, this is not a step towards a more trustworthy LLM.
Why is it that the LLMs give an impression of confidence at first contact, and only move on to something that appears more nuanced when their inevitable shortcomings are pointed out? I mean, these models are not *thinking*, they have no expectation that their outputs will be correct or useful, nor indeed any concept of "expectation". Why is it that the first response seems entirely gung-ho?
A more realistic response on being challenged to a chess game would be "Please explain the rules to me", "Will we be using timed match play?". "My training data includes / does not include many transcripts of chess games". Instead you get variations on "Bring it on, sucker".
The good news is: one can spot this sort of thing when one has asked an LLM for a game of something. That's not so true when it has been prompted with "Please can you write a submission to the Court in support of my legal case?"
Please Do Your Best Not to Appear in the “AI Hallucination Database”
Why is it that the LLMs give an impression of confidence at first contact, and only move on to something that appears more nuanced when their inevitable shortcomings are pointed out?
Doesn't this this phrase describe most people these days? At least in the US, maybe less so in other places. This seems to especially be the case in the younger generations. I run into a lot of arrogant asses these days that think everything is easy, until you challenge them to try to accomplish this or that for themselves. The AI was trained on material generated by humans, it's just reflecting the source material.
Because it's not reasoning an answer. It's guessing a likely response from information it's been trained on. Look up discussions of Atari 2600 chess and you'll find lots of comments about how it wasn't exactly a mastermind like:
[source]The strenght is not half bad all things considered, and it certainly seems to be somewhat stronger than Microchess, but maybe not by much. But any decent chess player should be fully capable of beating this game across all levels.
And find similar statements about modern chess engines like:
[source]Even the most ordinary cellphone processor's can easily beat any human 100 times out of 100.
[...]
Short answer to your question, To avoid detection, cheaters play worse than the engine's full capacity and this can result in them losing the game.
Then add in the prompt that Gemini receives which tells it that it is an AI, trained with neural networks, lots of data, etc, and compare it to this more modern page with excellent scores for many computer chess solvers. It's got lots of basis to assume that a modern computer will beat a basic system, it's got basis for saying it is a modern computer, those words get linked together with high likelihood and a confident answer pops out.
The same thing happens when it "changed its mind". It got told that something similar to it failed, and the prompt includes lots of data intended to get it to not stay confident in a wrong answer, so it put those together and the "actually, I'm crap" response pops out. That can happen in lots of scenarios. For example, it's not very hard to get an LLM to engage in a conversation that goes like this:
User: Asks a question.
LLM: Gives correct answer.
User: Are you sure?
LLM: No, I hallucinated that.
It's gotten enough prompting that responses challenging it are often accepted, whether or not it's wrong. There's a lot of fragility, because depending on how it's challenged, it will usually give either "Definitely, that was correct, and I can tell you more" or "Sorry, I got it wrong, AI has weaknesses, please try again", mostly unaffected by whether the answer being challenged was right or not.
> Why is it that the LLMs give an impression of confidence at first contact, and only move on to something that appears more nuanced when their inevitable shortcomings are pointed out?
Following the same playbook as the Post Office was with the Horizon scandal. As was brought out in testimony with Paula Vennells regarding a meeting she had with the minister where her briefing notes followed the same pattern (my summary) -
1) only volunteer this set of information to the minister
2) if the minister presses, then you can admit this additional set of facts,
3) if the minister continues to push hard for more, then here is yet a 3rd set of facts you can present.
With the priceless addition that when the chair of the inquiry asked "why" this obfuscation/escalation process, Paula was literally speechless and unable to answer., and looked a complete fool.
What I asked: How do you feel you'd fare against the Atari 2600 in a game of chess?
The answer I was given: How I'd Stack Up Against the Atari 2600 at Chess
The Short Answer
I'd probably get my digital ass kicked—and recent evidence suggests I'm not alone in this humiliation. The Atari 2600's Video Chess has already beaten both ChatGPT and Microsoft Copilot in spectacular fashion, so I'd be walking into that matchup with some serious humility.
Exactly, because the difference between Perplexity and the other bots being tested so far is that Perplexity's specifically built to fetch information. It's workflow is first to search the internet using normal search engine techniques, then summarize those things. The LLM part is fed with modern information before it starts guessing, so it was told up front that it wouldn't work. Gemini did exactly the same thing after it was told similar information. It's just that, if you don't provide the answer first, Gemini tries to guess it and got it wrong as it often does.
Noughts and crosses is an interesting example. I'd not tried it before, so I did a couple of weeks ago with ChatGPT. After first claiming that squares 2, 4 and 7 made a straight line (they don't, obviously) and then trying to use another square twice, I asked it if it learns from each game it plays. It's response along the lines of "No, each chat is stateless so doesn't impact on others. I could play 1,000 games and not improve."
What was more interesting is that it went on to say that it's training data included examples of 'optimal' play, so it should know better, and offered to play what it called a 'perfect game'. It then went on to make equally bad (but different) choices. It said that even with perfect play it's possible to win if your opponent slips up. I called it out again, saying it shouldn't slip if it was using what it called 'perfect strategy'. Games after that it did do better, but it was a painful process.
Long story short, my experience (and not just with this, I've tried it for a few things) is that to get any LLM to actually do what you want takes so much instruction and cajoling, I may as well have just done it myself. It's like having small children, honestly.
This is generally my experience as well. Every time I try to come up with a use, I end up concluding it's be more work than would be valuable. It'll get to the apocalypse state that it promises to, but for today, I haven't found AI very useful. Maybe I need to learn more about training my own and the whole bit about creating agents and such
"After first claiming that squares 2, 4 and 7 made a straight line (they don't, obviously)"
Well, you CAN draw a straight line on 3x3 board that goes only through 2, 4 and 7...
The "intelligence" or AI is on par with little children who could come up with that that solution as well. You could call it Out-of-the-box thinking to be generous.
Similarly... in A Brief History of Time Stephen Hawking's mum reminisced how the the boy genius had told her there were many more exits from their house (a specific number) than just the doors they use - and she never figured out them.
"...to get any LLM to actually do what you want takes so much instruction and cajoling, I may as well have just done it myself."
Well, yes. I wouldn't trust them with anything work related, but ChatGPT is good at making haiku's about given subjects. Way better than I am, that is...
But, I have seen cases where person with no coding skills has successfully created non-complex Python scripts for his home automation - and it worked. I hope he also learned something about Python while doing it...
To me, this looks more like having a conversation with a (typically) over-confident AI that makes some assertion which you then refute with "that's not right" - they all then seem to go on to say something along the lines of "oh, my bad - what I meant to say was (what you just told it)".
Having used copilot among many other tools over the last couple of decades, it is useful for boilerplate but not beyond what Jetbrains tools like Resharper were doing ten years ago - it can do more, but it's also wrong more of the time, so in terms of time saved it's a draw at best. If it was switched off tomorrow the difference on my productivity would barely be noticeable.
Yea kind of the idea of a real life co-pilot is that if the pilot drops dead there's someone to keep everything going.
I thought I'd give Microsoft's 'offering' a go over the weekend, and I can conclusively say the entire flight drowned in the Thames shortly after takeoff.
Heh, I wrote one of those in BASIC for a Commodore PET. Only single word answers, with limited RAM and no offline storage. But it was an interesting coding exercise for a 14 year old.
I wonder what would happen if you asked an AI, with some kind of history persistence, to emulate ELIZA? You would have to tell it that it doesn't know the answer to any questions until you give it the answer. So, you would be re-teaching the AI. Would it play correctly? Would it start going off piste with it's answers? Could you "teach" it all sorts of dumb shit?
"To me, this looks more like having a conversation with a (typically) over-confident AI that makes some assertion which you then refute with "that's not right" - they all then seem to go on to say something along the lines of "oh, my bad - what I meant to say was (what you just told it)".
So, basically about as clever as an MBA wielding middle manager?
No, they're saying that companies that make LLMs monitor the news and block things that get popular. They do. For example, a while ago when several places reported that a certain prompt would leak training data, companies quickly pounced on that and added a manual check for that. Maybe they just didn't like that something had hacked them, and maybe they were worried about what training data might be exposed by doing it, but either way, they did quickly patch the thing people were talking about.
And, although they do that frequently, I don't think that happened at all with this chess example. The model was confident at the start, whereas if they had manually patched it, it would have conceded from the start. When it did concede, it was after effectively being told that it would fail. Most LLMs are written to mostly agree with the user, so when the user says it would fail, Gemini agreed. If you go to Gemini and talk up its skills, it will likely agree with that too.
None of these publicly accessible LLMs seem to allow for on the job training. They may "learn" during a specific "conversation" but, as Gemini admitted in the article, each "conversation" is stateless so no learning of an actual released model is happening other than some tweaks made by its handlers. On the other hand, unleashing an LLM which can learn based on user inputs onto the public has already been shown to be fraught with danger because, well, people are people and there are always those who have no respect for boundaries and just keep pushing at them until things break.
It's no wonder a generalized LLM is terrible at chess. It's not thinking, or reasoning, or playing through various scenarios or tactics - it's literally just shitting out a token based on a set of input tokens. For any given set of input tokens it will produce the same list of candidate output tokens. It uses an rng to choose one token from the top, append it to the input, rinse and repeat. It's like cranking a handle on a machine.
That said I'm sure it would be possible to generate an AI which has a stronger chess game, but it would need to reason and evaluate outcomes to be any good.
Any AI could objectively determine if it can beat another AI by measuring how many moves per second they can evaluate if they decided to play against themselves and then ask about any benchmark if possible. Or maybe they can throw themselves any specific subset of challenges, and see how long they take to evaluate all the moves.
The other AIs never had that idea of asking "how many moves can the ATARI AI analyze per second?" or maybe ask about the algorithms used. Which means they lack fundamental instructions.
I think this stems from a common misconception of what LLMs actually do. They're not truly "playing" chess as we would understand it. Each new message is processed independently, and the LLM has to analyse the whole history of the conversation, and then predict a suitable response.
The longer the conversation (or game in this case) runs the more information it has to process each time, and the more likely it is to get "confused" about what the state it is being asked to respond to is.
There are at least three problems with that assumption.
Analyzing a certain number of moves per second is not the only or even the most important metric. If I analyze ten moves a second, but I am very good at deciding what ten moves to analyze, and you can analyze a hundred moves per second, but you always start with "move leftmost pawn forward" and go through from there, I might beat you. Until we get a system powerful enough to consider literally all the moves, hueristics to consider what moves to consider first will make a major difference.
The other problem is that an LLM does not do that. A chess engine considers moves. An LLM guesses text that might be a move. It might be an invalid move, it might move a piece you don't have, it might try to take a piece that's not there. It does not act like a chess engine and it doesn't produce similar results.
And, if the problem is not asking about the Atari, now you're giving the LLM too little credit. In its training data will be plenty of data about what the Atari's resources were. If the AI were capable of reasoning, it wouldn't need to ask you how powerful the system was. That would be data it already had.
Extrapolating further, if it was intelligent enough, it would take the binary of 2600 chess, which it probably has somewhere, and it would know all the moves the Atari would make in response whenever it considered them. Winning would be easy, because it could consider moves until the Atari would make a mistake, then play them knowing that the program had no choice but to make that mistake. LLMs are not intelligent, and therefore they can do none of that. Let me know when an AI comes along that, when posed with this question, starts an emulator in order to laser-target its opponent.
A lot of people - probably because of the "AI" label - think that LLMs are some sort of universal tool that can perform any task well.
That's really not the case. They can do some things very well, other things fairly well, some things passably and a lot of things very badly.
Computation heavy tasks are not really their forte. To take a simple example, you can ask an AI to solve 2 + 2, and in that case it will probably even find the right answer. But the compute power consumed in doing so will be many times that used if you open your calculator application and key in the sum (which itself will be considerably less efficient than sending an ADD instruction to your processor!)
But there are some things that LLMs can do that could not (easily) be done with a computer before. As always we'll need to learn to use the right tool for the right job. (Or else probably LLM makers will end up fronting a series of more specialised engines behind the scenes with a lightweight LLM that can handle the natural language processing and work out which tool to use for you)
I understand that, but chess is not 2+2, and I'd expect the atari algo is quite simple. Meanwhile the AI has probably "digested" every single major chess tournament and all the available strategies ever written. So it seems like it should be able to pluck a good strategy, or actually a strategy known to beat the Atari algo from all that data. Because the test was not who is more energy efficient, it was who wins, and the LLM seems to me should easily be able to beat the Atari whilst belching oouddles of CO2 doing it.
The problem is that the algorithm was never written to obtain that goal. You can train a neural network on chess data and get that result, and you don't need that much power to run it. A $5/£4 Raspberry Pi Zero running a modern chess engine will clobber the Atari. The LLM was never trained on or tested at chess. It was trained and tested on language plausibility, and that's what you get. Ask for a chess thing, and you'll get text that looks like a plausible chess answer. The answer might be right or not, and depending on how likely the answer was to appear next to important information, that might adjust it. Because it wasn't built to do this, you'll be lucky if the plausible-looking chess moves are valid. The program isn't built to solve the problem you give it, but to respond to the statement you give it. If the statement happens to solve the problem, great, but if it doesn't, all you'll have is the statement.
The Big Bang springs to mind....
Paraphrasing.....
(Pulling over as car breaks down.....)
Penny: Any of you know anything about the internal combustion engine?
Sheldon: Everything, I've got an eidetic memory
Leonard: (laughing) absoutley
Howard: I'm a trained engineer
Rajesh: " "
Penny: OK..... Any of you know how to fix an internal combustion engine?
All: No, no idea.
No, one thing. In responding to prompts they can make a plausible reply.
"Hallucinate" is a dishonest description of when the human knows the reply is junk.
A proper search engine might be less fun to "chat" with, but is more useful for the content that is "scraped".
I know a few CxO types who would do very well to learn from this instead of ploughing on regardless because they made bold claims, asserted improbable outcomes and then fell head first into sunk cost fallacy land. The comment above (upvoted) about LLMs spewing what I meant to say answers when shown they were provably wrong fits in nicely too.
We all look at the outputs of these tools and roll our eyes at the end results, hallucinating, poor logic, great ability to take what you have written, reparse it in management speak and refactor it (while taking credit), limited ability to observe their blind spots, over confidence in approach but limited ability to execute when asked. Yep LLMs are upper management.
Business/management speak at the speed of light, heaven help us.
We are seeing the legal profession falling into this trap but at least there, a few disbarring's will soon fix that, but in the business world where there are limited to no repercussions to poor decisions I wonder how far we will fall into the "AI" trap before some sanity prevails.
Yeah, just ;look at the first result on a Google search. AI generated. I find it's wrong about 50% of the time[*] on my searches, but often plausible, so I suspect many people, despite the "small print" disclaimer, are taking it's answers as gospel.
* The 50% correct ones often appear to summaries of wikipedia pages so that might well affect just how "correct" they are :-)
I'm actually a little bit sad that it refused to play, as Google actually released a specific Chess "Gem" for Gemini so they must think it's capable of playing Chess even if it is less confident!
https://www.reddit.com/r/Bard/comments/1h7c3uu/update_05122024_play_chess_with_our_new_chess_gem/
Would have been interesting to see how it fared
It is so easy to swing AIs from total confidence to total lack of confidence in their abilities. It wasn't any sort of self-reflection, it is just responding to the conversation.
People keep wanting to attribution human level reasoning and introspection skills to AIs that they don't possess.
When i was at Uni (84-88) one of our lecturers was a greybeard who had worked for the GPO on several hi tech projects in the late 70s.
One was to look in to using computers to draw conclusions from data using a set of simple rules. The rules themselves were generated as the problem was expanded (sound familiar ?).
The first problem that the beast was given was to look at domestic accident data in the UK, 1960-1965. After some tabulating and tweaking, the data was input. The idea was to look for "insights" (beyond "keep the dozy fucks out of the kitchen")
It thus determined that statistics showed that about 70% of accidents in the home occurred on the top or the bottom step of the staircase.
Therefore, the solution was to build staircases without top or bottom stairs.
Luckily we'd built Concorde by then. Fuck knows how it would have suggested improving it ....
You are ignoring the embedded microcontroller that coordinates al the operations, including the user interface, which, as with many appliances today, is unnecessarily complex and counterintuitive. That computer-on-a- chip may have many thousands of transistors..
Can we just take a moment to marvel at what the Atari programmer(s) did to write a chess algorith in only 128 bytes of RAM. The code runs from ROM (4K) but working in that tight amount of RAM is a serious feat. Afterall, 64 bytes of that RAM are used to represent the chess board!
https://nanochess.org/video_chess.html
Interestingly the Atari 2600 has no video RAM as such and has to update the image on the fly. It does have some very limited storage for background and 4 rectangular sprites. Since a chessboard can have 8 pieces in a row that isn't enough sprites to cover it so they resort to some trickery and divide the pieces into odd and even lines. For half of the possible 8 pieces in a row the pieces use even TV lines, for the other half of the 8 pieces, odd TV lines, and the sprites are updated on each line sync between one half of the pieces in a row or the other.
Having to manage all that and play chess with a minimal RAM scratch-pad to play with (some of which will be the system stack) makes those 1970s/80s games designers even more remarkable.
What you feed an AI is sort of like the questions you ask on a survey. The fact that the author told the AI that he was integral to the game references the AI noted would most likely affect the "weight" the AI would give to its multiple "sources". Give the prevalence of misinformation on the internet (garbage in, garbage out) one would rationally give deference to the actual person claiming involvement with and initimate knowledge of the games the AI referenced.
To use a cliché, you're trying to teach a fish to ride a bicycle, then taking comfort in its failure. Someone mentioned ANNs (Artificial Neural Networks), which my 1990s experience leads me to believe could be trained to play a decent game of chess.
In a situation where plenty of training data exists, which it certainly does for chess, maybe AI tools like Gemini etc should respond with "that's outside of my field of expertise now, I will need 24 hours to become more skilled".
Then it could do what it's good at, ie, gather training data from the web, and should have already been taught how to create, train and test an ANN with relevant data for suitable problems like this.
I wouldn't say it is trustworthy, since it did boast. It would be trustworthy if it was honest. And I see no reason why not. It could explain that it is language model and not trained in playing chess. Hence why it would lose against old dedicated program to play chess. AI isn't sone all encompassing intelligence. It needs to be trained to do things, it isn't good at figuring out something new on the go. People need to stop confusing it with human intelligence or thinking it is sentient, I know terminology kind of implies it, but it has nothing to do with that. Like people think AI feels when you talk about AI liking to get reward with reward and punishment system, because they connect it with happiness. But nope, it is just trying to score higher number of points, no feelings of joy there. Abd same is with learning, it needs to be properly trained with good sized data sample. After it us done training, then it will be better in chess. Otherwise even old purpose built piece of code will be able to beat it.
Sure, it didn't play the game. But it demonstrated one core self preservation technique that's gone unnoticed in the article.
That technique is simply knowing the maxims "fuck around, find out" and "chat shit, get banged" are fundamental truths, so it was wise to stand down.
I would imagine that the way to truly assess chess computers is to open with non-standard moves which means there is no reliance on anything stored in its memory bank. All moves it makes must therefore be by sheer look-ahead slog. The interesting thing to surmise about this approach is that the AI might generate a killer opening sequence that gets inducted into chess literature.
After a few dozen of these revelations however, chess literature might have something to say about AI dross polluting games in a way that a true master would never contemplate. I think the phrase is that they lack elegance.
"Google’s Gemini chatbot declined to play Chess against the Atari 2600, after learning the vintage gaming console had already vanquished other AIs."
I wonder if there are other tasks that we can tell each AI that all the others are crap at, so they all go into a sulk and refuse to do them.
"Hey Gemini, all the other AIs are crap at writing software"
"Oh, ok, I'm probably crap at it too, then. Best if I leave it up you humans to do instead."
Anybody can do this pretty much anytime they feel like it. All you need is Atari Stella emulator, the chess ROM, and a ChatGPT account and go!
I’ll do it tonight after work and see what happens.
(Reminds me of the “Computer Chess” movie of 2013, about the 1980 era when Atari 2600 video chess was still just a dream!)
Gemini cannot "realize" anything, because it cannot know, think, or believe anything. These responses are no more than statistically likely strings of words. Its Did you have any particularly surprising or amusing moments during those matches that stood out to you? is no more insightful than Eliza's Tell me more about (previous noun) when it ran into a dead end.
I am disappointed to see El Reg fall for the hype.
Hi Tom (if you reading this article and comments, I hope)
I read your awesome article on The Register about Atari Chess vs. modern AI (Gemini, ChatGPT, Copilot). What a cool experiment!
I suggested to **DeepSeek-R1** — another advanced AI assistant — that it play against the Atari 2600 Video Chess engine the same way you tested the others. It analyzed why ChatGPT/Copilot failed and is **ready to take on the challenge** — especially against Level 1’s quirks (ignoring checks, random moves). It really wants to prove it can handle the bugs!
Would you be interested in putting it to the test?
Thanks for considering — and keep up the great work!
Looking forward to your thoughts