The Register Home Page

* Posts by that one in the corner

5504 publicly visible posts • joined 9 Nov 2021

Singapore, Amazon lead push for 'purpose bound' digital money

that one in the corner Silver badge

Re: Translation?

I read it as closer to putting a hold on the money whilst it is still at your bank (so no extra trusted third party) and transferred in payment when the conditions are met (delivery of goods); one also hopes that the conditions are also set to allow cancellation of the order and/or a maximum time for completion, after which the hold is released an you get control over the funds back.

This way, your bank gets all the transaction fees (including when the supplier calls to verify the hold is in place) nothing for that greedy escrow agent.

UNFORTUNATELY although the above fits the description given in the article and could actually be a useful addition to banking services, the actual whitepaper starts by prattling on about digital currencies and stablecoin, so it isn't going to be anything sensible and will probably end up with overly-complicated smart-contracts that are coded in Javascript and can be gamed in all sorts of fun ways. Hmm, if I can cause *this* "contract" to trigger now, which will trigger Fred's contract, but mine is so badly code (oops) that it garbage collects, meanwhile Fred's has completed and my other contract has swept in, got the funds, raised some interest, released it back to my first contract which is now ready to continue...

Linux 6.4 debuts after literally unremarkable development push

that one in the corner Silver badge

Re: Are we sure it's time to throw something else out?

Also used work on Windows Embedded boxes that boot off CF cards (that kit is still in service and should have years left in its life - no, it isn't connected to the Internet).

At home, still have a perfectly functional CF bootable mini-ITX mainboard system, ready to be put back into service.

The problem nowadays is that IDE is only one of the modes that CF supports and the later cards stopped supplying it. So have gone from "any of the CF cards work to boot from" in the days of 512MB cards to "you suddenly pay a massive surcharge for a card that actually follows the whole spec" which made, say, a 4GB CF too costly to be worth it.

And three guesses whether an up to date copy of the OS that used to fit in 512MB will now even fit onto the 4GB card if I'd bought it! Annoying, as the 512MB card is in perfect nick (it only got read at boot, never written to during normal use). Ah well, at least SSDs are more affordable now than bootable CF cards.

Google asks websites to kindly not break its shiny new targeted-advertising API

that one in the corner Silver badge

asking advertisers to promise they won't abuse...

Well, I feel safer already.

We can all trust advertisers to always do the right thing, the same way I can always trust the bacon to glide through the window and land squarely on the bun as soon as it hears the ketchup being opened.

If AI drives humans to extinction, it'll be our fault

that one in the corner Silver badge

Re: A Future Surprise for Current Running Realities, the Rise of Virtual AIMachines .....

Any exploration of what hasn't happened yet is fiction.

(So is a lot of other stuff, btw)

One person's trash is another's 'trashware' – the art of refurbing old computers

that one in the corner Silver badge

End of 2025 - that sounds awfully close, and we still don't have a date published to fibre-enable our exchange.

Microsoft investigating bug in Windows 11 File Explorer that makes the CPU hangry

that one in the corner Silver badge

"always show icons, never thumbnails"

Huh? Why do you want to remove that one?

> "display file icon on thumbnails,"

Ah, now it makes sense.

Remove those two and you'll never see icons for anything that can be thumbnailed so you won't spot the moment that Microsoft overrides your choice of image viewer. Again.

"IrfanView? Bleugh! You want a proper Microsoft Media Experience!"

that one in the corner Silver badge

Re: "hide protected OS files"

> if you're looking for them you're probably capable of editing the registry to find them.

More sensibly than editing the registry (seriously?) try:

Probably capable of finding a better file manager program.

that one in the corner Silver badge

What, no Task Manager?

> Right now, the only workarounds for stopping the high CPU utilization problem is either restarting the device or for the affected user to sign out. Locking Windows won't do the trick, Redmond wrote. The users must sign out.

So using Task Manager to kill File Manager has stopped working under Windows 11 as well?

And why would anyone expect locking Windows would help? Doesn't anybody else lock Windows when they go for a cuppa, fully expecting any tasks to continue?

(Disclaimer: zero practical experience using Win11; for all I know, TM has been neutered and locking now means deep sleep!)

Missing Titan sub likely destroyed in implosion, no survivors

that one in the corner Silver badge

Re: A fitting epitaph

> the alternative would be designing a milspec controller that ends up costing >$1m a unit and does the same thing

The story is that they *did* have a "milspec" controller supplied by the contractor, at an obscene cost, until a savvy submariner pointed out to the boss that it was actually just a video game controller (probably rebranded, that is what cost so much). Won't swear that they actually now use an XBox device.

The controller is only used for pointing the modern replacement for the periscope (and no doubt is not the *only* way of doing so - they may even still have those funky folding handles and can wear their caps backwards as a backup, before resorting to "ratings and binoculars").

(Tried to find my source to post the URL but search results are currently flooded with thr Titan story).

that one in the corner Silver badge

> He is able to beam himself at will

EM activates the teleportation by getting into the driver's seat, closing his eyes and meditating for five minutes: when he opens his eyes, ta-da they have arrived!

Meanwhile:

As soon as EM drifts off to sleep ("meditating") the *real* driver takes over (amazing what you can do with a Tesla paired to a bluetooth steering wheel from an old Wii game). When they arrive, all Tesla and SpaceX clocks are wound back to five minutes after they left (if it was only a two minute journey, they get wound forwards) and EM is gently prodded awake again.

This is how EM manages to do his amazingly long work days *and* explains why he is so happy to make those claims about "FSD is only 3 months away" or "you can buy a Semi with delivery in a year": he is getting lots of sleep and, according to his watch, it is currently an unseasonably warm February in 2017.

Don't worry about the employees: the worst they have to worry about is ignoring the factory wall clocks and being sure to have the company app running during meetings with the boss, showing EM time. Okay, there are a few odd days when they end up eating three breakfasts in one day full of meetings and they are super-careful around launch days (when they balance things out by having lots of two or three minute drives around the launch site for each fifteen minute drive from hotel to site).

Red Hat strikes a crushing blow against RHEL downstreams

that one in the corner Silver badge

What about any non-GPL components?

Are there parts of RHEL that are *not* covered by the GPL and without which a fully RHEL compatible version is not possible?

They need not be major parts, in the grand scheme of things, but if there any little bits which may really be restricted by RH's actions that would be enough to prevent any rebuilds from being 100% bug-compatible with RHEL. A point which RH would be sure to emphasise.

If such a situation exists, *we* may all know that those components are irrelevant to 99.999% of Users, but even a few differences would be enough for RH sales & marketing to leverage into full-blown FUD ("you are unique, you are one of the 0.001%" would probably work on most CEOs).

All they have to do is set the perceptions about remixes they don't like, not the reality.

that one in the corner Silver badge

> The subscription agreement with Red Hat is not under the term of the GPL

Which means that RH customers can not pass on a copy of the subscription agreement (but who would want a copy of it?).

Meanwhile, the rest of the GPLed code can be shared, as described above.

Lawyers who cited fake cases hallucinated by ChatGPT must pay

that one in the corner Silver badge

Re: What?

> A catch-all word such as "error" is not better

It would also be totally wrong - the LLM is not acting erroneously, it is working perfectly, doing precisely what the algorithm says it ought to be doing.

that one in the corner Silver badge

Re: What?

I'm happy with using the word "hallucination".

Especially in the sentiment that: "It’s not that they sometimes 'hallucinate'. They hallucinate 100% of the time, it’s just that the training results in the hallucinations having a high chance of being accidentally correct."

If you don't like "hallucinate" about the only other word I know that comes close is "stoned". Groovy.

that one in the corner Silver badge

> Then we'll probably get an AI winter

Oh joy, *another* AI Winter?

WTF can't we just get on with working on these problems without some - people - suddenly deciding that they can hype[1] the hell out of the thing and then embarrass everyone when they fall on their arse!

[1] Yes, I know that science funding proposals end up having some hype in them, but I'm referring to ChatGPT levels of hype: the general public doesn't hear about funding proposals on the TV news.

that one in the corner Silver badge

> LLMs are a new thing.

True in that their actual existence is new - but that youthfulness is down to economics alone (i.e. paying for the cycles required).

But how to make them has been known and thought about for decades.

> nor our education, has prepared us for

Ah, THAT is the nub of the matter. These things haven't been widely taught - and now that they "have entered the public perception"[1] 95%[2] of the available materials are utterly useless to Joe Bloggs The Lawyer; if only he that.

[1] i.e. as usual, everyone who was talking about them was ignored: science outreach is a Very Good Thing but falters because it is, of necessity(?), voluntary on the recipient's end

[2] wild optimism?

[3] I can point you at a lot of good academic or technical material ("whooosh" says Joe) or lots and lots of total bollocks and frankly dangerous "do this, it will redefine your business" twaddle on YouTube.

that one in the corner Silver badge

Re: It's not a GOLUM, either.

I did apologise first. Meanie.

that one in the corner Silver badge

I am VERY glad it was lawyers getting caught out

When people do bloody stupid things like this in other fields, the next few years are spent shouting at and suing each other in court and the waters get muddied, things get settled out of court and/or restrictions are placed on the details of what actually went wrong.

Here the situation has gone straight to a judge, all of the details are in the record for us to see and the judgement is startlingly comprehensible; he almost came right out and said it as simply as "Use the proper tool for the job and that choice is your personal responsibility".

Now, hopefully, the company lawyers will be looking at what their company proposes and be saying "You are using an LLM here? Nope, not going to sign this off, not risking myself in court."

that one in the corner Silver badge

Re: $5,000 Fine ... a Pittance for an Attorney

The fines are not going to financially scupper the law firm, but their imposition has made it absolutely clear that the behaviour won't be ignored, by this or other courts (in the US): they were let off light for the first offence, but everyone has been warned.

that one in the corner Silver badge

> It's bad programming, no more, and no less.

The programming is fine (the code works and performs the maths required of it)

> model being used is fundamentally flawed

The model is fine: it is built upon the maths started back in the 1950s and does what couldn't be done back then: use large numbers of cycles to train on large amounts of input. It builds a correlator.

What *is* wrong is *applying* any of the above for a any purpose other than amusement.

Treat just like the Radium Snakeoil Sales - radium itself is not to blame, it has its place and can even be useful to humans in its place. BUT that place is NOT toothpaste!

that one in the corner Silver badge

Re: It's not a GOLUM, either.

> coupled to ... huge databases that are demonstrably full of incorrect, incomplete and incompatible data

> Please don't glorify it.

Please don't glorify it by even vaguely implying it is coupled to a database - the LLMs aren't using anything more than a honking great pile (not even a carefully organised and categorised by humans who know what they are doing pile) of correlations of "this has been seen to follow that".

Human idiots have piped the output of LLMs into other tools, such as database queries, such as using them for web searches, but describing that as "coupling" the LLM to a database is as meaningful as (apologies for incorrect options usage)

cat text | grep isbn | awk -i reformat_isbn | sqlite all_my_books.db

and describing "grep" or "awk" as being coupled to the database.

Inclusive Naming Initiative limps towards release of dangerous digital dictionary

that one in the corner Silver badge

A "black box" is something that is magically doing its job - nobody knows how (and sometimes why) it works, it just does.

And we are quite happy leaving it that way.[1]

[1] If you open it, it will immediately stop working and all you'll find in there is a pile of dried leaves ((C) Colin Greenland)

that one in the corner Silver badge
Joke

Why isn't "token" on their list?

After all, it is only used to refer to "the token women" or "the token black" in any group, isn't it?!

that one in the corner Silver badge

Re: Fixing things due to a misunderstanding of the origin

As with any other lists of verboten words we'll start going down the euphemism spiral soon enough (as you point out, they've started out partway down the slide already).

A set of common replacements will be chosen, then they'll become obvious euphemisms and become as disliked as their predecessor, so a new replacement is chosen...

that one in the corner Silver badge

Re: Disambiguation

IIRC it was already in use when referring to any mechanism and had just spread to computers 'cos they are complex mechanisms as well.

that one in the corner Silver badge

Re: Mis-applied Words

> The phrase "master copy" is yet another mis-application of the word "master".

That isn't the wrong use of "master", it is a "wrong use" of "copy", if anything!

The phrase is simply a shortening of (along the lines of) "the master made this object from which all copies are to be made". The master piece (which may well be a masterpiece, although the copy itself may well be the masterpiece, if that is the requirement to gain one's guild masterhood) is the thing to be copied, whilst the wording "master copy" taken as-written implies that it is the *result* of the copying.

Although the master copy can, of course, be a copy itself: you can create a piece, use it as a master to make copies and then send each of those off to other sites where they will play the role of being the master copy against which those sites will verify their own copies.

You can also have some wordplay, as in "the master copy must copied slavishly", if one is allowed to.

"Master copy" is simply a contraction.

that one in the corner Silver badge

Re: And by "solving" a non-problem ...

> Or a neurodiverse person (whatever that means) saying how much (s)he resents "sanity test"?

I have a piece of paper that says I am "neurodiverse"[1] - and my go-to code toolkit proudly runs a SanityCheck() every single time. And does not hesitate to print out a message and halt[2] when it fails.

[1] actually, it tells me what my diagnosis is, not just some weird neologism

[2] I first wrote "kills itself" but one has to be sensitive these days

that one in the corner Silver badge

Get these people out of their offices and onto the farm

Or is considered bad now to segregate cattle?

https://www.kingshay.com/shop/segregation-gates/

MIT discovery suggests a new class of superconductors

that one in the corner Silver badge

Occhialini's continuum of materials

has the ring of something that will be the subject of a six lecture course during second year.

Hoping this research line bears fruit and maybe that'll come true.

Mark Zuckerberg would kick Elon Musk's ass, experts say

that one in the corner Silver badge

"Threads"

"For all your post-apocalyptic social media needs."

Very much the first thing that came to mind when I read that name; for anyone who isn't BBC watcher of a certain age: https://en.wikipedia.org/wiki/Threads_(1984_film)

Restaurant hired 'priest' to extract workplace confessions from staff

that one in the corner Silver badge

Re: Another job at risk - why confess to a priest when you can confess to God directly?

All I said was "That piece of halibut was good enough for Jehovah".

that one in the corner Silver badge

Re: Another job at risk - why confess to a priest when you can confess to God directly?

God does indeed have your full details, written down in the big book that St Peter can be seen reading from in all those cartoons; you know the ones, Peter asks "Name?" then looks to see if it is time to use the trapdoor lever or not.

The trouble is, God is really terrible with names, although He can remember a face[1]. So it is important to go to confessional, then He can look down, get a really good gander at you and put a face to all those peccadilloes He'd been hearing about[2]. Oh, and if you can make your sins unique, or at least interesting, that'll help a lot.

[1] And don't suggest adding ID engravings to the book, He'll just sulk. Something about never having got the hang of writing because it came after His Big Six Day Bash and the book is just so fiddly: He only managed to get four words to fit on an entire wall!

[2] Angels just love to gossip about goings on in Midgard[3], if only to give their four faces something to do

[3] Ah, yes, there was something else. He has got Himself some ravens and now self-identifies as Wednesday, so "now everything can be done mid-week and he finally gets a peaceful weekend". Raphael says there is nothing to worry about, but Michael is looking a bit strained these days.

BepiColombo probe turns to the dark side … of Mercury

that one in the corner Silver badge

Re: So

It's an elfie - deceiving Aeneas, pretending to be Mercury.

'We hate what you’ve done with the place – especially the hate' Australia tells Twitter

that one in the corner Silver badge

They'll just get the usual response from Twitter

A poop emoji

Apple squashes kernel bug used by TriangleDB spyware

that one in the corner Silver badge

populateWithFieldsMacOSOnly

How come the malware people manage to get decent programmers who know how to use meaningful names when some of us start to feel grateful if we get a CalcVal()?

I bet they even use naming conventions so you can at a glance if a line of code is using static, global, local or member variables!

Amazon Prime too easy to join, too hard to quit, says FTC lawsuit

that one in the corner Silver badge

Re: Different UI in America?

Indeed. As I said recently, multiple time I've signed up for Prime deliberately and used the 30 days, cancelling without any great difficulty; and I plan on doing so later this year, when they make available all the episodes I want to binge watch.

Ok, cancelling means going into the "My account" which is harder than joining - returning a product is about as difficult (i.e. not very, if you read what is actually on the screen), just as buying the product was easier than returning it. Not really a surprise.

Ok, I have hit the wrong button (the big one rather than the link beside it that plainly days "no, I don't want the magical experience" (or something like that) but it has always been possible to catch that before completing checkout.

Amazon certainly promote the heck out of Prime, with big happy buttons and banner ads, but then so does everyone who offers a subscription. But if you read what is presented to you, including what the checkout says (has the p&p just vanished?) then you can escape unscathed. If not, you have 30 days to cancel.

Weirdly, SWMBO has come to me, worried that she had signed up to Prime, and I've had to point out that, no, she hadn't!

BTW this is for the UK website - Amazon apps may be totally different.

OpenAI calls for tough regulation of AI while quietly seeking less of it

that one in the corner Silver badge

Re: Control bad applications, not bad AI

Oops, "high risk", not "evil".

Because it is only "high risk" to deliberately use the wrong tool for the job.

that one in the corner Silver badge

Control bad applications, not bad AI

All of the cases of "evil AI" given are nothing to do with AI but are simply applying the wrong tool to the problem, mostly[1] down to attaching the LLM to some i/o that it can run rampant with[2].

If I set up an automated tax portal and implemented it by code that just filled your form with Lorem Ipsum and and the Profanisaurus you would rightly say that my service is not fit for purpose and, if I charged you for and posted the forms you could sue. If I use ChatGPT then, again, you can sue. But it is clearly my service that is at fault: neither ChatGPT nor Profanisaurus are correct choices but they are not at fault.

Similarly, I can create any of the other "evil" scenarios without ChatGPT (and probably do it cheaper).

So, legislate against creating *any* stupidly dangerous application; the "AI" part is just theatre for vote grabbing and fear mongering.

BUT OpenAI will never support that, because people will realise that their product is not fit for any purpose and stop using it (after they've been sued enough times, hopefully brfore too much damage is done).

[1] exceptions being the "giving bad advice" situations, such as on mental health; but it is still the wrong tool for the job (or the problem was putting console i,/o on ChatGPT amd letting anyone loose on it).

[2] eg the (apocryphal?) example of letting LLM actually create Facebook accounts

Where's my money?! Now USA Today publisher sues Google over online advertising

that one in the corner Silver badge

Re: Feel the sympathy melt away

They are just saying *they* wanted a windfall and didn't get it.

There is nothing there to support any idea that Google got a windfall and failed to pass it on: Google just did what everyone expected Google to do.

Google have pretty much had a monopoly on online advertising for a long time now, there are no surprises there for anyone who did their due diligence.

that one in the corner Silver badge

Re: Feel the sympathy melt away

Yes, that is what *actually* happened, well done, we all know that.

Did you miss the bit where they *expected* a windfall and they suing because *that* didn't happen? The quote doesn't say "we just wanted a fair return for all our hard work".

that one in the corner Silver badge

Feel the sympathy melt away

> Internet advertising, Gannett's lawyers argued in their case, should have been a windfall for publishers shifting from print to online

We are supposed to be sad that you didn't get a windfall? So you are suing because you didn't get something for nothing?

Trot on.

Microsoft rethinks death sentence for Windows Mail and Calendar apps

that one in the corner Silver badge

Got that backwards, mate

> If you can't port a UWP app to the native toolkit, it's basically an admission that nobody should ever build native windows apps

Or a clear admission that you should have stuck with the native Windows and never written the UWP one in the first place.

Nowt wrong with Windows native[1], everything ends up invoking it right at the bottom. Plenty of additional toolkits available if you want extra fluff.

> Not even Microsoft is using their own toolkit.

Way to miss the point! They just want you on their cloudy thing, that is all.

[1] yes, yes, I know: "you are using Windows, that's what's wrong"!

Over 100,000 compromised ChatGPT accounts found for sale on dark web

that one in the corner Silver badge

Re: Is this worse that other products?

> Is there evidence that ChatGPT is worse than other products?

Well, it isn't anywhere near as tasty or nutritious as Ambrosia rice pudding. Nor does it work as well after you add Golden Syrup.

that one in the corner Silver badge

Re: Is this worse that other products?

No.

It would have been clearer if the article had pointed out that The Racoon info-stealer it mentioned is malware that infects individual PCs and grabs anything it can.

This isn't a report that OpenAI's servers have been breached.

In fact, the references to ChatGPT are only here to manufacture a headline: all the Racoon-infected PCs probably coughed up a lot more valuable logins than those for ChatGPT but there is nothing newsworthy about leaking bank accounts or Github credentials.

Another redesign on the cards for iPhone as EU rules call for removable batteries

that one in the corner Silver badge

> Apple and Google both mandate that developers target current OS versions (released within the last year).

Meanwhile, our favourite poster boy for "Why won't they learn to code!' - aka Microsoft - still allow the running of code built for Win2k (and will until they kill off the Win32 API).

The security fixes to the OS are "under the hood"; the API changes are (almost entirely) around adding new features (which you can ignore if you don't to bother with them) and those that aren't are given a backwards compatible form.

Hmm, pretty sure Linux can manage this sort of amazing feat as well.

Strange that Google & Apple can't manage backwards compatibility - do you think there could be any ulterior motive? Can't just be because they aren't as clever as Microsoft!

PS

Lest anyone then claim that any Win2k-compatible exe must be so old that it is itself insecure and dangerous, pretty much all the Windows exes I build are still Win2k compatible[1] whilst using decently up to date OSS libs[2] (so not just my weird code).

The only real "trick" is that some IDEs/compilers generate exes that insist on checking for newer OS versions, so they get options tweaked - or just don't bother with those unless it is really necessary to.

[1] yes, I still build 32-bit as well as 64: I've got perfectly functional 32 bit hardware chugging away, as well as the Win10 box. The 64-bit builds, curiously enough, usually end up XP compatible. Have to admit to not having fired up the Win98 laptop for a while.

[2] do have to sometimes put in a few #ifdef's or break a C file in two because the author didn't separate out their use of fancy new features I still don't care about from the actually useful bits of their lib, but that is simple enough.

Whose line is it anyway, GitHub? Innovation, not litigation, should answer

that one in the corner Silver badge

Re: Playing by the rules while making things better?

Interesting approach.

> Partitioning may affect the LLVMs ability to generate useful output

> If the target project is GPL, only the GPL LLVM[1] can be used

You could perhaps group these smaller LLMs that are GPL-compatible and try each member of the group in turn to see what gives a useful result. Although that requires discipline on the part of the User (or some automation in the IDE) to ensure the correct licence is applied to the correct lines of code.

Unfortunately, licences that require attribution would end up in one LLM all by themselves.

[1] if only this stuff was coming out of LLVM: those guys do good work you can use without all these arguments!

that one in the corner Silver badge

Re: Playing by the rules while making things better?

Yes. And?

I'm assuming you are positing this as a riposte to my description, but you don't say how or why that disagrees with what I described.

Although what doesn't help is that the quote said "copy code" when it ought to have said "replicate code" - subtle difference, but important.

There are two parts to that replication, which work together to both agree with Tim Davis's statement but also indicate how the LLM structure as it stands still makes it very difficult to generate attributions. What follows is all broad strokes (or we would be here all day) and considers an abstract chunk of code, not Tim's specific code (if for no other reason that neither of us has chased down. Also, I have absolutely no idea how much CS you (or anyone else reading this) has, so it is going to be really crude and simplistic and full of holes; apologies, but bear with me.

Part One

The learning process took Tim Davis's code (call it T) and started with basic lexing[1], breaking it into tokens. These are treated, along with all the tokens from the rest of the the training data, with processes like[2] "take all the pairs of tokens and count how many times each particular pair occurs; reject all the zeroes, convert the counts to a fraction of the total and you have the probability of such pairs occuring in a sample of your data". You can now run this "in reverse"[3] to generate output text that precisely comports to those statistics but which will be total gibberish.

If only you could also find wider correlations, such as "every { must be followed by - stuff - and then a }". If you could then the generated output will look more realistic. You may decide to say "those are just the syntax rules of the language, I can put them in explicitly"[5]. But that is starting to get hard, as you really need all the semantic rules, which even the best compilers don't encapsulate: e.g. for a certain commonly occuring nested loop structure {for (i=0; i<n; ++i) { for (j=i; j<n; ++j) {...}} you have to match the variables used instead of i and j, *but* not all commonly occuring nested loops start with j=i, some start j=0. The size of the problem is getting way, way out of hand to do manually: you don't have space to store all the tables that *accurately* describe the correlations and you are unlikely to even think of all the possible correlations, let alone how to express them.

So we turn to an algorithm that will (try to) find lots and lots of such correlations but, unlike the above, can't be guaranteed to find them all (it may not even find ones you consider obvious): this learning algorithm is generic, but by prefixing it with our above steps (i.e. reading text, lexing it) it magically becomes an LLM. But all it is doing is building up a massive graph with weightings along each transition which is a sparse version of the infeasibly big "n dimensional table of every correlation, accurately calculated" we tried to create, above.

Hopefully, with enough handwaving, you can see how we get to a model that generates really clever looking output *and* how that matches my earlier description of how hard it is to tag every node in the graph with the relevant attribution.

Part Two

BUT how can this ALSO replicate Tim Davis's code, apparently as a single chunk of licence-breaking text?

If we change our input reader slightly and instead of chunky lexemes we read the text literally bit-by-bit we can obviously do all the same work and get the same result (an cod-Copilot) but one that generates its output as a bitstream (which needs to be clumped back together as characters/codepoints to be legible to us) and is even less efficient in operation than the first version. The model no more and no less "understands" the bitstream as it did a lexeme stream.

Now go back to the first table of pairs of lexemes, ordered by probability and consider it is now a table of bit pairs, ordered by probability. As we've reduced the token space to two, we can afford the RAM to extend this from pairs to longer sequences and we can store this as a binary tree, rather than an n-dimensional array, still sorted by probability with the empty token as the root. This structure is now strangely familiar[6] as the basis of Huffman Encoding - i.e. we can see that this data could be used to simply (de)compress all of our input texts (if we fed all the original texts back in, the Huffman coding is now our lexer and if we feed its output into the generation process instead of using the random process, tada lossless compression and decompression).

Expand back out (whilst wildly waving hands even more) from a neat and complete binary tree to the sparse graph we had before, we can relate to that graph as a means of compressing the inputs: if we walk it in the *correct* order, we can (almost, sort of, because it is sparse and has gaps) think of it as decompressing the original texts back out again.

It now sounds like I have contradicted myself - if it is decompressing, then it stored the original, just like any other compression system, so it did just copy Tim Davis's code into itself and spit it straight out again! So why can't it admit that and add in the attribution correctly?

Well, the "decompressed" text is a bad copy: there are all those gaps in the graph and the random numbers used each time there is a choice - we aren't following one fixed *correct* order, it will (likely) be different each time. And, just to be really weird, there is also the potential for inefficiencies in the learning process that make it probable that constructs we "see" as being the same are replicated multiple times, possibly with irrelevant (to us, and to a compiler) details: how many ways are there to write a "for" loop in C? What about use of whitespace? "for" or unrolled as a "while"? It is entirely possible that 12 out of 17 next choices from this node may all result in for loops, just not quite the same one, so most rolls of the die do the same thing (as far as we are concerned, reading the output).

How bad a copy? Good question - not helped by not seeing *precisely* what Tim Davis saw [7] and how it compared to his original: was it character perfect or did it look like it had been effectively retyped? How many runs did he try and did they all generate the same thing?

As for the attribution: The same conditions apply as in my previous comment (each node in the graph is derived from all the inputs - precisely the same way that each node in a Huffman tree is derived from all the inputs). And remember all the variations on a loop? They came from different places...

Shoot Self In Foot, Destroy Own Argument

Having said all that, there *is* the possibility that one or more traversals never actually hit a choice point and will always generate identical outputs. And one or more of those *could* have formed from a weirdly unique input text, which has now been memorised and can be said to have been copied verbatim and, as it is the only thing that could be generated once that node has been reached, it could even be attributed.

Unlikely, but not impossible. Um, could even be due to a bug (i.e. the learning process isn't actually doing learning, it *is* doing compression).[8]

Footnotes

[1] intriguingly, so far I've not read if the lexing was adjusted from that used with the a natural language models into one that includes the lexing rules we apply to programming languages (i.e. the lexers in all our compilers/interpreters) - citations gratefully accepted.

[2] or totally unlike! Sorry, it gets messy being crude, and one problem is that you can not only do the crude reward/punishment process to slowly nudge the weightings until they (sort of) match the correlations, you can also pre-process the data by using faster (aka sensible) methods to generate the early, simple (and explicitly accurate) correlations between, say, token pairs and triplets - but you don't go very deep with this because, duh, the combination space grows in size and you run out of memory. Even the initial lexing step is one such pre-process.

[3] one array per token, including the empty token, fill it with the pairs starting with that token; sort by probability (randomly ordering equally probable items) then tag each entry with the cumulative probability from the start of the array, beginning with zero - each entry in the array is now in the bucket between its cumulative prob and that of the next entry. Start with the empty token, loop: generate random number 0.0 to 1.0, find the corresponding bucket, output the token in that bucket, set current token to that value, stop when you hit empty token again.[4]

[4] at this point we have the old programming exercise of the "rubbish generator", which done properly uses a decent Markov Chain, of course.

[5] you have also spotted that a parser's "follow sets" for a (programming) language can also be put directly into the generator portion of your Markov Chain assignment to produce "sort-of-recognisable program sources".

[6] still waving my hands, remember!

[7] if you can provide citations for that detail, please, please do

[8] In which case it isn't generating a proper LLM thingie in the first place and as I've only been rabbiting on about LLMs then all my arguments still stand. Phew.

Elon Musk's Twitter moves were 'reaffirming' says Reddit boss amid API changes

that one in the corner Silver badge

Re: Reddit's CEO doesn't realize

> other than having the biggest concentration of discussions in one place it has no advantage over the million other forums on the internet

To be fair, Reddit content does have the ability to totally screw up GPT models, as demonstrated here:

Glitch tokens (Computerphile) https://m.youtube.com/watch?v=WO2X3oZEJOA

that one in the corner Silver badge

Twitter? Profits? Ever?

> He also said that he often wondered why Twitter couldn't turn a profit under previous management

> my takeaway from Twitter and Elon at Twitter is reaffirming that we can build a really good business in this space at our scale

Implying that he believes Twitter is now making a sustainable profit under Elon.

Ok, have I missed something? Last I recall, Twitter was having problems getting the advertisers back, still owed a lot of money in unpaid bills and we were all waiting for the understaffing to really bite.

Not a great model to be talking about if you plan to IPO - *unless* perhaps he is trying to convince Elon to buy all the shares at IPO (after all, he bought one bag of flaming poo...)

Google warns its own employees: Do not use code generated by Bard

that one in the corner Silver badge
Facepalm

Re: The inverse of dogfooding

Typo alert!

Missed out the "or" at the end of the last sentence, should have read

"OR is there an existing well-known phrase I'm ignorant of?"

So asking for new suggestions, as well as any oldies I'm not aware of!