Re: Am I the only one ...
I think the CSAM ingestion accusation is a bit of a stretch. Most diffusion models can generate fairly photographic images of clearly impossible things by bodging together known concepts (e.g. A pink elephant in a rice paddy or a Ferrari drawn in the style of Da Vinci).
To do these things, the model would have to have been training data that allowed it to identify an elephant, determine what colour pink is, have source material of rice paddies, and Ferraris, as well as at least a representative sample of the corpus of work by Da Vinci.
Similarly, to render a naked person, it needs to have had at least some source material representing what people look like under their clothes, or it's not going to know what genitalia look like.
I've not seen the supposed output of Grok (I've never even used Grok, since I won't touch Elmo's stuff on principle), so I don't know how realistic or convincing the output is, but the reporting would suggest that it is accurate enough to cause concerns, which implies that it has been fed accurate training data.
Once again, a statistical model cannot produce output that does not meet the patterns of its input. In simple terms, it can either provide an "average" of the data it is provided, interpolate between two data points, or extrapolate from a series. If it hasn't been trained on images of naked people, then this would be a case of extrapolation, and I simply do not believe that an unthinking piece of computer software would correctly extrapolate these images. The conclusion, then, is that it is either "averaging" from a large number of images of naked people, or interpolating between, say, the image of one person clothed, and another naked. You cannot interpolate between two points without knowing both points. This is, of course, a simplification, but the thing here is that "AI" cannot create anything new, and anyone who tells you it can is swindling you.