Re: I, for one, welcome...
Certainly !
221 publicly visible posts • joined 22 Jun 2012
pure gold LOL ". ... Of course, I'd also suggest that whoever was the genius who thought it was a good idea to read things ONE F*CKING BYTE AT A TIME with system calls for each byte should be retroactively aborted. Who the f*ck does idiotic things like that? How did they noty die as babies, considering that they were likely too stupid to find a tit to suck on?"
I recently dabbled with ollama on CPU only 64gb old cheap server loading for example Gemma3 37GB model for cpu infrrencing. It _is_ slow but say 10 minutes for complete answer is not eternity either.
From what I understand ollama can work hybrid, load as much to gpu vram as possible, and the rest to system RAM.
This way you could still use gpu, and models bigger than its VRAM
I am still noob, mind you
Tested on Ubuntu 22.04