That Cactus Needle sounds quite cool. I mean, if you're not interested in generative stuff (making up some volubile 'literary' text outputs) then it makes sense you don't need the MLP FFNs that commonly make up 2/3 of 'model' parameters [ https://github.com/cactus-compute/needle/blob/main/docs/simple_attention_networks.md ]. Even vibed code and cyberattacks come from strapped-on agents rather than the genAI bits of mammoth LLMs (istm) so ...
Keeping only the attention-subsystem-oriented syntax analysis (quey processing, 'prompt' decoding, NLP) looks promising for starting to remove some of the ponderous excess baggage currently stuffed-up in this most distending of techs, on the way to 0-token where feasible (eg. most everywhere).
The TFA-linked Gemma 4-E2B-it piece already shows one can reduce full LLM size down to 1/5ᵗʰ of prior sweltering efforts (27B to 5B) with no reduction in 'performance', so taking out parts that rank from useless to just plain annoying should be perty much SOP ATM. Especially since all benchmark 'results' shown whenever a new 2x, 4x, 10x heftier model is introduced always end-up rather saturated, suggesting the improvement relative to smaller prior tools is actually evermore just plain minimal, iiuc (i.e. from saturated, to saturated, again, wtf).
TL;DR such cactus needle might well puncture the AI bubble and put girdled procedural 'agents' (aka rather normal code) back in the forefront afaics. A sort of waste not, want not, counterpunch! ;)