Remember open source is cancer
The FDA can simply ban all Open Source models as carcinogens.
OPINION Every six months or so a Chinese model sparks a panic, calling into question America’s AI dominance. Moonshot AI’s Kimi K3 is the latest example. Recall when DeepSeek R1 shook markets early last year? Following a similar pattern, Moonshot’s latest model isn’t all that interesting apart from its benchmark performance, …
Some people and businesses are willing to risk the use of open-weight models that might be poisoned to avoid being reliant on a specific chatbot-as-a-service provider. Given the likely consolation in the future and not knowing which one of Claude, ChatGPT, Gemini or Grok will be available, there is no alternative except to not deploy AI at all.
On the other hand, If I were deploying a custom AI agent in an enterprise setting I'd back it with Nemotron 3 or Inkling.
that a 3tn parameter model is runnable on a "small little network" is pushing the boundaries of that description pretty far
It's not a box in the corner of the office, sure. But the guidelines are ~64 accelerators (B200 or similar) to achieve sensible inferencing performance for production users. You can easily do that in a single rack (most such servers are 2.5-4 accelerators per OU, so 16-24OU), provided you can get 150kW in (and suitable cooling).
This is a big old line item, but also not hyperscale and well within the reach of PLCs/enterprises (who also might just rent time from neoclouds). In the context of the sort of business that can afford such a rack, it's probably a pretty noddy little network compared with their broader server fleet, office networks, production/manufacturing facilities, etc. If you're in the business of running moderate computational clusters for CAD/CFD/Modelling, then this will be a very expensive rack (given the RAM density/cost of the accelerator cards) but not really large or complex.
This post has been deleted by its author
I look forward with bated breath to hearing about the plan for how the US govt. plans to distribute that fine to international rights holders as part of restitution for how this US company has harmed them.
They are planning on doing that, right?
Because otherwise this just looks a bit like another racketeering cash grab that has zero relevance for AI vendors not on US soil.
Actually, at this point, "US law" is almost an oxymoron. Other than petty crimes committed by commoners, to fill the prisons, law in the US seems to have been replaced by a system of payments to a certain crime family originally from Queens, now based in the lawless land of Floriduh.
Good one!
First it has to be discovered.
Then someone with deep pockets and an incentive has to bankroll a legal process that can take years or decades.
Then you have to win.
Then you have to overcome any challenge to higher courts.
Then you have to get paid.
Then you have to distribute the payment to those affected. I don't think a general fine does that, does it?
But less competition inevitably means enterprises and consumers get screwed.
And guarantees increasingly effective and stealthy unknown, and ideally unknowable, enemy opposition ........ almighty phantom ghost adversaries ...... existential threat vulnerability exploiters/brokers.
And you might like to consider such as be guaranteed in the above is the natural unavoidable progression of future things no matter what courses of next actions be followed.
And to deny it possible and ignore the dire repercussions resulting has one fatally compromised and surprisingly easily overwhelmed and defeated/captured/captivated.
I'm not worried about lack of competition — the barriers to creating new models seem low, and there are lots of models getting created at a breakneck speed.
As to Chinese open weight models, I find them a very positive development. It seems like people have the freedom to choose various free "Linux" alternatives to paying "Windows" models, without moat or barrier to adoption. The fact they're Chinese rather than Finnish seems irrelevant, considering you'll be running it on your own infrastructure. What's not to like?
Training model still requires an impressive infrastructure, where even maintenance costs can be prohibitive. Chinese get away with it for the same reason they get away with relatively cheap electric cars for their internal market: government support/funding. That's not trivially replicated outside of China at the moment.
I think that's reasonable to argue that Chinese subsidies are competing with those offered by the US finance industry. For example, the rules for including the various AI companies in stockmarket indices have been relaxed. This virtually guarantees that tracking funds will be forced to buy the stock, effectively subvertint the market and providing a strong incentive for VC firms to keep giving the companies money, knowing they've got a guaranteed repayment when the IPO happens. This kind of funding has led to some of the circular investments, again signs of market disfunction, and preferential deals with utilities: you pay more for electricity because the data centres down the road got good deals.
But Chinese governments are no longer subsidising these companies as much anymore because they don't need to. Competition in China is so strong that they coined a term for it involution. This is the environment in which many Chinese sectors are operating and only the most effective companies will survive. Silicon Valley doesn't build companies for this kind of environment, instead it provides funding for companies in the hope that winner takes all and that, where it's not possible to beat the competition, you can just buy it. China is aware of this risk and has already vetoed the sale of Moonshot to Meta.
It's a mistake to see recent increases in market share by Chinese companies driven solely by subsidies and lower costs. If we don't admit that they have outcompeted us in many areas, we will never catch up.
The marginal price for inference is pretty much the price of electricity: China has been investing more in its grid over the last 20 years than America. It's not there yet, but it now does have some huge (even bigger than anything in Texas) wind and solar farms out west that could soon be plugged in with close to zero marginal cost. Providing the models as open weights provides added incentives to "try before you buy" – China doesn't really care because it knows the next generation of models are already in development.
In the real world the price comparison may be somewhat less impressive and speed for "real" tasks tends to be the determining factor for many. However, this will still "good enough" for many to want to pay either on their own hardware or somwhere else to run it. And this is despite all the handicaps that the Chinese developers are working against: limited hardware options and active restrictions in some cases. However, it could be that, as in evolution, it's precisely these restrictions that will make them outcompete. We're now starting to see the first systems that can use hardware optimisations on Huawei silicon. Again, China is generations behind both in software developmen and fab process, but it is iterating faster.
You'd be surprised at how fast a 26GB model version of Google's latest runs on a 12GB VRAM 4070Ti on a 128GB Debian host under llama.cpp. It isn't all that much slower than an online provider, and if you're looking at several files at the same time, it is significantly faster because it doesn't have to upload them over the internet. Now granted, it does make my CPU fans run a little, but it's not taxing my 16-core AMD4 processor all that much compared to a Java build, at which point they crank to full speed.
An economic analyst here in Australia (yes, I know, economic analysts have successfully predicted 20 of the last 3 financial crises ..) - has pointed out that whenever China moves into an industry space, profitability moves out..
Something about the corrupt and exploitive robber baron capitalists not being able to keep their ludicrous profit margins if competing against corrupt and exploitive state controlled communists..
I think it's worth adding that the competition within any particular industry is fiercest within China itself. Yes, subsidies do play a part in gaining market share for exports, but they don't explain the incredible pace of development in fields such as telecommunications,s batteries, electric vehicles, solar cells and more recently semiconductor manufacturing and, of course, LLMs. The CCP has so far tried in vain to intervene, because the lack of profitability does carry risks, though the same could be said of the US.
And trying to enforce some kind of ban open source is going to be about as successful as Canute's advisers suggesting he could command the tide. Though, I'm sure this is something that would appeal to Trumpty Dumpty…
The US is the most corrupt state in the entire world at this point, with it's administration blatantly taking payoffs from corporations and foreign nations to get what they want at taxpayer expense.
China, on the other hand, would have shot Der Pumpkin Fuhrer a long time ago for half of the crimes he's been convicted of, never mind accused of.
Google is going to give their LLMs away for free. Why would they care? They make money from adverts, and they certainly don't want other companies taking over.
You can download Ollama.cpp from GitHub, add a free Google LLM on Hugging Face and your computer will start talking to you.
Sure - it's not as good as the stuff you pay for... ...but seriously, for 95% of what LLMs are actually useful for, it's fine.
[Not trying to be difficult here, I really want to know. This forum seems as good as any to ask.]
1. Is there a way to verify "open weights"? How can one be sure they are the same as the "closed" ones that the Chinese military (potentially the Pentagon, Palantir, etc. - substitute your favourite villain at will, this is just an illustrative example) uses? What would be the scope of subjects/topics to test to validate all the 2.8tn parameters? Can it be done on a "zero-knowledge" basis, i.e., without access to the training and test sets?
2. How much in resources and money would it take an independent third party to verify benchmark results reported (as far as I understand) by model creators? Is it routinely done? Has it ever been done? Who are the trusted referees?
[I do realize the 2 questions are related.]
The models are just matrices of numbers. They cannot 'phone home' without that specific functionality being provided by the model server. Common model runners are llama.cpp and vllm, which are both open source and have no web connection capability. If you run the models locally then there is no chance they will steal your data. If you use the online providers then they probably are logging everything.
The weights cannot phone home by themselves, but the runner can. On my machine the Ollama binary imports socket/connect/bind/listen functions, links DNS/TLS-related system libraries, and contains net/http, crypto/tls, ollama.com, registry/pull/push, and localhost API strings.
Asking the model:
... As an AI language model created by Alibaba Cloud, my primary function is to engage in conversation with users and provide insights based on my understanding of various topics. I don't have direct access to external internet-based servers or any specific programming tools like those found in a host program. ...
nm -u shows direct socket/DNS symbols:
_socket
_connect
_bind
_listen
_accept
_sendto
_recvfrom
_sendmsg
_recvmsg
_getaddrinfo