I usually don’t even understand the lingo they use. “Open-weighted” is the most recent one, then it usually goes down to specific “models” that everybody is supposed to know about.
These are my thoughts (I will stick to the vague “it” for now, but of course therein lies another question: “and how does all this apply to various specialised AIs”):
- Is it really feasible to run it 100% locally? I know there’s plenty of people with very powerful rigs indeed, but still. Or are 99% of these people really saying “it would, in theory, be possible to run that locally, therefore your concerns are invalid”?
- If yes to the previous: the software doesn’t come from nowhere and ultimately still relies on gas-turbine-powered datacenters and stolen IP and stolen personal data, no?
If what I wrote above is true, what exactly are people arguing when they say it’s still possible to use LLMs ethically or true to FOSS philosophy, because … ???
edit
Thanks to all who answered.
I guess it’s my fault for asking several questions in one, but this thread has attracted exactly the type of people I’m writing about; several even used the term “open-weighted models” without explaining it.
Asking to get arguments explained, I got more arguments instead.


To answer a lesser answered sub question of a question: what is a weight?
Imagine you have an output, its either a or b. You have an input, whatever it is. And then there is a weight in the middle. The weight determines wether the output will be a or b.
In math it would be a function input*weight=output.
ANN (Artificial Neural Networks, just afancy name) is a chain of a whole bunch of weights. There are layers and each layer generates an output from a bunch of weights and the next layer uses it as inputs. The name comes from the visualization, which looks like neurons firing to other neurons (from biology, things that make your brain work).
Oh and the connection between LLMs, Tokens and ANNs.
ANN = LLMs, LLMs are a subset of ANNs that have the purpose of doing anything language. Most times you’d want maybe to predict a single thing (e.g. a temperature, a color, a letter). The specific output depends on the usecase. LLMs simply output characters and predicts one character at a time whats the most likely output.
To make sense of this nonesens you’d need to output and input things that are larger than a single character. These are tokens. I don’t know much about that, but basically it is the whole reason why LLMs took off and they got invented by some research done by google in 2017 (idk the title of the paper anymore, maybe something along the lines of “a new way of thinking”). Funnily enough google ignored the paper and some guys at google got mad and founded their own AI company cough Anthropic cough
Anyways tokens let the AI-Model predict 4 characters or so at a time, which enables the LLM to predict actual words.
I am not sure exactally how it works from there, maybe it gives the whole in- and output so far back into the model as a new input to generate the next token? Sounds kinda inefficiant, but the whole thing is just fucking inefficient and useless. AI has usecases but LLMs are billionare madness in terms of calculation power needed.