← Back

What you actually downloaded

Open weights are a binary you can patch but not rebuild. The argument worth having is not about that word, it is about the file underneath it.

Every model you pull down ships with two files that decide what you can do with it. One is the model card, which everyone reads. The other is the license, which almost nobody opens until legal asks.

The gap between those two files is where most of the confusion about “open source AI” lives, and it is a more expensive gap than the terminology argument people usually have instead.

Weights are a binary

If you have shipped software, you already understand what you downloaded. Weights are a compiled artifact. Billions of floats in a .safetensors file, executable by anything that knows the architecture, and completely opaque to reading.

Everything you can do with a binary, you can do with weights. Run it locally. Redistribute it. Patch it, which is what fine-tuning is, modifying behaviour without touching the source. Strip it down, which is what quantization is. What you cannot do is rebuild it, audit how it was made, or change a decision that was baked in during training.

We settled this argument once already. A binary you can download for free is freeware. Source is source. The open source movement spent the nineties making that distinction stick, and the AI industry borrowed the winning word without the thing the word described.

Where the analogy breaks, and why that matters

A binary compiles from source deterministically. Same input, same output, every time.

A model does not. Training is stochastic, sensitive to hardware, ordering, seeds and a hundred implementation details nobody writes down. Hand someone your full dataset and your entire training pipeline and they still will not reproduce your weights. They will get a cousin.

So open source AI is strictly harder than open source software. “Ship the source” is not a coherent request here, which is exactly the problem the people defining this had to solve.

What the bar actually is

The Open Source Initiative, the same nonprofit that defined open source for software, published its Open Source AI Definition on 28 October 2024. Four freedoms, to use, study, modify and share, plus three things a release must include: training and inference code, the parameters, and what it calls Data Information.

That last one is the compromise, and it is lower than people assume. OSAID does not demand the training corpus, because much of it legally cannot be republished. It asks for a description detailed enough that a skilled person could build a substantially equivalent system, plus a list of the public data and how to get it.

Even at that height, almost nothing clears it. In OSI’s validation exercise, five systems passed: Pythia from EleutherAI, OLMo from Ai2, Amber and CrystalCoder from LLM360, and Google’s T5. BLOOM, StarCoder2 and Falcon came close and were held back by license terms rather than disclosure. Llama 2, Phi-2, Grok and Mixtral did not pass. OSI treats this as a test of the definition rather than a certification scheme, so there is no badge, and every new release has to be judged on what actually shipped with it.

Llama, Qwen, DeepSeek, Mistral, gpt-oss, Gemma. All binaries.

And none of that is what will hurt you

Here is the part that gets skipped, and it is the part with money attached.

“Open” describes the artifact. The license describes your obligations, and the two have almost nothing to do with each other. A model can be nowhere near OSAID compliant and still be safer to build a company on than one that is closer, because the terms are what a court reads.

Those terms vary per model, not per lab, and they move between versions. Qwen’s open line, gpt-oss and Mistral Small ship under Apache 2.0. DeepSeek R1 and Phi-4 use MIT. Gemma spent years under a custom Gemma Terms of Use with acceptable-use carve-outs, then moved to plain Apache 2.0 with Gemma 4 in April 2026. Mistral keeps some models permissive and others under a non-production license. Reading last year’s blog post about a vendor tells you nothing about this year’s release.

Llama is the one to read line by line. The Community License is not OSI-approved and carries four real conditions. You must request a separate license from Meta if your product passes 700 million monthly active users, granted at Meta’s discretion. You must display “Built with Llama.” Any model you derive from it must have a name starting with “Llama.” And you cannot use Llama outputs to train a non-Llama model, which quietly rules out the most common synthetic data workflow in the industry.

None of that appears in a launch thread. All of it appears in the file.

Why the paperwork stays vague

There is a reason the data documentation, specifically, is the thinnest part of every release.

Courts have started splitting the question in two. In Bartz v. Anthropic, the judge found that training on books could be transformative enough to count as fair use, but that keeping a library of pirated copies was not. The case settled for $1.5 billion. NYT v. OpenAI is still live, Getty is still fighting Stability.

The exposure is not the training, it is the provenance. Weights are hard to interrogate. A documented data pipeline is not. OSAID asks for exactly the artifact that is currently being subpoenaed, which is most of why so few labs produce it.

The argument worth having

Whether a model is open source or open weight is a real distinction and a mostly academic one. If your goal is to run the thing yourself, tune it on your data and stop sending prompts to a company that may compete with you later, a binary was always enough.

The distinction that will actually reach your company is one file lower. We spent the nineties learning to ask what license something ships under before we build on it, and somewhere in the last three years we started accepting a word on a launch post instead. The word is marketing. The file is a contract.

Open it before you ship.