The licence is the specification
Open weights are not open source, and the difference decides what you are allowed to ship.
10 minOpen-Source AI Models
Most models described as open are released under bespoke terms written by the lab that trained them. They are not OSI-approved licences, and they routinely carry conditions that a procurement team will care about long after the benchmark scores stop mattering.
We read the current terms for eleven widely deployed open-weight models. No two were the same, and four contained obligations that would surprise an engineer who assumed the label meant what it means in software.
The clauses that bite
- Usage thresholds that convert a free licence into a negotiated one above a certain number of users.
- Attribution requirements that extend to the product interface, not just documentation.
- Restrictions on training other models on the outputs — which affects distillation directly.
- Jurisdictional carve-outs that make deployment in some markets a legal question rather than a technical one.
Why this is getting harder
Distillation has made provenance a chain rather than a fact. A model fine-tuned on synthetic data generated by a second model inherits constraints from both, and the second model's terms are frequently silent on the case.
Our legal review took longer than the evaluation. That ratio is now normal and nobody budgets for it.
The practical step is to record the licence alongside the checkpoint hash in whatever registry you use, and to re-read it at each upgrade. Terms change between versions of the same model family more often than teams expect.
What to watch
Whether any major lab moves to a genuinely standard licence. Several have said they are considering it; none has done it.