Mozilla: Open-Weight AI Closes Gap, Lags on Deployment
Recurring report from Mozilla, The state of open source AI, v1.0.1, July 2026, introduced by a letter from Raffi Krikorian (CTO): seven sections, an interactive site, and a downloadable report.
By **Mozilla** — éditeur du rapport// Source stateofopensource.ai ↗/Reading 2 min/.md// Auto-verified translation
#Mozilla#state of open source AI#open weights#open weights#open source AI#OSI definition#training code#data documentation
Recurring report from Mozilla, The state of open source AI (v1.0.1, July 2026), introduced by its CTO Raffi Krikorian.
The thesis opens the first section: « The model layer has commoditized. Value accrues to the harness above it. » Inputs that have become commodities lose their pricing power, and the majority of production workloads run well below the frontier ceiling.
A harness tuned tightly to one lab's weights… degrades on anyone else's model, so the tighter the tuning, the less swappable the weights underneath. Lock-in arrives as a side effect of optimization.
— **Mozilla** — éditeur du rapport , stateofopensource.ai
Capability state. On the Artificial Analysis Intelligence Index, the best closed model scores 61 (Claude Opus 5), the best open model 57 (Kimi K3), fourth overall; on the Epoch Capabilities Index the gap is six points, "about one release cycle", with overlapping confidence intervals. The frontier is sawtooth: open leads in frontend code, contests terminal agentic work, and clearly cedes ground on professional knowledge work.
The usage shift. The share of OpenRouter tokens routed to open weights rose from a negligible level to a majority by mid-2026, with the seven highest-volume models all open — but the report notes that by request count, closed providers still lead, the open lead being a token-volume lead concentrated in coding and agentic workloads.
The central finding: « Open ships easy. Open deploys hard. » 79% of developers use open models versus 71% closed, with half using both; but only 53% of open teams reach production versus 63%, and the gap widens with company size, which rules out an explanation by resources. The stack map confirms it: two cold columns across every layer, standardization and enterprise readiness.
The harness is the new frontier.« The agentic harness is another user agent » — the browser's role replayed one layer up. And the lock-in mechanism is stated precisely: a lab's harness, tuned to its own weights, degrades on everyone else's, so « the tighter the tuning, the less swappable the weights underneath. Lock-in arrives as a side effect of optimization. »
Sovereignty is framed as a right to exit, illustrated by Fable 5's nineteen-day blackout over export controls: « You can switch off a model. You cannot switch off a copy already running on a machine you hold. »
Mozilla advocates for what it measures. Scrupulous captions and a self-stated reversal watchlist make the data usable; the framing remains a thesis.
Key takeaways
Date / source.The state of open source AI, Mozilla, v1.0.1, July 2026, introduction by Raffi Krikorian (CTO). Third-party sources credited (Artificial Analysis, Epoch AI, OpenRouter, LMArena) and proprietary survey with SlashData (n = 1,410 on adoption barriers).
Key framing.« The model layer has commoditized. Value accrues to the harness above it. » Justification: « Commodity inputs surrender pricing power », and the majority of production workloads run well below the frontier ceiling. ### The terminology distinction, worth keeping | Term | What it covers | |---|---| | Open model | downloadable weights, executable and modifiable on hardware you control | | Open weights | parameters under a permissive license, without training code or data documentation — « which describes most of what this report measures » | | Open source AI (OSI sense) | additionally requires training code and enough data information to reconstruct the system | The report is titled open source AI but overwhelmingly measures open weights. It says so; the takeaways won't. ### Capability gap, two consistent instruments | Instrument | Closed | Open | Gap | |---|---|---|---| | Artificial Analysis Intelligence Index v4.1 | 61 (Claude Opus 5) | 57 (Kimi K3) | 4 pts; K3 4th out of 586 models, 3 of the top 11 open-weight | | Epoch Capabilities Index | 162 (GPT-5.6 Sol) | 156 (K3) | 6 pts ≈ one release cycle, overlapping confidence intervals | On price, K3 sits 3.6 points off the top for roughly a third of the price. ### The sawtooth frontier Open leads in frontend code (K3, 1,679 Elo, six domains out of seven on LMArena Frontend Code Arena); contests terminal agentic work (88.3 versus 88.8 on Terminal-Bench 2.1; K3 wins Program Bench, SpreadsheetBench 2, and BrowseComp, loses FrontierSWE 81.2 versus 86.6); cedes ground on professional knowledge work (Fable 5 leads by 92 Elo on GDPval-AA v2, the largest gap among shared benchmarks, and Moonshot concedes it). Hence: « Match the model to the job and you need the frontier for less than you think. » The report warns that scores use different scales and mostly come from vendor-run tests — « treat them as directional ». ### The most misquoted figure Open weights rose from a negligible base to a third of OpenRouter tokens by late 2025, then to a majority by mid-2026, with the seven highest-volume models all open-weight (72.4% of top-20 volume held by open-weight ranks 1-10). But: « By request count, closed US providers still lead. The open lead is a token-volume lead, concentrated in coding and agentic workloads », and the shares cover only the top 20. A majority of tokens is not a majority of use cases. ### Deployment over capability « Open ships easy. Open deploys hard. » 79% of developers adding AI use open models versus 71% closed, and 50% use both (29% open only, 21% closed only): the two categories are complementary, not rival. But 53% of open teams reach production versus 63% closed, and the gap widens with size — closed 54% → 73%, open 53% → 57%. « Scale rules out a resources explanation. Enterprises can buy their way through closed deployment. Open deployment waits on tooling that remains unfinished. » Named barriers, all operational: infrastructure cost, security and compliance, maintenance, deployment complexity. The stack map (48 components, 9 layers, 10 criteria) confirms this from another angle: two consistently cold columns across every layer — standardization and enterprise readiness. ### Section 5: the harness as user agent The analogy is the browser, « code on the user's side negotiating with servers on their behalf », replayed one layer up. The harness is « where production difficulty concentrates, and where the open-vs-closed, owner-vs-renter contest restarts ». Five mapped layers: Govern (stateful policy, registry and lineage, budget and revocation), Surface (AG-UI/A2UI, x402/AP2/UCP), Action (E2B/Daytona/Modal sandboxes, permission and identity — « the unsolved write surface » —, eval and observability), Reach (MCP, A2A, memory), Control (LangGraph, CrewAI, AutoGen, LlamaIndex). The phrase « the unsolved write surface » echoes the diagnosis in [[valente-zalewski-beyond-zero-enterprise-security-ai-era-2026-07-20]]: read is solved, write is not. ### The lock-in mechanism « The model is eating the harness. » On every model where a lab's own harness and an independent harness coexist, the former now wins, the 21.8-point gap having compressed to about 3 at the top. Hence: « A harness tuned tightly to one lab's weights becomes a fitted component of that lab's product. It degrades on anyone else's model, so the tighter the tuning, the less swappable the weights underneath. Lock-in arrives as a side effect of optimization. » Lock-in doesn't need to be a strategy: optimizing is enough. Compare with the portability contract in [[janakiram-agent-platform-portability-contract-2026-07-20]]. ### Traction and economics LangChain 126,000+ stars and 60% developer share; MCP at 97M monthly SDK downloads and 10,000+ active servers within a year, +4,750% in 16 months, handed to the Agentic AI Foundation in December 2025. Governance gap: only ~21% of enterprises report mature agent governance. On pricing, inference dropped 50x in 36 months at GPT-4 level (versus 2.6x for bandwidth during the dotcom era), the frontier price having fallen 112x since GPT-4's $45. On OpenRouter (May-Sept. 2025), closed held ~80% of usage and ~96% of revenue — at quality parity, it costs roughly 6x more per call, implying an estimated ~$24.8B in unrealized annual savings (Nagle-Yue study for the Linux Foundation). ### Sovereignty and the right to exit More than 70 active national strategies, and « the strategic case for open is the ability to leave », backed by the cloud precedent ($90-120K to exit a petabyte from S3, 37signals down from $3.2M to under $1M, 80% of enterprises repatriating). « Closed model APIs reproduce the same trap… Open weights are exit rights. »The nineteen-day blackout makes the argument concrete: June 9, Anthropic ships Fable 5 and Mythos 5 → June 12, Commerce applies export controls with immediate effect, barring access to any foreign national, including Anthropic's own employees; nationality being unverifiable in real time, both models go dark for everyone → June 26, partial restoration of Mythos for ~100 US critical-infrastructure organizations → June 30, lifted → July 1, Fable 5 restored. Then July 16, Moonshot opens K3's API. « Access can be revoked and restored. A weight release cannot be withdrawn once the files are distributed… You can switch off a model. You cannot switch off a copy already running on a machine you hold. » ### China Qwen surpassed the next eight organizations combined in Hugging Face downloads by February 2026; Chinese open-weight models rose from under 2% of OpenRouter tokens in late 2024 to over 45% of weekly traffic by April 2026, ~61% among the ten most-used models. DeepSeek claims 26,000+ enterprise accounts and was part of 58% of new AI startups' stacks in 2025, even as at least eight jurisdictions restricted the hosted service: « Enterprises ban the hosted app and adopt the weights anyway. » ### The K3 distillation affair, by the report's own categories | Level | Content | |---|---| | Confirmed | Anthropic's February 2026 disclosure — ~24,000 fraudulent accounts, over 16M exchanges of which 3.4M attributed to Moonshot, against earlier Claude models; statements from Kratsios (July 22) and Bessent (21) | | Signal | Greenblatt (Redwood Research): K3 self-identifies as Claude in a way statistically hard to explain as noise — but it names a model predating the affair, and self-identification is a known artifact of training on web text | | Proof | absent — « No logs and no forensic package. Moonshot denies. » | The report notes that weight publication now makes behavioral forensics possible. The allegation should never be presented as established. ### The watchlist and its method Four families of signals (capability/adoption, harness, market structure, trust/safety), each paired with its own reversal condition — « Reverses if: the lab-harness lead widens, or a closed platform sets the permission standard first ». Stating in advance what would prove the thesis wrong is what sets this document apart from an ordinary advocacy piece. ### A note on citation caution Mozilla is measuring a subject it advocates for. The rigor of its captions and the presence of reversal conditions make the data usable; the framing — commoditization achieved, openness inevitable, « open won » — is a thesis, not a finding.
Key figures
57 on the Artificial Analysis Intelligence Index v4.1, versus 61 for the best closed model
the model layer has become commoditized and value is shifting up to the agentic harness
— Mozilla
the open-weights lead is a token volume, not a request count, concentrated on coding and agentic use
— Mozilla
no log or forensic record supports the distillation allegation, which Moonshot denies
— Mozilla
Kimi K3 aurait été entraîné par extraction covert à grande échelle depuis Fable 5
— Michael Kratsios
The knowledge graph extracted from this fiche — 8 entities, 25 relations.
In this graph :The state of open source AI · Mozilla · Raffi Krikorian · open-weights · écart opérationnel · verrouillage par optimisation · surface d'écriture non résolue · droit de sortie