Skip to content

root / tags / moe

#MoE

2 fiches

Tools & Platforms Auto-verified translation

Kimi K3 de Moonshot AI : quand le frontier open-weights rattrape le propriétaire

SFEIR's engineering-cabinet analysis ("an engineer's reading") of the **July 16, 2026** launch of **Kimi K3** by the Chinese laboratory **Moonshot AI**: an **open-weights, frontier-class model** whose provider claims **~2.8 trillion parameters**, a **one-million-token context**, and **weight release before July 27, 2026** (likely under a Modified MIT license, as with the K2 lineage). Thesis: capability once thought reserved for proprietary giants (Anthropic, OpenAI, Google) is becoming available **in open weights, at a discount price, from a Chinese lab**. SFEIR — despite being an **Anthropic and Google Cloud partner**, and thus "with no interest in oversell­ing a Chinese model" — adopts a cardinal **methodological caveat**: on launch day, **no official, complete benchmark table** exists; specs (2.8T, Kimi Delta Attention, +25% training efficiency) and scores are **vendor-stated** or drawn from **community arenas**, "to be treated as claims, not measured facts." The new architecture (**Kimi Delta Attention**, hybrid linear attention; decoding claimed up to **6.3x faster** at 1M tokens) breaks with the K2 cadence (K2 Jul. 2025 → K2.7 Code Jun. 2026, a flagship every two months); two variants accompany the launch (**K3 Max**, **K3 Swarm Max**), with forced sunsetting of the kimi-k2.5/moonshot-v1 series on **August 31, 2026**. **The real weapon is price** (~$3/M input, $0.30 cached, $15 output per secondary sources): a frontier open-weights model at this level **pulls the whole price-performance curve down** — the commoditization of the model layer, accelerated by open source. But the decisive singularity is not a score: it is **reversibility**. A frontier open-weights model turns a consumed API (vendor dependency) into an **option** (self-host, portability, exit from lock-in), at the cost of heavy infrastructure to host 2.8T parameters. SFEIR's view: **open-weights changes the question, not just the answer** — no longer "which model is best/cheapest?" but "how much of my system am I willing to make dependent on a vendor I don't control?". The right posture remains a **routed portfolio** (one model per task, one model per constraint), with Kimi K3 adding a **"reversibility" column** to the decision grid. The "AI Only" conviction stands unchanged: the model is a commodity, the durable advantage lies in the engineering around it (Context Engineering, harness, cost governance, ability to change one's mind). The figures still need validating "on your own" — your repositories, your data.

#Kimi K3#Moonshot AI#Yang Zhilin

SFEIR (voix éditoriale du cabinet)

AI Coding Agents & Skills Auto-verified translation

The Batch n°350 — How Coding Agents Accelerate Different Types of Software Work (Andrew Ng) + GLM-5.1, Digit chez Schaeffler, anti-data-center revolt, assistant axis

Andrew Ng's editorial in The Batch #350 sets out an **acceleration hierarchy for coding agents** by type of software work: **Frontend (max) > Backend (moderate) > Infrastructure (low) > Research (minimal)**. The rationale rests on implicit *verifiability* (fluency in TypeScript/JavaScript plus an autonomous agent–browser test loop on the frontend) and on the LLMs' blind spots (corner cases / security / DB migrations for backend, opaque network tradeoffs for infra, irreducible hypothesis formation for research). The issue is rounded out by 4 structuring news items: **GLM-5.1 (Z.ai)**, a 754B/40B-active-parameter MIT-licensed model capable of autonomous tasks lasting 8 hours (SWE-Bench Pro leader at 58.4%); **Digit (Agility Robotics) at Schaeffler**, the first industrial deployment of humanoids (5'9"/143lb, $10–25/h vs $20/h for a human); the **anti-data-center revolt** (~$64B blocked May 2024 – March 2025, Maine moratorium on 20MW+ facilities, molotov cocktail at Sam Altman's home); and the **"assistant axis"** (Christina Lu, MATS / Oxford / Anthropic), which reduces persona drift and jailbreaks (Qwen3 32B: 83%→41%; Llama 3.3 70B: 65%→33%) without degrading IFEval/GSM8k/MMLU-Pro/EQ-Bench.

#Andrew Ng#The Batch#DeepLearning.AI

Andrew Ng (édito principal — fondateur DeepLearning.AI, Stanford, ex-Google Brain, ex-Baidu) ; rédaction The Batch (DeepLearning.AI) pour les sections actualités