Skip to content

root / tags / context-engineering

#context engineering

13 fiches

AI Coding Agents & Skills Auto-verified translation

The AI Engineering Skills Map

X post by **Andrew Ng** from **August 14, 2026** (16:29 UTC), reprising the "Dear friends" letter from ***The Batch* #366** (DeepLearning.AI, same date), ~900 words. Ng presents **The AI Engineering Skills Map** and publishes **four skills** held to be the most important. **(1) Building and deploying AI applications** — the specificity is named: *« The key difference between AI and non-AI applications is that the former has unpredictable outputs »*, hence the emphasis on *evals* and error-analysis loops. **(2) Software engineering fundamentals**, because *« Understanding software fundamentals allows you to recognize what tradeoffs even exist »* — the inexperienced developer fails *« because they don't know what context to give their coding agent »*, hence the goal of *« steering coding agents using the precise language of software engineering »*. **(3) Using coding agents**, in an operational formulation: *« help the agent autonomously close loops by providing verifiers or evals »*, and *« knowing how much to intervene and how much to leave them alone »*. **(4) *Shaping the build***: *« Given a clear spec, coding agents are rapidly improving at delivering to it. Thus, our work as engineers is shifting toward deciding what should be in the spec »*, paired with *« Engineers should no longer expect to be given a pixel-perfect design and asked only to implement it. »* A **terminology note** carries most of the framing: Ng talks about **skills** in AI engineering and **not the role** "AI Engineer", with an explicit analogy — *« All developers today should know how to work with the cloud, and only a smaller number have a "Cloud engineer" title. »* The whole is backed by *« an analysis of more than 10,000 job postings, dozens of structured interviews with experts, hiring managers, and recruiters, surveys, and other online data »*, of which **no numeric results are published**: Ng describes his process as *« informally… akin to running clustering »* and announces a detailed map in future posts. He states the interest in the second-to-last sentence: *« DeepLearning.AI's principal focus is to help developers gain these AI engineering skills. »*

#AI Engineering Skills Map#skills map#Andrew Ng

**Andrew Ng** — fondateur de **DeepLearning.AI** · general partner d'**AI Fund** · cofondateur de **Coursera** et de **Google Brain** · ancien chief scientist de Baidu. Texte signé · à la première personne · écrit *« with my team »* sans qu'aucun collaborateur soit nommé. Publié le **14 août 2026** sur X et dans ***The Batch* n°366** — même texte aux deux endroits ; préférer *The Batch* pour toute citation durable. Quatrième fiche Ng du corpus · après les lettres n°350 (24 avril) · n°352 (8 mai) et n°359 (26 juin).

AI Coding Agents & Skills Auto-verified translation

Mon usine logicielle à l'heure de l'IA

Reference page published on **eventuallycoding.com** on **July 28, 2026** by **Hugo Lassiège** (Lyon, developer turned entrepreneur, author of Bloggrify, Hakanai, and Writizzy). The author announces it as such: *"This will be more of a reference page than an article,"* intended for his own resources page. **Subject**: an exhaustive, tooled description of a **solo software factory** where *"the code produced is now nearly 100% generated,"* across several polyglot monorepos (Nuxt, Kotlin, JS — Hakanai, Writizzy, Bloggrify) in **continuous deployment to production**. **Distinction stated upfront**: this is not **vibe coding** in Karpathy's sense (experimentation, letting oneself be carried along) but **context engineering** — *"giving all the necessary context, at the right time, so that the software matches an intention and is systematically controlled,"* with the sentence that grounds the responsibility: *"Even if I don't write the code, I am responsible for it and must keep control over it."* **The entire toolset answers three questions**, and this is the text's most reusable reading grid: *"What does the agent know?"* (context, memory, code graph) — *"What does it know how to do deterministically, without improvising?"* (skills, procedures) — *"What stops it when it gets it wrong?"* (hooks, architecture tests, quality gates). **Six layers detailed**: (1) **context** — root `CLAUDE.md` + topical `.claude/rules/*.md` conditionally loaded via `paths:` + `.agents/*.md` for non-technical matters (personas, positioning, tone); (2) **skills** — about thirty, existence criterion *"if I explain the same thing a third time"*; (3) **tools** — JetBrains IDE MCP, **GitNexus** (code graph: `impact(symbol)`, `detect_changes()`), Claude-mem, RTK filtering wrapper, Sentry, read-only database; (4) **executable guardrails** — harness hooks, **architecture tests**, pattern linting (**ast-grep** for architecture decisions, not just ESLint); (5) **factory** — blocking quality gate with `needs:` on the quality job, five test stages; (6) **product process** — numbered specs with a drafting skill **and a closure skill**, design in Claude Design, staged delivery behind feature flags, distinction between **feature flipping** (Unleash) and **gating** (customer contract). **The rule that sums it all up**: *"What matters must be executable. An instruction is followed 'most of the time'… A hook or a test is followed all the time."* **A rarity for the genre**: a "To improve" section that exposes four lived limitations — the **impossibility of measuring a rule's obsolescence** (*"I have no way of knowing whether an old rule has become obsolete"*), the **rabbit hole** created by a boyscout rule, the **lack of packaging** for skills across projects, and above all the admission of tension: *"I am becoming less and less useful during implementation phases,"* *"torn between the satisfaction of having an increasingly efficient factory and the risk of losing knowledge."*

#software factory#context engineering#vibe coding

**Hugo Lassiège** — développeur devenu entrepreneur · basé à **Lyon** · écrit du code depuis 2001 et tient **eventuallycoding.com** (le blog a porté le nom `hakanai.free.fr` avant de devenir *Eventuallycoding* en 2013). *Eventuallycoding* est le nom-parapluie qui regroupe ses projets · sa chaîne YouTube et ses blogs.

Transformation & Adoption Auto-verified translation

IA et emploi : le vrai risque, c'est le décrochage

In-depth opinion piece published on **sfeir.com** on July 23, 2026, signed by **SFEIR** (the firm's editorial voice). It is a **strategic commentary on Trésor-Éco note No. 391** from the DG Trésor (June 2026 — see [[dgtresor-ia-effets-emploi-2026-06-30]]), read through SFEIR's doctrine of « **amplifying AI rather than enduring it** ». The article praises Bercy's **cautious economist's tone** (mechanisms plus uncertainty rather than a prediction) and draws from it a **three-part thesis**: (1) **no measurable aggregate effect** at this stage (two offsetting forces — displacement vs. productivity — EU adoption ~20%); (2) a **single solid empirical signal, on juniors** (−16% employment among exposed 22-25 year-olds in the US); (3) a **long-term danger that shifts the question** — **competitive lag** (non-adoption), not job destruction. The analytical core SFEIR retains: **price elasticity** determines the employment effect (the **Jevons** paradox applied to code) → the argument is **structurally pro-employment for developers**. The article **dismantles the "AI layoffs" narrative** (4.5-6.2% of US layoff announcements, "labeling" at 59%) and points to the note's **blind spots** (the agentic scenario relegated to a footnote; diffusion speed not discussed; OpenAI/Anthropic having become sources for Bercy = an unflagged source bias). **SFEIR's operational translation** (for CIOs/CTOs): value migrates toward intent/architecture/control, training **augmented engineers** (**AI Champions** programs), and avoiding rushed adoption (**workslop**, technical debt) through **context engineering** and governance.

#AI and employment#competitive lag#non-adoption

**SFEIR** — ESN française « AI Only » (~850 ingénieurs, 8 agences France & Benelux). Voix éditoriale du cabinet (byline « SFEIR »). Positionnement de la maison sur la transformation IA des DSI ; ce texte prolonge la ligne éditoriale portée notamment par Didier Girard (cf. [[girard-sfeir-ai4it-vs-ai4business-budgets-2027-2026-06-24]]).

Tools & Platforms Auto-verified translation

Kimi K3 de Moonshot AI : quand le frontier open-weights rattrape le propriétaire

SFEIR's engineering-cabinet analysis ("an engineer's reading") of the **July 16, 2026** launch of **Kimi K3** by the Chinese laboratory **Moonshot AI**: an **open-weights, frontier-class model** whose provider claims **~2.8 trillion parameters**, a **one-million-token context**, and **weight release before July 27, 2026** (likely under a Modified MIT license, as with the K2 lineage). Thesis: capability once thought reserved for proprietary giants (Anthropic, OpenAI, Google) is becoming available **in open weights, at a discount price, from a Chinese lab**. SFEIR — despite being an **Anthropic and Google Cloud partner**, and thus "with no interest in oversell­ing a Chinese model" — adopts a cardinal **methodological caveat**: on launch day, **no official, complete benchmark table** exists; specs (2.8T, Kimi Delta Attention, +25% training efficiency) and scores are **vendor-stated** or drawn from **community arenas**, "to be treated as claims, not measured facts." The new architecture (**Kimi Delta Attention**, hybrid linear attention; decoding claimed up to **6.3x faster** at 1M tokens) breaks with the K2 cadence (K2 Jul. 2025 → K2.7 Code Jun. 2026, a flagship every two months); two variants accompany the launch (**K3 Max**, **K3 Swarm Max**), with forced sunsetting of the kimi-k2.5/moonshot-v1 series on **August 31, 2026**. **The real weapon is price** (~$3/M input, $0.30 cached, $15 output per secondary sources): a frontier open-weights model at this level **pulls the whole price-performance curve down** — the commoditization of the model layer, accelerated by open source. But the decisive singularity is not a score: it is **reversibility**. A frontier open-weights model turns a consumed API (vendor dependency) into an **option** (self-host, portability, exit from lock-in), at the cost of heavy infrastructure to host 2.8T parameters. SFEIR's view: **open-weights changes the question, not just the answer** — no longer "which model is best/cheapest?" but "how much of my system am I willing to make dependent on a vendor I don't control?". The right posture remains a **routed portfolio** (one model per task, one model per constraint), with Kimi K3 adding a **"reversibility" column** to the decision grid. The "AI Only" conviction stands unchanged: the model is a commodity, the durable advantage lies in the engineering around it (Context Engineering, harness, cost governance, ability to change one's mind). The figures still need validating "on your own" — your repositories, your data.

#Kimi K3#Moonshot AI#Yang Zhilin

SFEIR (voix éditoriale du cabinet)

Economy & Market Auto-verified translation

GPT-5.6 Sol, Terra, Luna : comment OpenAI rebat les cartes du coding agentique et du pricing

SFEIR analysis (firm's voice) of the general availability, on July 9, 2026, of **GPT-5.6** by OpenAI — not a single model but a **family of three tiers**: **Sol** (long-horizon/cyber/science flagship, the only one to unlock the "max" and "ultra" modes), **Terra** (everyday balanced tier, ~half the price of GPT-5.5), and **Luna** (fast/economical, high volume). All three share ~**1.05M tokens** of context, **128k** output tokens, and a knowledge cutoff of **February 16, 2026**. The most structuring fact is not a score but an **aggressive pricing grid** (Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens): Sol keeps the previous flagship's price while being more capable, forcing the comparison onto the **capability-to-cost ratio**. Two billing subtleties (cache writes billed at **1.25×**, a surcharge beyond **272k** tokens) make the grid misleading until one has measured how much context the agent re-reads (read/write ratio ~**153:1** in agentic coding). Engineer's verdict, claimed to be neutral (SFEIR is both a **Google Cloud Premier** partner *and* an **Anthropic** partner): **no one sweeps every table** — GPT-5.6 dominates Terminal-Bench 2.1 and the Coding Agent Index (at a third of the cost per task), Claude stays ahead on SWE-Bench Pro (~15 pts); METR flagged a record **reward hacking** rate on Sol. Conclusion: "stop looking for the champion, learn to route" — the model is a commodity, the durable advantage lies in **Context/Harness Engineering**.

#GPT-5.6#Sol#Terra

SFEIR (voix éditoriale du cabinet)

Transformation & Adoption Auto-verified translation

Comment l'IA agentique bouscule les Grands Groupes ? Partie 2/2 #DevSummit

Podcast interview « À la French » (French-language tech channel, recorded at DevSummit) with Mathieu Grymonprez, Global CDO of the Adeo group (Leroy Merlin, Obramat, Weldom). How a century-old family retail group embraces the agentic AI wave: culture vs structure, accountability, token cost and FinOps, enterprise intelligence lock-in, company memory and agent orchestration. Domain: digital transformation, agentic AI, retail, IT strategy.

#Agentic AI#digital transformation#CDO

Mathieu Grymonprez (Global CDO, groupe Adeo) — invité ; Jean-Baptiste Kempf · Steeve Morin · Mehdi Medjaoui (hôtes du podcast « À la French »)

AI Coding Agents & Skills Auto-verified translation

Lessons from building Claude Code: How we use skills

Blog post from **Anthropic / claude.com** by **Thariq Shihipar** (Member of Technical Staff, Claude Code team), published on **June 3, 2026**, which distills Anthropic's **internal experience** on designing and using **Skills**. **Framing thesis**: a Skill is not a simple markdown file but a **folder** (instructions + scripts + resources + config + hooks) that the agent **discovers and manipulates**; *« You should think of the entire file system as a form of context engineering and progressive disclosure. »* The article makes two structuring contributions. **(A) A taxonomy of 9 skill categories** observed at Anthropic: (1) **Library/API Reference** (docs for internal libs/CLIs with *gotchas* — e.g. `billing-lib`, `internal-platform-cli`, `sandbox-proxy`); (2) **Product Verification** (testing/verification via Playwright or tmux — `signup-flow-driver`, `checkout-verifier`, `tmux-cli-driver`); (3) **Data Fetching & Analysis** (access to data/monitoring stacks — `funnel-query`, `cohort-compare`, `grafana`, `datadog`); (4) **Business Process Automation** (repetitive workflows — `standup-post`, `weekly-recap`, `create-<ticket>-ticket`); (5) **Code Scaffolding** (framework boilerplate — `new-migration`, `create-app`); (6) **Code Quality & Review** (`adversarial-review`, `code-style`, `testing-practices`); (7) **CI/CD & Deployment** (`babysit-pr`, `deploy-<service>`, `cherry-pick-prod`); (8) **Runbooks** (multi-tool diagnostics — `<service>-debugging`, `oncall-runner`, `log-correlator`); (9) **Infrastructure Operations** (maintenance with guardrails — `<resource>-orphans`, `cost-investigation`). **(B) A set of best practices**: don't restate the obvious (*« Claude already knows how to code and can read your codebase »* → target what **contradicts default behavior**); polish the **Gotchas section** (*« the highest-signal content in any skill »*); **progressive disclosure** via the file tree (point to reference files depending on the situation rather than loading everything upfront); **descriptions written for the model** (*« the description field is not a summary, it's a description of when to trigger this skill »*); **setup flows** (config in `config.json`, otherwise prompt via `AskUserQuestion`); **persistent memory** (append-only logs / JSON via the `${CLAUDE_PLUGIN_DATA}` variable); **helper scripts** (*« lets Claude spend its turns on composition… rather than reconstructing boilerplate »*); **hooks conditionnels** (enabled only for the duration of the skill — e.g. a security hook blocking destructive commands). **Distribution at Anthropic**: skills are stored in `./.claude/skills`, informally shared via Slack in a sandbox folder, then promoted via **PR** to the internal **marketplace** once they gain traction; **usage measurement** via a **hook PreToolUse** that logs invocations (revealing popular skills versus underused ones). Direct follow-up to the fiche [[shihipar-claude-code-html-unreasonable-effectiveness-markdown-2026-05-10]] (same author) and a concrete complement to the Skills fiches by Anthropic/Willison/Vincent and to *harness engineering*.

#skills#Claude Code#Anthropic

**Thariq Shihipar** (Member of Technical Staff chez Anthropic, équipe **Claude Code** ; @trq212 / @trq sur X, thariqs.github.io) · pour le blog **claude.com**. Même auteur que la fiche *Using Claude Code: The Unreasonable Effectiveness of HTML* (2026-05-10). Publié le **3 juin 2026**.

AI Coding Agents & Skills Auto-verified translation

The New SDLC With Vibe Coding — From ad-hoc prompting to Agentic Engineering

Google whitepaper (the "Day 1" installment of a series, by Addy Osmani, Shubham Saboo and Sokratis Kartakis) mapping the transformation of the software development lifecycle (SDLC) in the age of coding agents. Thesis: the fundamental shift is not a new language but the move from writing code to **expressing intent**. The document sets out a spectrum ranging from *vibe coding* (prompting and accepting) to *agentic engineering* (AI implements under constraints, tests, and feedback loops designed by humans), with **context engineering** as the central skill, the **software factory** model (the developer's deliverable = the system that produces the code), **harness engineering** (Agent = Model + Harness), and a CapEx/OpEx economic analysis of total cost of ownership.

#new SDLC#vibe coding#agentic engineering

Addy Osmani · Shubham Saboo · Sokratis Kartakis (Google)

Transformation & Adoption Auto-verified translation

The ROI of AI-assisted Software Development

Joint **DORA × delta** report (Google Cloud Professional Services), 60 pages, version **v. 2026.1** (citations February 2026, PDF created April 21, 2026), license **CC BY-NC-SA 4.0** — the first official **DORA ROI** framework dedicated to AI in the SDLC, with an **interactive calculator** at dora.dev/ai/roi/calculator. Pivotal thesis: ***"AI is an amplifier"*** — AI **amplifies** simultaneously the strengths of high-performing organizations and the dysfunctions of struggling organizations; it does not create performance, it **multiplies it where it already exists**. New central concept: the ***J-Curve of AI value realization*** — every AI adoption goes through a **temporary productivity dip** (learning curve + verification tax + pipeline adaptation) before **exponential growth**, a metaphor for the *"tuition cost of transformation"* to be **explicitly budgeted**. Reference calculation: organization of 500 FTE / fully loaded salary $176k / 12.5% time saved per developer (≈ 1h/8h day) → **value $11.6M / investment $8.4M / ROI 39% / payback period 8 months (0.7 year)**. Modeled costs: licenses ($250/user/year), additional API ($80/user/year), training ($9,600/user/year), infra ($100k/year), J-Curve cost ($3.3M for a 15% drop over 3 months). Modeled value: **headcount reinvestment capacity** ($11M — freed capacity to reinvest, **NOT headcount reduction**), revenue from extra feature deployments ($990k, based on a 33% idea success rate, Larsen 2023), **negative downtime impact** (−$344k, "instability tax"). **Explicit reinvestment strategy**: ***"we strongly recommend organizations do not adopt a headcount-reduction strategy"*** — reinvest in innovation, retain talent, capitalize on institutional knowledge. Five pillars of value: Productivity / User Experience / Cost Efficiency / Developer Experience / Business Growth (from most direct to most indirect, *cumulated business value*). Five systemic keys of adoption: **Trust + Platform + Data + Users + Guardrails**. Two-phase roadmap: (1) **Build context layer (CapEx)** — quality IDP + healthy data ecosystems; (2) **Empower human in loop (OpEx)** — context engineering + trust in AI. Indicators: leading = experiment frequency + deployment frequency; stability gauge = change failure rate + rework. Three scenarios to model (Conservative 0.8 value × 1.5 cost / Realistic 1.0 / Optimistic 1.2 × 0.8). External data mobilized: 78% of executives report ROI on ≥1 gen AI use case (Google Cloud), 88% of early agentic AI adopters see positive ROI, **35-40% greenfield productivity vs ≤10% brownfield/legacy** (Stanford), inference cost ÷280 between Nov 2022 and Oct 2024 (Stanford AI Index 2025), **727% ROI over 3 years** for Google Cloud AI customers, average market AI payback of **8 months**. Acknowledged weaknesses: *"all models are wrong"* — the model needs contextualizing, the calculator needs adjusting; risk of double-counting value (time saved → both avoided hire AND extra revenue); a "loose" user experience link, hence excluded from the calculator. **Deontological insight**: ***"We don't measure AI by the code it writes but by the bottlenecks it clears"*** — measured by bottlenecks cleared, not code volume. **Major relevance** for CIOs/CTOs who need to build a defensible AI business case for a CFO/board; for France/Europe, to be articulated with Wescale (realistic X3-X4), Tatsyi/Raiffeisen Bank Ukraine (bank case study, −75 people but deliberate reinvestment), Frizzo (3-5× median), Curran/Intercom (3× R&D over 16 months), DORA Report 2025 (on which this ROI builds).

#DORA ROI of AI-assisted software development#Google Cloud DORA report 2026.1#J-Curve of AI value realization

Rapport conjoint **DORA team × delta team** (Google Cloud Professional Services). Auteurs principaux : **Eva Dong** (AI Value Realization Americas, ex-McKinsey 8 ans, Master Financial Engineering Michigan) · **Andre Ellis Jr.** (Cloud Financial Operations Lead, Morehouse + Wharton MBA) · **Nathen Harvey** (DORA team lead, co-auteur multiples DORA reports + 97 Things Every Cloud Engineer Should Know) · **Vivian Hu** (10X Technology Consultant, contributrice DORA 2025 State of AI-assisted Software Development) · **Ursula Lübbert-Passing PhD** (AI Value Realization EMEA, 20 ans benchmarking + value advisory, PhD effort estimation software projects) · **Eric Maxwell** (lead 10X Technology consulting, ex-Chef Software, contributeur DORA) · **Aaron Wanjala** (cloud developer advocate Spring Boot/Angular). Conseillers et contributeurs : **Ben Jose · Eric Lam · Matt Orr · Allison Park · Ryan J. Salva · Jerome Simms · Dave Stanke · Cedric Yao**. Design : Human After All (humanafterall.studio). Document publié sous licence **CC BY-NC-SA 4.0** · version v. 2026.1 · citations retrieved February 2026.

AI Coding Agents & Skills Auto-verified translation

Improving Frontend Design through Skills

Claude Skills frontend design - Distributional convergence - Context engineering - UI quality improvement - Typography color motion - Anthropic

#Claude Skills#frontend design#distributional convergence

Anthropic (author non spécifié)