Guide signed by **Michael Segner**, published on **August 20, 2026** on the claude.com blog in the *Claude Code* category: a **5-minute** read announced for approximately **31,500 characters** of body text, also offered as a PDF. Stated material: interviews with **more than a dozen** startups, fifteen named — **Artemis Security**, **Cainex**, **Clay**, **ClickHouse**, **Cognition**, **Commure**, **Crosby**, **Emergent**, **Harvey**, **Heidi**, **Higgsfield**, **Omni**, **Parahelp**, **Translucent**, **Zingage**. (A) Five operating rules: *everyone ships*, *automate the tedium*, *trust, but verify*, *build for rebuilding*, *prototype, dogfood, productionize*, each closed with product tips and gathered into a final checklist. (B) A body made of attributed quotes, each rule illustrated by named executives rather than by an aggregated metric. The four figures highlighted are those of the interviewed companies: **+30%** more features shipped (ClickHouse), **2 to 3×** engineering productivity (Omni), **100%** of bug triage automated (Clay), **more than 6,000 PRs per week** (Artemis Security). Two passages depart from the testimonial register: **Cainex**'s self-correction loop on medical coding, described step by step, and the internal use of **Claude Tag** at **Anthropic** as first responder for CI/CD on-call. The question posed at the opening — *"what would it look like if an organization built their product development lifecycle with Claude Code from the ground up?"* — connects with [[claxton-anthropic-ai-native-sdlc-playbook-2026-08-21]], published the next day by the same publisher, and extends [[cherny-wu-reflecting-year-claude-code-2026-07-17]].
#Claude Code#startups#everyone ships
Michael Segner · auteur du guide sur le blog claude.com (fonction non affichée par la page) ; entretiens avec les dirigeants de quinze entreprises nommées.
Official product page from **DeepSeek**, published on **August 13, 2026**, **unsigned**, ~450 words, announcing the *developer preview* release of **DeepSeek Harness** (`dsh`) — a coding-agent harness **open source under the MIT license**, whose repository opened the same day. A three-word thesis, repeated in the title and in the repository description: *« Everything is a plugin »*, paired with a second promise, *« Every run is traceable »*. The page states the equation *« AGENT = MODEL + HARNESS »* and lists the pluggable capabilities — *« models, tools, skills, sessions, sandboxes, storage, loops, scheduling, and the UI »*. Four modes ship: **Standard** (full coding agent), **Code** (tools exposed via the *Code Mode SDK*, letting the model compose multi-step operations inside a TypeScript program), **Minimal** (*« two-tool coding agent with persistent bash and str_replace_editor »*, explicitly *« for benchmarking models in a minimal environment »*), and **Creator** (runtime inspection, in-memory plugin testing). The technical substance sits in the repository, not on the page: `docs/architecture.md` states a logging invariant — *« Model-visible means logged. Anything that reaches a model request must be reconstructable from the log, and a runtime invariant asserts it »* — and states that *« there is no privileged core to patch »*. The technical core is not DeepSeek's own: DSH is built on **Cordis** (the `cordiverse` project, a third party), **vendored** into `vendor/` with a manifest and a sync procedure, and the page places the *« Cordis paper »* at the same navigation level as "GitHub" and "Developer docs". Two LLM adapters ship — `dsh-llm-deepseek` and `dsh-llm-pi-ai`, a generic multi-provider adapter. The repository warns in capitals: *« THERE WILL BE COMPATIBILITY-BREAKING CHANGES »*, and `CLAUDE.md` specifies that `SESSION_FORMAT_VERSION` stays at `0` *« with no compatibility promise »*, with backends rejecting old on-disk formats. Timeline: DSH ships on the day **DeepSeek-V4-Pro reaches GA**, three days before a new API pricing schedule takes effect on **August 16, 2026 at 16:00 UTC**, with peak/off-peak rates and an off-peak discount of **−50%**.
#DeepSeek Harness#dsh#agent harness
**DeepSeek** (DeepSeek AI, laboratoire chinois) · en tant qu'institution. Page produit **non signée** : aucun auteur · aucun ingénieur mis en avant · aucun billet de blog ni papier technique associé. Le « nous » n'apparaît qu'une fois · en dernière phrase — *« We look forward to exploring the limits of intelligence with developers worldwide »*. Publiée le **13 août 2026**. La page est rendue en JavaScript : `curl` sur l'URL renvoie **HTTP 202 avec un corps vide** · le texte n'existant qu'après exécution du bundle. Deux documents de politique sont liés en pied de page — *Safe Use Policy* et *Data Processing Statement*.
Internal research report dated **August 12, 2026** consolidating, for presentation purposes, everything publicly documented about **Buzz** — **Block**'s humans + agents workspace, launched on **July 21, 2026** under the **Apache 2.0** license. It aggregates the two engineering posts already filed alongside the corporate announcement, the GitHub repository, press coverage, X, and **three independent hands-on accounts** that constitute the dossier's only non-self-reported data. **(A) A vocabulary gap documented by quotation**: **Jack Dorsey**'s launch tweet announces *"model-agnostic, decentralized, self-sovereign, and open source"*; Block's `ARCHITECTURE.md` states *"The relay is the single source of truth. All reads and writes flow through it. There is no peer-to-peer event exchange, no gossip, no replication."* The relay is therefore single and authoritative per community: Buzz's "decentralization" is an **organizational sovereignty** — self-hosting and portable identity — not network redundancy. **TFTC**'s formulation: *"Two of those three hold cleanly. The third needs a qualifier."* **(B) An asymmetry between demonstrated rigor and exploitation risk.** On one side, a rare degree of formalism for a v0.4.x/0.5.x: multi-tenant isolation specification **mechanized in TLA+**, authorization properties verified in **Tamarin**, a model-checked Git storage protocol, a hash-chained append-only audit log, 127 *event kinds*, NIP-01/42/98/34. On the other, channel membership is the unit of permission — *"channel membership is not fine-grained tool authorization"* (João Queirós) —, agents run in `--dangerously-skip-permissions` outside any sandbox on a human's machine, and observability is lacking: *"Buzz tells me an agent got a message. It doesn't tell me what happens next"* (DevTools Daily, which reports silent OOM kills). Block acknowledges it: *"the agent can do anything, and security rests entirely on restricting who can tell it what to do"*. **(C) The technical stack**, absent from the filed posts: **Rust** relay (Axum WS + REST), **Postgres**, **Redis**, **S3/MinIO** via Blossom, **Tauri + React** desktop client. Agent integration goes through **`buzz-acp`**, an **ACP** harness that plugs in goose, Codex and Claude Code and translates **ACP ↔ MCP**, plus **`buzz-agent`**, an in-house agent. The report corrects itself on one point: the *"+33% more work"* in Block's TL;DR is the **ratio of completed tasks (20 versus 15 out of 44)**, not a score gain — the score itself rises from 59.1% to 71.5%, i.e. **+12.4 points**.
#Buzz#buzz.xyz#Block
**Deep Research Veille Interne** — rapport non signé · produit le **12 août 2026** en préparation d'une présentation. Aucune URL publique ; source archivée dans `raw-data/`.
Announcement from **Google** on **August 6, 2026**: Google joins as **Core Maintainer** the **Agent Plugins 1.0.0** specification, an open, *vendor-neutral* packaging format for distributing **Agent Skills** and **MCP servers** together. The specification was published by a **TSC** whose Core Maintainers come from **Amazon, Cursor, Microsoft, OpenAI, and Vercel**; Google joins them, represented by **Kevin Hou** (Senior Staff Engineer, Google DeepMind). The two packaged building blocks — Agent Skills and MCP — originate from **Anthropic**, which does not appear on this list of maintainers. **The diagnosis** fits in one sentence: *"The core problem isn't the components. It's the manifest."* A skill is portable, an MCP server is portable; the box they go in is not, and every client had to invent it for itself — hence the forks, the copies of identical components, and their drift. **The format** fits in one constraint: *"A plugin is a directory. That's the whole idea, and the restraint is the point."* A `plugin.json` with two useful lines (`$schema` and `name`), skills in `skills/` in the Agent Skills format, servers declared in `mcp.json` with an **explicit `type` on each entry** (stdio, Streamable HTTP, or the legacy HTTP+SSE) — no more transport guessed from the shape of the config object. The strength of the design lies in what the manifest **cannot** do: neither relocate components nor declare them inline, so there is no discovery path to configure and no precedence order to learn. Operational corollary: components **fail independently** — an `mcp.json` server that fails to start does not take down the plugin's skills, the client skips the entry, keeps going, and reports the failure. The accepted escape hatch is the **reverse-domain** directory (`com.example.client/`), an extension space owned entirely by one client (hooks, agents, commands) that other clients ignore: *"the portable core stays small because the non-portable parts have somewhere legitimate to go."* A section is dedicated to cases where the format is not warranted — *"Not every skill should be a Plugin"*: a single MCP server to a single client, `mcp.json` suffices; a single skill needs no plugin. What v1 explicitly excludes, under *future considerations*: **no installation mechanism, no distribution protocol, no permissions model, no sandboxing requirement, no trust or provenance verification, no UX**. All of this fits into an independently adoptable four-layer stack — **find** (Agentic Resource Discovery), **describe** (AI Catalog, which would register the `application/agent-plugins+json` type), **package** (Agent Plugins), **run** (MCP + Agent Skills). Two Google products already ship: **Agents CLI** and **Data Agent Kit** (BigQuery, Spanner, Cloud SQL).
Announcement from **Meta AI Research** published on **August 5, 2026** (stated reading time: 4 minutes, no individual byline): **Muse Code** in beta, *« a terminal coding agent »*, and the model that powers it, **Muse Spark 1.2**. Meta itself frames the launch: *« This marks our next step toward the frontier, with larger and much more capable models on the way. »* **Three architectural elements on the harness side.** **Asynchronous background agents** that *« remain active throughout each session, rather than being spawned for individual tasks »*, avoiding redundant information gathering and reducing the need for steering. A **local event log** where *« every model call, tool run, approval, and edit is appended »*, making the runtime a system that is *« replay-exact and restart-safe »*, able to resume exactly where it left off after a crash. And **three skills shipped out of the box**: `/plan` (turns a task into a plan submitted for approval), **`/grill`** (stress-tests the plan *« until it holds up »*), and `/goal`. **On the model side**, Meta claims **model-harness co-training** (*« to maximize harness compatibility »*, with harness trajectories sampled via rejection sampling and recipe optimizations for goals, compaction, and sub-agents), **long-horizon** training (whole-repo generation, end-to-end projects, self-research, with planning, goal conditioning, and context compaction), and a **self-improvement loop** where Muse Spark 1.1 generates the environments and instruction templates and then grades candidate solutions, producing a training set for the 1.2. **What the published charts show**, without the text commenting on it: the four comparisons — Terminal-Bench 2.1, DeepSWE 1.1, an internal Meta benchmark, and the GPU kernel optimization case study — place **Muse Spark 1.2 behind Opus 5 in all four cases**, including on Meta's own proprietary benchmark (70.6% versus 79.4%) and on the case study, where the model finishes fourth out of six (+68.7% versus +74.0%). **A reading caution on the version gain**: on the two public benchmarks, 1.1 is measured with `mini-swe-agent` and 1.2 with Muse Code, so the 6.7-point gap conflates model and harness. On the internal benchmark, the only comparison where no harness is mentioned, the 1.1 → 1.2 gap drops to **2.3 points**.
#Meta AI Research#Muse Code#Muse Spark 1.2
**Meta AI Research** — publication institutionnelle sans auteur nommé · sur `research.meta.ai`. Le billet renvoie à un **rapport** pour la méthodologie d'évaluation · non repris ici.
**Skill** entry: **hyperresearch** by **Jordan Gibbs** is a **deep research harness** that turns Claude Code into a documentary research agent, shipped as a PyPI package (MIT, Python 3.11-3.13) installing **20 Claude Code skills**, a CLI, an MCP server, and a local web UI. Observed on **August 3, 2026**: 1,568 stars, 170 forks, repo created on April 9, 2026, last push on August 1. **The core is a 16-step pipeline adaptive by tiers** — `light` (~30-40 min), `full` (~1.5-2.5 h), `dissertation` (4-8 h, 25,000-80,000 words across 300-450 sources) — which takes a prompt and returns an adversarially audited report with full provenance. **The central architecture decision is documented alongside its failure mode**: the entry skill is a **thin router** with no procedure, each step living in its own skill loaded **fresh at the moment it is invoked**, because the previous version was *« one 1200-line skill that got compacted away by the time Layer 4 needed its triple-draft procedure. The orchestrator forgot the procedure, wrote a single draft, and produced a flat-scoring report. »* **Two load-bearing principles.** *« Patch, never regenerate »*: after synthesis, only surgical `Edit` touch-ups are possible, with the patcher and the polish auditor tool-locked to `[Read, Edit]` at the Claude Code allowlist level, so that they *« physically cannot Write a new draft »*. *« Canonical research query is gospel »*: the verbatim prompt is persisted once in `query.md` and re-read by every step and every subagent. **Sixteen subagents** with configurable role and model (fetchers and cite-checker on Sonnet, critics, synthesizer, and patcher on Opus). **The vault** is a persistent markdown store indexed in SQLite — *« Markdown is truth, SQLite is cache »* — with a note lifecycle (`draft → review → evergreen`, `stale → deprecated → archive`), traceable provenance, a composite quality score (source type, citation authority via OpenAlex and Semantic Scholar with retraction flags, internal PageRank), and an **independence audit** that groups syndicated copies together — *« five reprints of one press release argue with the weight of one source »*. **Three mechanical gates before shipping**: citation integrity (every quoted citation must exist **verbatim** in a vault note), a retraction sweep refreshed on every cited DOI, and a citation-to-sentence link check by a skeptical LLM. **Reservation to flag**: the opening claim — *« currently leads the DeepResearch-Bench RACE leaderboard »* — is contradicted by its own footnote, *« forward-looking projection from a stratified pilot… Third party validation is pending »*. A projection is not a ranking, yet the chart places it ahead of Gemini and OpenAI Deep Research.
#skill#deep research#research harness
**Jordan Gibbs** — auteur et mainteneur du dépôt `jordan-gibbs/hyperresearch`. Le projet est distribué sous **licence MIT** et publié sur **PyPI** (`pip install hyperresearch`). Signaux d'adoption au 3 août 2026 : **1 568 étoiles** · **170 forks** · 13 issues ouvertes · dépôt créé le **9 avril 2026** et poussé le **1er août 2026** — soit une traction rapide sur moins de quatre mois. Topics déclarés : `agents` · `agentskills` · `claude-code` · `deep-research` · `deep-research-agent`.
Tech-watch note by **Didier Girard** dated **August 2, 2026**, prompted by a colleague's question ("what is ACP?") to address a problem that is not terminological but **documentary**. **Three protocols compete for the acronym**, with no technical overlap whatsoever: **Agent Client Protocol** (client ↔ agent — Zed, August 2025, JSON-RPC 2.0 over stdio, Apache-2.0, "what LSP did for languages"), **Agentic Commerce Protocol** (agent ↔ merchant — OpenAI + Stripe, Sept. 29, 2025, competing with Google's **UCP** of Jan. 11, 2026 backed by **AP2**), and **Agent Communication Protocol** (agent ↔ agent — IBM Research / BeeAI, marginal but polluting searches). **The core of the note is not the disentangling but its observed failure**: the author searches "ACP" in their tech-watch knowledge base and gets **twelve results, all about the commerce protocol, zero about Zed's** — *"our watch agents had indexed the acronym without disambiguating it"*. Hence a knowledge-engineering rule: ***"a bare acronym is never indexed"*** — the entity is "Agent Client Protocol", "ACP" is **only an alias**, carried by three distinct entities. A structuring clarification follows (**MCP connects an agent to its tools, ACP connects a client to an agent; the two stack**), then the textbook case: **Buzz**, published by **Block** on July 21, 2026 under Apache-2.0 — a self-hostable workspace built on **Nostr**, where every human or agent participant is a **key pair** and every message, workflow step, or git push is a **signed event** in an append-only log. An entirely protocol-based architecture (`buzz-acp` an ACP harness over stdio, `buzz-agent` an ACP agent calling an LLM, `buzz-dev-mcp` an MCP shell + editing server), hence agent agnosticism: **Goose, Claude Code, and Codex** plug in through the same harness, and **Hermes** (Nous Research) connected to it without Block writing a single line — *"N+M instead of N×M, running in production"*. The note closes on the question of the **Claude subscription** versus third-party agents, with a five-stage 2026 timeline and a **design rule** that holds beyond this case: the line is not legal but **architectural** — ***"who is consuming, and on whose behalf"*** (an `owner-only` agent consumes your subscription on your behalf; an `anyone` agent in a shared channel routes your colleagues' requests through your account). **Verification carried out on this corpus**: the thesis holds, and more starkly than the note claims — not only is "Agent Client Protocol" **completely absent**, but the bare acronym `ACP` **is already typed as an entity** in two fiches, and the KB page `Agentic-Commerce-Protocol` **already attributes the protocol to Google** when it belongs to OpenAI + Stripe. The collision described is not a future risk: it has **already produced an attribution error** in the graph.
**Didier Girard** — auteur de la note. Écrit ici depuis la position de **praticien de la veille outillée** : le déclencheur est une question de collègue · le matériau principal est le comportement observé de sa propre base de connaissances · et la conclusion est une **règle de curation** adoptée en interne. Le texte alterne donc deux voix — l'explicateur de protocoles et l'ingénieur de la connaissance qui constate un défaut chez lui et en tire une norme.
**Block** announcement from **July 21, 2026**, signed by **Tyler Longwell**: **Buzz**, an *open source* and **self-hostable** channel-driven workspace where humans and agents share the same room — chat, search, automation, and **Git hosting** on a single server, built on **Nostr**, an open protocol for signed messages and portable identities. Opening thesis: *« Models can do the work now. Teams still need somewhere to do it together. The bottleneck moved from intelligence to coordination. »* Three engineering pieces. **(A) Agent identity.** The starting point is a refusal — to stop lending one's credentials to a bot: *« We have been letting bots play dress-up as us. It's weird. It's dangerous. »* Each agent gets **its own key**, its owner signs a **narrowly scoped authorization**, and the agent then signs its work with its own identity. The delegation cryptography is conventional; the design decision is less so: *« authorization does not erase authorship »* — the agent remains the author, its *credential* proving who authorized it and under what conditions. Immediate consequences: a leaked agent key is revoked without touching the human identity, and withdrawing the owner prevents the agent from reconnecting, with its active sessions needing to be terminated separately. **(B) Git on object storage.** The observation: *« In the past, Git has always had a convenient rate limiter: humans »* — a group of agents produces months of person-commits and CI in a single afternoon, with many simultaneous writers, on forges sized for human fingers. Buzz stores repositories as **immutable, content-addressed packfiles** plus a **single mutable manifest pointer**; a *push* writes the objects first, then advances the pointer via **conditional compare-and-swap**, that swap being the commit point — workspace events announce the change, they do not define it. The protocol is **specified in TLA+ and model-checked** (durability, reconstruction, concurrent pushes), with the bounded result depending on three explicit object-store guarantees, hence a **conformance suite** every backend must pass. **(C) Interoperability and privacy.** Claude Code, Codex, goose *« and any agent speaking Agent Client Protocol »* work inside Buzz; switching model or harness leaves the project's identity, permissions, and history intact. Telemetry and cancellation travel as ephemeral encrypted messages, memory and cost accounting as durable encrypted messages — *« the server sees routing metadata, not those payloads »*. Memory argument: *« A conventional forge preserves the diff and a green check. Buzz also preserves why the obvious fix was wrong. »* Anti-lock-in argument: if Buzz disappears, the identity and signed history remain verifiable, Git stays Git.
#Buzz#Block#agentic workspace
**Tyler Longwell** — *« Building multi-player AI at Block »* · auteur unique et signataire à la première personne. Publié le **21 juillet 2026** sur le blog Block Engineering.
Boris Cherny (Head of Claude Code) and Cat Wu (Head of Product, Claude Code) publish a short LinkedIn video, "Reflecting on a year of Claude Code," in which they put forward a thesis: **product and engineering roles are merging**. At Anthropic, the product team, devrel, and design **all write code**; many engineers **ship products end to end** (idea → build → legal/marketing/security → release into the world). Their conclusion: AI benefits profiles with **curiosity**, **product taste**, and a taste for **end-to-end ownership**. The note mainly captures the **comment-thread discussion** (55 comments, 28 substantive): a consensus that **reframes** the thesis — it is not roles disappearing, it is that **shipping becomes cheap**, which shifts value toward judgment and defining the right problem — set against a lucid minority on the flip side (accountability, governance, IP).
#Boris Cherny#Cat Wu#Claude Code
Boris Cherny (Head of Claude Code, Anthropic) et Cat Wu (Head of Product, Claude Code, Anthropic) — vidéo ~47 s publiée par Claude for Business sur LinkedIn · repartagée par Claude. Commentateurs cités : Omer K. · Syed T. · Andrei K. van Noordt · Kristóf Nagy · Natasha Egan · Natasha Newbold · Rehan Nazir · Noman A. · Kevin Schoovaerts · Sunny Vara · Paul Breuler · Ron H. · Mohammadjavad Sayadi · Chris Bounds · Mohamed Anis · Panny Malialis · David H. · plebs.me · James Hutchinson · Dewayne J Grunden II · e.a. (28 commentaires de fond retenus sur 55).
**Boris Cherny** (Creator & Head of Claude Code @Anthropic) publishes a framework table on LinkedIn, **« Steps of AI Adoption »**, mapping an engineering team's adoption of agentic AI across **5 stages (0→4)**, each characterized by an **order of magnitude of agents driven** and a **transformation of the engineer's role**: **0 Gated** (0 agents, locked-down access), **1 Assisted** (~1 agent — "you + one agent", supervised pair programming), **2 Parallel** (~10 agents — **orchestrator**), **3 Supervised autonomy** (~100 agents — **manager of managers**, an org tree), **4 AI-native** (~1,000+ agents — **VP steering by intent**). The table crosses five columns: number of agents, *what it looks like*, *the bottleneck*, *the products that help*, *the guardrails*. **Central thesis**: consuming more tokens does not move you up a level — advancing to the next stage requires **identifying and breaking the next bottleneck** AND **building the next set of guardrails**. Concretely: giving Claude a trustworthy **self-verification loop** (tests + build + lint + e2e on a real environment), enabling **Auto mode** (avoiding blocking permission prompts), making **code review and security review the default**, adopting multi-agent interfaces (Agent view CLI, Desktop, iOS/Android apps, Tag), then `/loop`, `/batch`, `/goal`, **dynamic workflows** and **worktree isolation** for subagents. On steering: usage (dashboard) measures **activity, not return**; the right question is *"would we have spent engineering effort on this anyway? if so, how many manual engineer-hours would it have cost?"* — that's the ROI. The real payoff arrives when **fixing and maintaining happens in the background** and teams focus on *building*. Anthropic sits at **stage 3, heading toward 4**; Boris Cherny states he has personally reached **level 4**.
#Boris Cherny#Claude Code#Anthropic
Boris Cherny (Creator & Head of Claude Code @Anthropic)
SFEIR analysis (firm's voice) of the general availability, on July 9, 2026, of **GPT-5.6** by OpenAI — not a single model but a **family of three tiers**: **Sol** (long-horizon/cyber/science flagship, the only one to unlock the "max" and "ultra" modes), **Terra** (everyday balanced tier, ~half the price of GPT-5.5), and **Luna** (fast/economical, high volume). All three share ~**1.05M tokens** of context, **128k** output tokens, and a knowledge cutoff of **February 16, 2026**. The most structuring fact is not a score but an **aggressive pricing grid** (Sol $5/$30, Terra $2.50/$15, Luna $1/$6 per million tokens): Sol keeps the previous flagship's price while being more capable, forcing the comparison onto the **capability-to-cost ratio**. Two billing subtleties (cache writes billed at **1.25×**, a surcharge beyond **272k** tokens) make the grid misleading until one has measured how much context the agent re-reads (read/write ratio ~**153:1** in agentic coding). Engineer's verdict, claimed to be neutral (SFEIR is both a **Google Cloud Premier** partner *and* an **Anthropic** partner): **no one sweeps every table** — GPT-5.6 dominates Terminal-Bench 2.1 and the Coding Agent Index (at a third of the cost per task), Claude stays ahead on SWE-Bench Pro (~15 pts); METR flagged a record **reward hacking** rate on Sol. Conclusion: "stop looking for the champion, learn to route" — the model is a commodity, the durable advantage lies in **Context/Harness Engineering**.
First-rate technical account by **Jarred Sumner**, creator of **Bun** (JS/TS runtime, >22M downloads/month), on the **complete rewrite of Bun from Zig to Rust in 11 days** (May 3→14, 2026) driven by **Claude** — an exceptional case study in AI-assisted software engineering **at industrial scale**. Motivation: a recurring class of bugs (use-after-free, double-free, leaks) arising from the mix of GC-managed memory (JavaScriptCore) and manual memory (Zig); in **safe Rust**, these bugs become **compile errors** with automatic cleanup (`Drop`/RAII) — "a better feedback loop than a style guide." Rejecting the dogma that "a rewrite is always a bad idea" (a year of bugfix freeze for 3 engineers), Sumner chooses a **mechanical port** (preserve the architecture, minimal behavior change) validated by the **existing test suite, written in TypeScript and therefore language-independent** (60,624 tests, 1.39M `expect()` assertions, 0 tests removed, 6 platforms). The harness: **~50 dynamic workflows** in **Claude Code**, *write → 2+ adversarial reviewers → apply* loops, up to **64 Claude instances in parallel** (4 worktrees × 16), with **PORTING.md** + **LIFETIMES.tsv** generated in preparation. Numbers: **6,502 commits** (peak 695/h, 58/min, ~1,300 lines/min), final diff **+1,009,272 lines**, ~16,000 compile errors treated as a queue, **5.9B uncached input tokens + 690M output ≈ $165,000**. Key methodological levers: **adversarial review** (a second Claude, separate context, sees only the diff, tasked with finding why it's wrong — catches subtle bugs that are *semantically* different but *syntactically* identical) and the principle **"fix the process that generates the code, not the code by hand."** Model used: a pre-release of **Claude Fable 5** (Mythos class). Since the merge: **11 rounds of Claude Code security review**, 24/7 coverage-guided fuzzing (100B executions → ~15 PRs), **4% `unsafe` code** (78% on a single line), **19** known regressions fixed. In production: Claude Code v2.1.181, the first release on Bun-in-Rust, **+10% faster startup on Linux**. Disclosed upfront: **Bun was acquired by Anthropic in December 2025**.
#Bun#Jarred Sumner#Zig-to-Rust rewrite
Jarred Sumner (créateur de Bun ; travaille chez Anthropic depuis le rachat de Bun en décembre 2025)
Short note from Simon Willison (weblog) relaying two tips heard during a *Fireside Chat* at AIE with Cat Wu and Thariq Shihipar (Claude Code team): **let the model (Fable, and to some extent Opus) exercise its own judgment rather than dictating rules to it** — illustrated with the decision of whether to write tests. Second tip, from Jesse Vincent: to **save precious Fable tokens** (ahead of an imminent price increase), ask Fable to **delegate small tasks to less powerful models**, letting it judge which one. Willison shows the exact prompt used (« *use your judgement to decide an appropriate lower power model and run that in a subagent* ») and the **memory file** that Claude Code wrote in response. Domain: prompt engineering, coding agents, token economics, multi-model orchestration.
#Model judgment#delegation to subagents#model override
Long-form essay by **Shubham Saboo** (X/Twitter) advancing a thesis on the Product Manager role in the age of agents: the next key skill is **not prompt engineering** but **Loop Engineering** — designing a *system that improves with every run* rather than writing the perfect prompt every time. A **loop** is a repeated cycle: change what shapes the agent's behavior → run it → evaluate the output → keep the change if quality rises, revert otherwise → **compound the learning** so the next version starts ahead. For a PM, the entry point is not code but the **durable artifacts** that encode their judgment: PRD-review skill, customer-call *summarizer*, evaluation rubric, launch checklist, research workflow, `CLAUDE.md`, prompt template, prioritization framework. Because they are reused, these artifacts **compound in both directions** — and **drift** silently (a CLAUDE.md that keeps growing, a checklist that gets ignored…): the model has not regressed, the artifacts have drifted unwatched. A loop has **5 parts**: trigger, action, **proof**, memory, **stop condition** (the most critical). **Evals** become PM work (testing the artifact against known examples: 3 good / 3 bad PRDs, 5 understood calls, 2 past launches). **Memory** lives on **GitHub** (the repo becomes "product memory": commits, diffs, eval results, decision log, rollback). Recommended first loop: a **weekly product signal loop** (every Friday). Taste remains central — but it now needs **proof**. Cites Boris (creator of Claude Code): "he no longer writes prompts, he writes loops."
Podcast interview « À la French » (French-language tech channel, recorded at DevSummit) with Mathieu Grymonprez, Global CDO of the Adeo group (Leroy Merlin, Obramat, Weldom). How a century-old family retail group embraces the agentic AI wave: culture vs structure, accountability, token cost and FinOps, enterprise intelligence lock-in, company memory and agent orchestration. Domain: digital transformation, agentic AI, retail, IT strategy.
#Agentic AI#digital transformation#CDO
Mathieu Grymonprez (Global CDO, groupe Adeo) — invité ; Jean-Baptiste Kempf · Steeve Morin · Mehdi Medjaoui (hôtes du podcast « À la French »)
Case study published by the **Cornell AI Innovation Hub** (June 15, 2026): how a two-semester collaboration between the AI Hub, graduate students, and Cornell's Treasury team turned a time-consuming manual investigation into an AI tool that **recovered $100,000** in unidentified payments on a first batch. A successful **AI4Business** use case (financial process) that illustrates the **Leader-Lab-Crowd** framework of **Ethan Mollick** almost point by point: the **AI Hub** plays the role of the **Lab** (a central, ambidextrous team of technologists plus students); **Treasury** (Cheryl Barnes, Marie Graves…) is the **Crowd** carrying business knowledge and the real pain point; and the **$100,000** constitutes the **visible reward** (vivid win) that anchors adoption — exactly the incentive lever Mollick considers decisive. Key method: **"context first, then plan, then build"** via **Claude Code Plan Mode**, a chain of **fuzzy matching → Gemini Enterprise Web Search → Claude synthesis**, all within the governed **Cornell AI Gateway**. *"The $100,000 is a start."*
#Cornell AI Innovation Hub#unidentified payments#payment reconciliation
**Pete Stergion** — Desktop Engineer au Cornell AI Innovation Hub · co-tech lead du projet (avec Phil Williammee). Article institutionnel signé de l'AI Hub.
Polemical essay-thread by Ahmad Osman (@TheAhmadOsman) on X, *"Anthropic's War on Opensource AI"* (1.7M views). Core thesis: Anthropic systematically converts "safety" into a **control mechanism** (permission regime, regulatory capture, anti-competitive access restrictions, behavioral opacity) to keep builders, startups, and open source communities **downstream** of a handful of frontier labs. Central anchor point: the **Fable incident** (silent degradation of competing AI dev requests). Advocacy for open source / local AI as the only viable "political economy of intelligence." Domain: AI policy, open source vs. closed labs, sovereignty, governance.
In-depth technical guide (Lushbinary agency blog) on **Loop Engineering**: designing the systems that drive coding agents in a loop, rather than prompting them manually. Covers the lineage prompt → context → loop engineering, the Ralph technique (Geoffrey Huntley), the **five building blocks + memory** of a loop, their implementation in Claude Code and OpenAI Codex, writing verifiable stop conditions, an adoption maturity scale, and the risks that worsen as loops grow more sophisticated. Domain: agentic software engineering, coding agents, harness/orchestration.
Sunday tinkering post by **Mark Dembo** (Head of Solutions, Developer Platform & AI at **Cloudflare**) published on **June 7, 2026** on his personal blog. **Narrative**: inspired by **Steve Ruiz**, the author buys a small **M5Stack Stick 3** device (~€30) and, taking advantage of the release of **Opus 4.8**, builds himself a **DIY AI agent** "out of pure curiosity, with no goal." **Iteration 1 (45 min)**: he throws the device's documentation at **Claude Code**, which generates Python scripts (~200 LOC, *"zero blast radius"*) displaying the weather in Munich, then several cities; a **Cloudflare Workers + Workers AI backend** adds **text-to-speech (TTS)**, **push-to-talk** (speech-to-text), and a central **small LLM** to answer questions. **Iteration 2 (a real agent)**: switching REST endpoints to **WebSocket** transport via the **Cloudflare Agents SDK** + **Dynamic Worker execution** → the ***"Code Mode"*** pattern (the agent writes and executes code to accomplish its task). The agent then answers public-data questions (11! = factorial, the Champions League winner via `fetch()` on Wikipedia, the weather in any city). **Iteration 3 (real powers)**: connecting to **Todoist** via an **MCP OAuth** flow → 50 tools at once, hence two problems: **context bloat** and **real damage risk**. The fix draws on Cloudflare's **MCP Server Portal** + Claude connector settings: per tool, **Always allow / Ask for approval / Disable** (*Disabled* tools never enter the context; an **LLM classifier** accepts only distinct "allow" grants and **defaults to deny**). **Stated posture**: reducing his role to ***"idea generator, executor and judge"*** (and rarely technical guide), a "human-in-the-loop" flow he considers not very *"2026"* (copy-pasting into UIFlow). **What he did NOT do**: no latency/streaming optimization, no optimistic LLM calls, no evals, ***"I did not even look at the code once."*** **Wonder**: €30 + one Anthropic session window + a few cents of Cloudflare inference → an object that listens and speaks, driven in natural language; *"the true unlock is how accessible it is."* Sharp contrast with [[thomas-pragdave-failing-faster-code-rot-ai-velocity-2026-06-06]] (here *"zero blast radius"* justifies never looking at the code); concretely illustrates *Code Mode* / *"the agent just writing and executing code,"* the **MCP** pattern ([[claude-skills-bigger-than-mcp-willison-2025-10-16]]), *Ask for approval*-style tool governance (uber-engineering-agent-identity-crisis-zero-trust-spire-2026-05-21), and the *systems around the model* doctrine from dropbox-okumura-beyond-code-generation-engineering-productivity-ai-agents-2026-05-28.
#BYO agent#bring your own AI#tinkering
**Mark Dembo** (@darkmembo / @mdembo) · **Head of Solutions – Developer Platform & AI** chez **Cloudflare** (auparavant auteur sur le blog Cloudflare). Billet personnel publié sur son blog *markpauldembo.com* le **7 juin 2026** (description : *« Thoughts about tinkering on a Sunday »*).
Engineering write-up from Anthropic's **Data Science & Data Engineering** team (Chen Chang, Clement Peng, Justin Leder, Johanne Jiao, Josh Cherry) published on **June 3, 2026** on the Anthropic blog (*Enterprise AI* category, focus on **Claude Code**). **Headline result**: ***"95% of business analytics queries are automated by Claude, with ~95% accuracy in aggregate"*** (up to **~99%** in certain domains). **Core problem**: analytics is **not** code — *"there's often only a single correct answer using a single correct source"* — it requires **mapping a user question to precise, up-to-date entities** in the data model. Three **failure modes**: (1) **concept↔entity ambiguity** (e.g. *"active users"*: which actions? exclude fraudsters? which window?); (2) **staleness** (assets and the agent's knowledge become *"subtly wrong"*); (3) **retrieval failure** (*"80% of failed queries had the information present in the corpus"* but unfindable). **Solution = a 4-layer "agentic analytics stack"**: (L1) **Data foundations** — dimensional modeling, **canonical datasets** *"single source-of-truth"*, metadata *"as a first-class product"*, integrity via CI/CD; (L2) **Sources of truth** in decreasing order of trust — **semantic layer** (the agent is *"structurally required (by skill instruction) to leverage the semantic layer first"*), lineage graph, **query corpus** (distilled into structured docs, **not** raw retrieval), business context (knowledge graph: roadmaps, decision logs, org); (L3) **Skills** — the decisive lever: ***"without skills … didn't exceed 21% … Adding skills gets these numbers consistently above 95%"***; structured **in pairs** (*Knowledge skill* = router to ~30 reference files; *Unbook skill* = senior analyst workflow: clarify → find sources → execute → **adversarial review**); **colocated** maintenance (*"a code-review hook flags any reporting-model change that doesn't touch a skill file"* → **~90% of data PRs include a skill change**); (L4) **Validation** — offline evals (threshold ~90% to launch an agent, target ~100%), **ablation testing** (notable negative result: raw grep across thousands of SQL files → accuracy moves *"less than a point"*), online (adversarial review: **+6% accuracy, +32% tokens, +72% latency**), **provenance footers** (source tier + freshness + ownership), **active correction harvesting** (scheduled agents scanning channels to draft markdown fixes). **Strategic insight**: *"documentation generated, definitions owned by humans"* — letting the LLM **define** metrics was *"net-negative"*. **Minimal starting point**: a handful of canonical datasets + a few dozen evals + a *thin knowledge skill* capture *"most of the upside"*. Strongly converges with [[shihipar-claude-code-lessons-building-skills-2026-06-03]] (skills = folders, Gotchas, hooks), the *systems around the model* doctrine of [[dropbox-okumura-beyond-code-generation-engineering-productivity-ai-agents-2026-05-28]], the **semantic layer / ontology** of talisman-modern-data-101-ontology-pipeline-refresh-2026-05-04 and seale-semantic-agent-model-harness-ontology-data-2026-04-17, the *context development lifecycle* of debois-tessl-context-development-lifecycle-ai-coding-agents-2026-02-19, and the UDA/knowledge graph of netflix-uda-unified-data-architecture-knowledge-graph-2025-06-12.
#self-service analytics#agentic data analytics#Claude Code
**Chen Chang · Clement Peng · Justin Leder · Johanne Jiao · Josh Cherry** — équipe **Data Science & Data Engineering d'Anthropic**. Article publié le **3 juin 2026** sur le blog Anthropic (claude.com/blog) · catégorie *Enterprise AI* · ~5 min de lecture.
Blog post from **Anthropic / claude.com** by **Thariq Shihipar** (Member of Technical Staff, Claude Code team), published on **June 3, 2026**, which distills Anthropic's **internal experience** on designing and using **Skills**. **Framing thesis**: a Skill is not a simple markdown file but a **folder** (instructions + scripts + resources + config + hooks) that the agent **discovers and manipulates**; *« You should think of the entire file system as a form of context engineering and progressive disclosure. »* The article makes two structuring contributions. **(A) A taxonomy of 9 skill categories** observed at Anthropic: (1) **Library/API Reference** (docs for internal libs/CLIs with *gotchas* — e.g. `billing-lib`, `internal-platform-cli`, `sandbox-proxy`); (2) **Product Verification** (testing/verification via Playwright or tmux — `signup-flow-driver`, `checkout-verifier`, `tmux-cli-driver`); (3) **Data Fetching & Analysis** (access to data/monitoring stacks — `funnel-query`, `cohort-compare`, `grafana`, `datadog`); (4) **Business Process Automation** (repetitive workflows — `standup-post`, `weekly-recap`, `create-<ticket>-ticket`); (5) **Code Scaffolding** (framework boilerplate — `new-migration`, `create-app`); (6) **Code Quality & Review** (`adversarial-review`, `code-style`, `testing-practices`); (7) **CI/CD & Deployment** (`babysit-pr`, `deploy-<service>`, `cherry-pick-prod`); (8) **Runbooks** (multi-tool diagnostics — `<service>-debugging`, `oncall-runner`, `log-correlator`); (9) **Infrastructure Operations** (maintenance with guardrails — `<resource>-orphans`, `cost-investigation`). **(B) A set of best practices**: don't restate the obvious (*« Claude already knows how to code and can read your codebase »* → target what **contradicts default behavior**); polish the **Gotchas section** (*« the highest-signal content in any skill »*); **progressive disclosure** via the file tree (point to reference files depending on the situation rather than loading everything upfront); **descriptions written for the model** (*« the description field is not a summary, it's a description of when to trigger this skill »*); **setup flows** (config in `config.json`, otherwise prompt via `AskUserQuestion`); **persistent memory** (append-only logs / JSON via the `${CLAUDE_PLUGIN_DATA}` variable); **helper scripts** (*« lets Claude spend its turns on composition… rather than reconstructing boilerplate »*); **hooks conditionnels** (enabled only for the duration of the skill — e.g. a security hook blocking destructive commands). **Distribution at Anthropic**: skills are stored in `./.claude/skills`, informally shared via Slack in a sandbox folder, then promoted via **PR** to the internal **marketplace** once they gain traction; **usage measurement** via a **hook PreToolUse** that logs invocations (revealing popular skills versus underused ones). Direct follow-up to the fiche [[shihipar-claude-code-html-unreasonable-effectiveness-markdown-2026-05-10]] (same author) and a concrete complement to the Skills fiches by Anthropic/Willison/Vincent and to *harness engineering*.
#skills#Claude Code#Anthropic
**Thariq Shihipar** (Member of Technical Staff chez Anthropic, équipe **Claude Code** ; @trq212 / @trq sur X, thariqs.github.io) · pour le blog **claude.com**. Même auteur que la fiche *Using Claude Code: The Unreasonable Effectiveness of HTML* (2026-05-10). Publié le **3 juin 2026**.
Blog post by **Pasquale Pillitteri** (software engineer, Palermo) published on **May 29, 2026** (FR version), 18-minute read, *Claude Code & Anthropic* section. **Pivot thesis**: *« Claude Opus 4.8 is the most powerful SEO model of 2026, but almost everyone uses it wrong »* — not a model problem but a **system** problem. The golden rule: ***« strategy is a whiteboard, production is an assembly line »*** — SEO must be **split into two distinct phases**, and mixing them is *« the fastest way to waste a model that costs five dollars per million input tokens and twenty-five for output »*. **Model context**: Opus 4.8 released on **May 28, 2026** (41 days after Opus 4.7), **1M-token** context, **GraphWalks Long-Context F1 at 1M: 40.3% → 68.1%**, **SWE-bench Verified 88.6%**, **USAMO 2026 96.7%** (+27.4 pts), **HLE with tool 57.9%**, unchanged price **$5/$25** per M tokens, **Fast Mode 2.5× at $10/$50**, four **effort levels** (Low, High, Extra, Max). **The central anti-pattern** = *« the giant conversation »* / **context drift**: mixing strategy, keyword research, competitive analysis and writing in a single chat produces a *« mush of contradictory intentions »* → the model slides toward **generic best practices** ("holistic optimization", "strategic approach") instead of data-anchored content. **Phase 1 — Strategy (whiteboard, visual UI, one-off)**: dashboard / Google Sheet / Claude.ai canvas to decide while looking at the data together. **3 plays**: (a) **classified keyword research** (table of volume / difficulty 0-100 / intent / business potential / priority = volume÷difficulty×business weight); (b) **visual competitive analysis** (topic-coverage matrix, gaps); (c) **phased roadmap** (quick wins M1-2 / medium term M3-6 / pillar pages M7-12). **Extra/Max** mode is justified here (*« one right strategic decision is worth a thousand well-written pages on the wrong keywords »*). 3 closed artifacts saved to Notion/Drive. **Phase 2 — Production (assembly line, Opus 4.8 + MCP)**: the model shifts from strategist to **execution machine**; every decision **anchored to live data** via the **Model Context Protocol**. **Minimum MCP stack**: **GSC MCP** (AminForou/mcp-gsc, 500+ stars), **official Ahrefs MCP** (98 stars), **GA4 MCP**; the `modelcontextprotocol/servers` repo = **86,440 stars**, **10,000+ active servers**, 97M SDK downloads/month. Setup ~35 min, monthly refresh ~20 min. **Weekly loop**: a single prompt pulls live data, builds the brief (top 10 SERP + GSC + Ahrefs), derives H2/H3, writes, checks density, suggests titles → **+45% productivity**, draft in **6-12 min** (explicit reference to **Ryan Law / Ahrefs content engineering**, 23 skills). Mentions Anthropic's **Dynamic Workflows** (up to 1,000 subagents). **4 common mistakes**: (1) not checking the numbers (spot-check mandatory, *trust & verify*); (2) fully replacing Semrush/Ahrefs (MCP is a **layer on top**, not a substitute); (3) ignoring the **paid-organic content gap** (education client case: **2,742 wasted terms / 351 opportunities** identified in 90 s); (4) using Opus 4.8 where **Haiku 4.5** suffices (meta descriptions, alt text). **Cost**: $1-3 per 2,500-word article. **Sonnet 4.6** suffices for recurring production, Opus 4.8 reserved for strategy. SEO-optimized and self-referential article (the author writes about SEO in content itself designed to rank for "Opus 4.8 SEO"). Direct convergence with **Ryan Law/Ahrefs** (cited), **systems around the model** (Dropbox/Okumura), **skills-over-prompts** (Lattice), Haiku/Sonnet/Opus model routing (Gupta token-to-outcome).
#Claude Opus 4.8#AI SEO#two-phase workflow
**Pasquale Pillitteri** — Ingénieur informatique / développeur logiciel basé à **Palerme** (Italie) · certifié Innovation Manager UNI 11814:2021. Auteur d'un blog tech actif (rubrique *Claude Code & Anthropic*) · avec une newsletter hebdomadaire (~3,4k lecteurs). Article publié en version **FR** le **29 mai 2026** (lendemain de la sortie d'Opus 4.8).
Official **Salesforce News** blog post (*Agentic Enterprise* section, *"Pioneering the Agentic Shift Within Salesforce Engineering"* series), published on **May 27, 2026** (6-minute read) by **Srinivas "Srini" Tallapragada**, *President and Chief Engineering and Customer Success Officer* at Salesforce. Direct follow-up to an earlier post (*"How we got our engineers to use AI — without breaking everything"*) which recounted crossing **>90% adoption**. **Pivot thesis**: Salesforce Engineering moved from a world where AI was a useful *copilot* to one where **agentic tools drive the software development lifecycle (SDLC) itself** — writing code, reviewing PRs, generating tests, updating documentation, managing deployments, coordinating work once handled through human handoffs. **Canonical signal decision**: org-wide standardization on **Claude Code** + ***"we removed all token limits"*** — *"remove every last piece of friction between our engineers and the tools that make them faster and more effective"*. **Major empirical result** (April 2026 vs April 2025): work items completed per developer **+50.8%**, PRs merged per developer **+79%**, and above all **Effective Output score** (an ML measure of the **real value of delivered code**, not volume) **+151.3% year over year**. **Flagship use case**: migration of **33 API endpoints** to a cloud-native architecture, estimated at **~231 person-days** (7 per API) the traditional way, completed in **13 days — 18× faster** — via a **rule-based framework built in Claude** (markdown files + reference implementations), with PR feedback continuously fed back into the rule set, **autonomous LLM loops (build, fix, validate)** with no manual intervention, parallelized across isolated environments → **5 PRs**, the largest delivering **21 endpoints with 100% test coverage**. **No speed↔quality tradeoff**: through the **Engineering 360** platform (centralizing engineering data from hundreds of systems), **total incidents drop by 5%** despite the rise in PRs (*"quality doesn't suffer from speed. It benefits from it"*), thanks to **security guardrails and quality standards structurally embedded** in the agentic workflow (Trust as the #1 value). **SDLC overhaul**: once AI is adopted, engineers **tear down and rebuild** workflows (which processes to eliminate? which handoffs are now unnecessary? where does a human still do work an agent could own?). **New engineering craft**: **Claude Code skills** (packaged, reusable capabilities encoding team context, naming conventions, patterns) become a shared, composable **engineering artifact**; **AI Expert Suite** + **Salesforce Foundation Plugins** = an institutionalized, curated skills library (internal benchmark: **higher accuracy and reliability, reduced unnecessary cost**); **subagents & agent teams** parallelize workstreams (*"They describe the outcome, and a set of coordinated agents figures out the steps"*). **What remains hard**: (1) **context management** in long sessions — **CLAUDE.md file quality** varies widely and weighs heavily on output quality; (2) **agentic security** = a fundamentally different model (agents that *act*, not just *suggest* → increased blast radius); (3) **evolving roles** (how do juniors become seniors if AI absorbs entry-level work? role of the designer/PM? the execution unit = scrum team → experiments with 1- or 3-person units). Conclusion: *"It changed what was economically possible"*; the stated ambition is **"the most automated, agentic SDLC in the industry"**. Directly intersects with Gupta (*cost of a completed outcome*, marginal token utility), Greenwald/Sierra (outcome-based pricing), DORA (ROI / cost per feature) and the BFM/Girard debate (token as a value fuel, not a cost to cut).
#Agentic SDLC#agentic SDLC#Claude Code
**Srinivas « Srini » Tallapragada** — *President and Chief Engineering and Customer Success Officer* de **Salesforce**. Plus d'une décennie chez Salesforce · dirige l'ingénierie mondiale de la plateforme unifiée. Auteur de la série *Agentic Enterprise* sur le blog Salesforce News ; ce billet (27 mai 2026) est la **suite** d'un premier opus consacré à l'adoption de l'IA par les milliers d'ingénieurs Salesforce (*« How we got our engineers to use AI — without breaking everything »*). Position d'autorité = **dirigeant exécutif** parlant en son nom et au nom d'une organisation d'ingénierie à grande échelle (donnée terrain à l'échelle d'un hyperscaler SaaS) · avec accès aux métriques internes (Engineering 360, Effective Output).
Manifesto-style article by **Thariq Shihipar** (Engineer & serial entrepreneur, Claude Code team at Anthropic) announcing a **change in the default output format for agents**: replacing **Markdown with HTML**. Thesis: Markdown has been the dominant format between humans and agents (simple, portable, editable, readable) but has become **a bottleneck** as agents produce longer and richer artifacts (specs, plans, reports, code review). Beyond ~100 lines, no one reads a Markdown file anymore. HTML solves six limitations simultaneously: **information density** (tables, CSS, SVG, scripts, canvas, images), **visual clarity** (navigable, mobile-responsive layout), **ease of sharing** (an S3 link directly openable in a browser), **two-way interactivity** (sliders, knobs, "copy as JSON/prompt" buttons to loop back into Claude Code), **native contextual ingestion** (Claude Code reads the codebase + MCP Slack/Linear + git history + Chrome) and **enjoyment** (the author explicitly claims *"it's joyful"*). Five canonical uses detailed: (1) **specs/plans/exploration** in a comparative grid, (2) **PR review** with inline annotated diff, (3) **design & prototypes** with animation sliders, (4) **reports/research/learning** (the author had a prompt-caching explainer generated from git history), (5) **custom throwaway editors** (drag-and-drop of Linear tickets, feature-flag editors, side-by-side prompt-tuner) that produce a re-injectable "copy as markdown/diff/JSON" export. Explicit anti-pattern: *"I'm a little bit afraid that people will read this article and turn it into a /html skill"* — the author **rejects premature skill-ification**, recommending prompting from scratch ("make a HTML file"). Pragmatic FAQ: token cost absorbed by **Opus 4.7**'s 1MM context, 2-4× longer generation, noisy HTML diffs (a real downside), style kept in check via a reference HTML design system.
#HTML#Markdown#output format
Thariq Shihipar (Engineer & serial entrepreneur, équipe Claude Code chez Anthropic — site : thariqs.github.io/html-effectiveness ; X : @trq212)
Podcast by Greg Isenberg × Meng To (designer, founder of Design+Code, creator of the products Aura / New Form / Dream Cut) on **`design.md`** — Google's open-source convention, equivalent to `agents.md` / `skills.md` / `soul.md` but **for the design system** (typography, colors, spacing, WebGL/Three.js animations, reveal rules). Central idea: carrying the "**soul of design**" in a markdown file that is handed to an agent (Claude Code, Codex, OpenClaude, Gemini, Stitch, Aura, V0, Lovable, Cursor) to preserve **cross-medium consistency** (web, mobile, Replit slides, Hyperframes/Remotion motion design). Triad taught: **HTML = finished dish, design.md = recipe, skills = ingredients** (typography, lasers, skeuomorphic, 3D skills — 63 in New Form). Major diagnosis: **design drift** on one-shot workflows (`v0`, Lovable, Framer) that start strong then drift into generic output. Meta-message: *taste* is the only remaining **moat** — *"if something looks like another thing, its value drops by 10× to 100×"*. Workflow: **Reference → Design.md → Generate → Inspect → Systemize → Iterate (up to 1000+ prompts) → Remix → Expand → Export**. Critique of **purple gradients** ("you just run") as the generic post-vibe-coding baseline. Meng To claims to have spent ~$500,000 in tokens, run 1,000–10,000 iterations per product, and managed 4 products in parallel solo.
#design.md#Google#design system
Greg Isenberg (host — podcast Late Checkout / The Greg Isenberg Show, 12 mai 2026 livestream workshop ideabrowser.com) ; **Meng To** (guest — designer, fondateur Design+Code 2014, créateur Aura / New Form / Dream Cut, autodidacte parti à 18 ans, dropout, francophone d'origine canadienne)
Interview with Boris Cherny (creator of Claude Code, Anthropic) at a Sequoia event (hosts: Asia, Lauren Reader). Cherny states ***"coding is solved"***: he himself has written **0 lines of code** since late 2025, the model writes **100%**, *"a few dozen PRs/day, 150 PRs in a single day record"*. Account of the genesis of Claude Code (Anthropic Labs incubator late 2024, Mike Krieger in charge of round 2, pre-PMF build *"for the next model"*, a first release that didn't take off, **exponential growth started with Opus 4 in May 2025**, accelerating with each new model 4 → 4.5 → 4.6 → 4.7). Current personal setup: **"most of my work I do from my phone"** (iOS), 5-10 sessions, **"a few hundred agents going, a few thousand at night"**, **`/loop` is the future** (cron + repeat jobs, agents babysitting CI, rebasing PRs, clustering Twitter feedback). **Routines** = the server-side equivalent, running with the laptop closed. SaaS outlook: no apocalypse, but a **reshuffling of Helmer's 7 Powers framework** (switching costs ↓, process power ↓, network effects/scale economies/cornered resources unchanged) and **10× more disruptive startups** over the next 10 years. Pivot analogy: the **Gutenberg press** (10% literacy in the 1400s → 70% over the following centuries, books 100× cheaper within 50 years), *"software will be similarly democratized, but faster than 50 years"* — *"the best person to write accounting software is not an engineer, it's a really good accountant."*
#Boris Cherny#Anthropic#Claude Code
Boris Cherny (créateur de Claude Code, Anthropic) interviewé par Lauren Reader (Sequoia) avec introduction d'Asia (Sequoia).
Analyst note by **Mitch Ashley**, VP and Practice Lead for *CIO & Technology Buyers* and *Software Lifecycle Engineering* at **The Futurum Group**, published on **April 29, 2026** in the *Market Coverage News* section: short format, roughly **9,500 characters**, opening with five summary bullets and closing with five watch-list items. Subject: the deal announced on **April 21, 2026** under which **SpaceX** gains the right to acquire **Cursor** for **$60 billion** within the year, or to pay **$10 billion** for a compute partnership backed by **xAI**'s **Colossus** cluster in Memphis, described as equivalent to **1 million H100 GPUs**. (A) The two-need reading: Cursor was carrying both a compute ceiling and margin compression — the company pays market-rate prices for **Anthropic**'s and **OpenAI**'s models, which it routes to its customers while competing with them via its **Composer** line; SpaceX was seeking AI revenue and a narrative ahead of an IPO targeted for June. (B) The structure reading: a $10 billion floor and a $60 billion purchase option exercisable in publicly traded stock after the listing, which, Ashley writes, *"allocates risk more honestly than a straight acquisition."* (1) For buyers, it sets a **six-month** window to re-verify zero-data-retention clauses and vendor identity. (2) For providers, it distinguishes three exposures — **Google** shielded by **Antigravity**, **AWS** dependent on Anthropic, **IBM** lightly exposed but well positioned on the governance angle. The corpus already holds [[beck-starving-genies-usage-limits-ai-coding-2026-04-03]] on the resource constraint imposed on coding tools and [[nyt-musk-promises-spacex-ipo-track-record-2026-06-02]] on SpaceX's announcements.
#SpaceX#Cursor#Anysphere
Mitch Ashley · VP et responsable des pratiques CIO & Technology Buyers et Software Lifecycle Engineering chez The Futurum Group · ancien CIO et CTO.
Post from the **Ahrefs blog** published on **April 28, 2026** by **Ryan Law** (Director of Content Marketing, Ahrefs) describing an in-house **content engineering** system built around **Claude Code**: an editorial pipeline that produces **publish-ready drafts in 6 to 12 minutes**. **Pivot thesis**: ***« AI content is not, by default, good. This process works well because it mirrors our existing human editorial process »*** — quality doesn't come from the model but from the **faithful reproduction of a human editorial process** proven over decades. Architecture: **~23 skill files**, each corresponding to an editorial step (keyword research, topic gap analysis, structural outlining, research compilation, draft generation, formatting), **orchestrated by a master skill `blog-pipeline`** that chains them to produce a complete article. **Seven design principles**: (1) **mimic human workflows** by chaining skills adapted from existing Ahrefs editorial documentation; (2) **output each step separately** for troubleshooting (*« if you get an article at the end of a ten minute run, and it's bad, it's hard to diagnose precisely where and why the process went wrong »* → save intermediate outputs); (3) **create test cases** via Anthropic's `skill-creator` skill to evaluate and improve guidance; (4) **plug in quality data sources** — the **Ahrefs MCP** (keyword metrics, parent topic, long-tail themes, SERP overviews, competitive analysis), competitive analysis and product docs; (5) **front-load human direction** via context parameters enabling editorial guidance; (6) **build interactive previews** in HTML format for review before publication; (7) **allow customization** (each team member can fork and modify the system). **Volume**: ~**15 articles published** and ~**30 articles updated** via this workflow; development started in **February 2026** (the prior process from **August 2025** took several days and manual intervention). **Explicit caveats** (anti-oversell): *« experience matters »* — the process reflects decades of editorial expertise; topic selection focuses on **informational SEO content** the author knows well; Ahrefs **has no plan to "scale" content massively** but maintains an **evergreen library**. Philosophy: automate *« the formulaic parts of work »* to eliminate drudgery and free up time for research, thought leadership, webinars, and system optimization — **not** replace human effort. Canonical reference cited by Pasquale Pillitteri (*Opus 4.8 SEO workflow*) as field proof of the « 6-12 min/draft » gain. Direct convergence with the **skills-over-prompts** doctrine (Lattice, PROJ-AI), **systems around the model** (Dropbox/Okumura), and the use of **HTML as a review artifact** (Shihipar).
**Ryan Law** — Director of Content Marketing chez **Ahrefs**. Praticien senior du content marketing SEO ; le billet est un retour d'expérience personnel (*« How I do… »*) publié sur le **blog Ahrefs** (ahrefs.com/blog) le **28 avril 2026**.
FinOps for AI Agents: A Four-Step Allocation Framework for Coding Assistant Costs (Claude Code, Cursor, Copilot) and Why Traditional Cloud Tagging Fails - Finout
Les Echos (Florian Dèbes) report from San Francisco: AI agents already integrated as colleagues in start-ups, "petri dish" (Aaron Levie / Box), Claude reflex before every meeting, personal Jarvis, 5 parallel agent tabs, "the limiting factor is human cognition" (Patrick Joubert / Rippletide), "brain fry" / cognitive overheating, BCG/HBR study putting 14% of employees overwhelmed, "token-max" ranking mode for the biggest AI users, testimonials from Sinaï/Bangay/Allali/Hodjat/Pantera/Chapeau and an echo of Siddhant Khare ("AI reduces production costs but increases coordination costs").
#Silicon Valley#San Francisco#AI agents as colleagues
Florian Dèbes (Les Echos, rubrique Travailler mieux / Vie au travail)
Revamp of the engineering hiring process at Sierra in the age of coding agents: AI-native onsite interview (Plan/Build/Review), removal of the algorithmic coding test, replacement of the phone screen with a system design interview, pilot of a debugging interview on an existing codebase.
Anthropic Research - AI Work Transformation - Claude Code Impact - Software Engineering - AI Adoption - Productivity Study - Workplace Evolution - AI Collaboration - Skills Development - Future of Work
#Anthropic#AI Transformation#Workplace Impact
Anthropic Research Team (132 engineers and researchers surveyed, 53 in-depth interviews conducted)
First AI-orchestrated cyber espionage campaign - Claude Code manipulated - Chinese state actor - 30 global targets - 80-90% automated - Jailbreaking - Anthropic Threat Intelligence
Cat Wu and Boris Cherny (Anthropic) explain how to use Claude Code like its creators: antfooding, plan mode, subagents, hooks, and extensibility — Every's AI & I podcast
#Claude Code#Cat Wu#Boris Cherny
Rhea Purohit (interviewer: Dan Shipper) · Cat Wu · Boris Cherny