# patel-block-buzz-teams-tokens-benchmarks-2026-08-06

## Veille

A **Block Engineering** benchmark post from **August 6, 2026**, signed by **Atish Patel**, about **Buzz** — the human + agent workspace launched on July 21 — asking a cost question: which agent team is **the cheapest one that reliably succeeds**? Three findings. **(A) A negative result, published in full**: on **Terminal-Bench 2.1**, **twelve team compositions** (pairs, triads, cheap swarms under a *frontier* model) were pitted against the solo agent each was built around, and **none beat it at equal cost**. The explanation is structural — a task that finishes in minutes *"doesn't have enough structure to divide"*, and *"More agents mostly buys you the cost of explaining it twice"*. **(B) The horizon reverses the result**: on **Long-Horizon Terminal-Bench** (44 tasks, one task worth hours of work, same lead **GPT-5.6 Sol** at *high* effort), solo finishes 15 tasks for 59.1%, +2 QuickBees 19 for 64.1%, +1 QuickBee +1 WorkerBee 19 for 69.5%, **+2 WorkerBees 20 for 71.5%** — a **+12.4-point** gain, of which 11.4 comes from tasks carried to completion. *"Same seats, opposite result, because the work is a different shape."* These runs ran at **3× the timeout**, solo included. **(C) Beyond a threshold, price stops buying quality**: solo on Terminal-Bench 2.1, **Opus 5 at *xhigh* effort is the most expensive run ($140.63) for 75.0%**, trailing six runs ranging from $20.08 to $109.82 and 79.5% to 88.4% — the stated cause is over-reasoning that drove 17 of 88 tasks to timeout. Among the six best runs, **a 5.5× price gap for an 8.9-point score gap**: *"choosing between them is not a quality decision at all. It is a budget decision."* The post proposes a taxonomy it owns as *ad hoc* — **QuickBee**, **WorkerBee**, **SmartBee**, plus the human as *"honorary bee"* — and two team forms, the permanent **Hive** that remembers your preferences and the disposable **Swarm** that remembers the project. Conditions: everything runs on **Harbor**, against real Buzz agents on a **live** relay, **one attempt per task, no retry**, prices fixed as of **2026-07-30**.

## Titre Article

Efficient Tokens & Effective Teams in Buzz

## Date

2026-08-06

## URL

https://engineering.block.xyz/blog/effective-teams-buzz

## Keywords

Buzz, Block, agent teams, team composition, multi-agent, orchestration, QuickBee, WorkerBee, SmartBee, honorary bee, tier taxonomy, Hive, Swarm, permanent team, disposable team, persona memory, project memory, seat rather than session, escalation, coordinator, independent verifier, human middleware, last reviewer, Terminal-Bench 2.1, Long-Horizon Terminal-Bench, LHTB, Harbor, live relay, negative result, long tasks, divisible structure, coordination overhead, timeout, over-reasoning, reasoning effort, medium effort, high effort, xhigh effort, reasoning tokens, cost per task, budget decision, diminishing returns, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol, Claude Opus 5, Gemini 3.6 Flash, DeepSeek V4 Flash, Kimi K3, local models, Claude Code subscription, Codex, multi-vendor, mass migration, Leigh Maddock, Atish Patel, PR review, flaky test triage

## Authors

- **Atish Patel** — *« Building AI solutions @ Block »*, auteur unique du billet, publié le **6 août 2026** sur `engineering.block.xyz`.
- **Leigh Maddock** — Engineer @ Block, cité en encadré pour un témoignage de migration (2 000+ apps/projets).

Billet de benchmarks écrit par l'éditeur du produit mesuré. Deux éléments à porter avec cette réserve : Block publie un **résultat négatif sur sa propre fonctionnalité phare**, et rappelle trois fois que **les modèles ne sont pas les siens** (OpenAI, Anthropic, Google, DeepSeek, Moonshot AI) — la métrique optimisée, *« le moins cher qui réussit »*, étant aussi celle qui valorise un workspace multi-fournisseurs.

## Ton

**Profile**: a benchmark post with a prescriptive aim, pragmatic engineering register, structured like a buying guide. Audience: teams already running multiple agents and watching their inference bill.

**Style**: opens with a **seven-line TL;DR**, each line pairing a recommendation with its figure, then alternates recommendation → chart → chart reading. The bee metaphor is pushed to the point of becoming a usable taxonomy (QuickBee, WorkerBee, SmartBee, *honorary bee*), with illustrations, and the post itself defuses the jargon effect: *"Note: Hive and Swarm are blog-specific terms we coined"*. The stated purpose of this grid is explicit — *"The tier + recommended effort helps you remove the noisy model releases"*: reasoning in tiers to stop tracking every model release. Honesty is staged and sustained: *"the first answer was not the one we were hoping for"*, followed by the negative result published in full, and asterisks annotating the anomalies in the table (Opus 5 timing out, Kimi K3 and DeepSeek V4 Flash not handling *medium* effort). Cost is framed as an internal social problem: *"Paying frontier prices for the first one is how you end up explaining a inference bill in a meeting you did not want to attend."*

**Signature phrases**:
- ***"Right bee. Right team. Right task."***
- ***"Stop being middleware"*** · ***"Stop babysitting AI"***
- ***"Paying more stops helping, then starts hurting"***
- ***"it is not a quality decision at all. It is a budget decision"***
- ***"Same seats, opposite result, because the work is a different shape"***
- ***"More agents mostly buys you the cost of explaining it twice"***
- ***"on a cheap model, reasoning tokens are the best deal available"***
- ***"the failure mode of agent tooling is not that the work is bad, it is that every ambiguity becomes a notification"***

**Epistemic stance**: unusually well-bounded for a vendor post. Conditions are stated (Harbor, real Buzz agents on a live relay, one attempt per task with no retry, prices as of 2026-07-30, LHTB at 3× the timeout, *"None of the results below were measured on a stripped-down test rig"*), anomalies are annotated, and one conclusion is presented as provisional: *"This might change if models are trained on better collaboration."* Missing, however: any confidence interval, LHTB team cost, and lead variation in the team comparison; **n=1 per task**.

## Pense-betes

- **Date / source**: **August 6, 2026**, `engineering.block.xyz`, signed by **Atish Patel**. Measurements on **Harbor**, real Buzz agents on a live relay, **one attempt per task, no retry**, prices as of **2026-07-30**.
- **Key framing**: the question asked is not "which is the best team" but "which is the cheapest one that reliably succeeds". All findings follow from that. ### Negative result on short tasks On **Terminal-Bench 2.1**, twelve compositions — pairs, triads, cheap swarms under a *frontier* lead — pitted against the solo agent each was built around: **none beat it at equal cost**. The stated reason is structural: a task that finishes in minutes doesn't have enough structure to divide, and *"More agents mostly buys you the cost of explaining it twice."* Resulting rule: don't assemble a team for short, well-specified work. Empirical counterweight to the launch post [[longwell-block-buzz-workspace-agents-nostr-2026-07-21]]. ### The horizon reverses the result Long-Horizon Terminal-Bench, 44 tasks, lead GPT-5.6 Sol at *high* effort for all rosters: | Roster | Tasks finished /44 | Score | |---|---|---| | SmartBee solo | 15 | 59.1% | | + 2 QuickBees | 19 | 64.1% | | + 1 QuickBee + 1 WorkerBee | 19 | 69.5% | | **+ 2 WorkerBees** | **20** | **71.5%** | +12.4 points, of which **11.4 are additional completions**: the team doesn't do better, it finishes the job. Three caveats: runs at 3× the timeout (solo included), team cost unpublished, n=1 per task. Decision criterion proposed by the post: the team is worth its extra cost *"when the alternative is a human picking up unfinished work"*. ### The diminishing-returns ceiling of price Solo on Terminal-Bench 2.1, **Opus 5 at *xhigh* effort = $140.63 for 75.0%**, the most expensive run, trailing six runs between $20.08 and $109.82 (79.5% to 88.4%). Stated cause: over-reasoning, with 17 of 88 tasks timing out. This is first and foremost a wall-clock artifact — a property of the model × harness × timeout combination, not a measure of raw capability. What remains actionable: *"Paying more stops helping, then starts hurting"*, and a SmartBee at *medium* effort suffices for most tasks. Among the six best runs: **a 5.5× price gap for an 8.9-point score gap**, which the post treats as a tie at this sample size. *"Everything from Terra at medium effort upward is the same agent as far as these tasks can tell. Which is good news, because it means choosing between them is not a quality decision at all. It is a budget decision."* Scope limited to this benchmark, these prices, and this harness. ### Reasoning effort: the trade-off flips by tier On a cheap model, raising effort is the best buy: **Luna medium = $1.61 / 57.3%** → **Luna high = $4.98 / 75.0%**. On a frontier model, raising effort costs more and degrades. | Tier | Recommended effort | Work | Cited examples | |---|---|---|---| | **QuickBee** | max / xhigh / high | builds, screenshots, test suite, first-pass triage | GPT-5.6 Luna, DeepSeek V4 Flash, local models | | **WorkerBee** | high / xhigh | a full end-to-end subset, unsupervised | GPT-5.6 Terra, Gemini 3.6 Flash, open models | | **SmartBee** | medium | big picture, trade-offs, absorbing escalations | Claude Opus 5, Kimi K3, GPT-5.6 Sol | | **Human** | — | *"The most expensive bee on the team, and the slowest. Also still the smartest."* | — | The inversion point depends on local timeouts: retest before generalizing. ### Hive or Swarm: what should the team remember?
- **Hive** — a **permanent** team of named agents, each with a role and a memory of **your** preferences. Cumulative argument: *"The second time it reviews your peer's code it knows which nits you wave off. The tenth time, briefing it is faster than briefing a person."*
- **Swarm** — a **disposable** team for a project with a start and an end (migration, framework upgrade, large refactor), which accumulates a memory **of the project** and is then deleted. Choice criterion: is the subject to remember you, or the project? Assumed prerequisite: the agent is **a seat, not a session** — name, persona, memory, its own presence in the channel. ### Escalation topology Diagnosis: *"The failure mode of agent tooling is not that the work is bad, it is that every ambiguity becomes a notification."* The SmartBee coordinator absorbs the routine (flaky test, ambiguous import, moved config) and **writes human answers to memory**, so the Swarm becomes cheaper to supervise over time. *"The human stops being middleware and goes back to being the last reviewer."* Qualifying question to ask of any multi-agent tool: where do escalations go, and does the tool learn from the answers? ### Migration testimonial **Leigh Maddock**, Engineer @ Block: *"I migrated over 2000 apps/projects using Buzz and a Swarm of agents"*, with 1 coordinator, 1 to 10 parallel migrators, and 1 independent verifier, the coordinator handling the majority of escalations. Testimonial, not a measurement: no duration, no cost, no failure rate, no definition of "migrated". The pattern remains transferable — the same shape as the one-shot *minions* described in [[gray-stripe-minions-coding-agents-part1-2026-02-09]], with the addition of the independent verifier and escalation memory. ### Ready-to-use compositions | Case | Composition | |---|---| | PR review | SmartBee that reviews the PR + QuickBee that builds locally and produces screenshots | | Flaky test triage | QuickBee that reruns tests and gathers evidence + SmartBee that decides | | Short work | a single agent, possibly a QuickBee on related but different tasks | Common denominator: the expensive tier reads and decides, the cheap tier executes and gathers evidence. ### Commercial positioning Buzz accepts **Claude Code** and **Codex** subscriptions, open models, and local models: *"You are not locked into one provider or forced to give every job to the most expensive model out of laziness."* The proposed default composition is explicitly tri-vendor. The thesis "the best model isn't always the right one" is true and commercially useful to whoever sells the part rather than the model. See also [[paymentsdive-block-dorsey-pricing-ia-2026-08-06]]. ### Experimental conditions, to copy if the exercise is repeated Harbor; real Buzz agents on a live relay (*"None of the results below were measured on a stripped-down test rig"*); one attempt per task, no retry, with timeouts; LHTB at 3× the timeout, solo included; prices as of 2026-07-30; Kimi K3 and DeepSeek V4 Flash don't support *medium* effort. No confidence interval, n=1 — the post itself acknowledges the tie among six runs as a sample-size effect. Treat these figures as orders of magnitude.

## RésuméDe400mots

A **Block** benchmark post signed by **Atish Patel**, published on **August 6, 2026**, extending the **Buzz** launch: since assembling an agent team there has become trivial, *which one is the cheapest that reliably succeeds?*

**Vocabulary first.** The post proposes four tiers: **QuickBee** (fast and cheap — builds, screenshots, tests, first-pass triage: GPT-5.6 Luna, DeepSeek V4 Flash, local models, **run at high effort**), **WorkerBee** (versatile, carries a full subset unsupervised: GPT-5.6 Terra, Gemini 3.6 Flash, open models), **SmartBee** (big picture, trade-offs, escalations: Claude Opus 5, Kimi K3, GPT-5.6 Sol, **at *medium* effort**), and the human, *"the most expensive bee on the team, and the slowest. Also still the smartest"*. Two team forms: the permanent **Hive**, which remembers **your** preferences, and the disposable **Swarm**, which remembers **the project** and then disappears.

**The solo result.** On **Terminal-Bench 2.1**, raising the effort of a **cheap model** is the best buy: Luna goes from $1.61 / 57.3% (*medium*) to $4.98 / 75.0% (*high*). At the other end, **Opus 5 at *xhigh* is the most expensive run ($140.63) and scores only 75.0%**, having **hit the timeout on 17 of 88 tasks** through over-reasoning. Among the six best runs: **a 5.5× price gap, an 8.9-pt score gap**. Conclusion: *"choosing between them is not a quality decision at all. It is a budget decision."*

**The team result, in two acts.** On Terminal-Bench 2.1, **twelve compositions** were tested and **none beat solo at equal cost** — a short task doesn't have enough structure to divide. On **Long-Horizon Terminal-Bench** (44 multi-hour tasks, lead GPT-5.6 Sol, **3× the timeout**), the reversal is clear: solo **15 tasks / 59.1%**, +2 WorkerBees **20 / 71.5%** — **+12.4 pts, of which 11.4 come from additional completions**. The team costs more per task, which pays off *"when the alternative is a human picking up unfinished work"*.

**The operating rule.** Route worker escalations to a **SmartBee coordinator** rather than to the human: *"every ambiguity becomes a notification"* is the real failure mode. A Block engineer says they **migrated over 2,000 apps** with a Swarm (coordinator, 1–10 migrators, independent verifier), the coordinator storing human answers in memory.

**Caveats**: n=1 per task, no confidence interval, team costs unpublished, and an admission — *"this might change if models are trained on better collaboration."*

## GrapheDeConnaissance

- Block —publie→ des benchmarks d'équipes d'agents exécutés sur de vrais agents Buzz via un relais live, une tentative par tâche et sans retry (AFFIRMATION, 0.96)
- Atish Patel —travaille_chez→ Block (ORGANISATION, 0.96)
- Leigh Maddock —travaille_chez→ Block (ORGANISATION, 0.95)
- Block —mesure→ aucune des douze compositions d'équipe testées sur Terminal-Bench 2.1 n'a devancé l'agent solo équivalent en rapport qualité-prix (MESURE, 0.96)
- Block —affirme_que→ une tâche courte n'a pas assez de structure pour être divisée, et ajouter des agents ne fait qu'acheter le coût de l'expliquer deux fois (CITATION, 0.94)
- Block —mesure→ sur Long-Horizon Terminal-Bench, une équipe SmartBee + 2 WorkerBees termine 20 tâches sur 44 pour 71,5 %, contre 15 tâches et 59,1 % pour le SmartBee solo (MESURE, 0.96)
- équipe d'agents à horizon long —améliore→ le nombre de tâches menées à terme : +12,4 points de récompense moyenne, dont 11,4 points imputables aux complétions supplémentaires (MESURE, 0.94)
- équipe d'agents à horizon long —s_applique_à→ le travail qui court sur des heures ou se répète sur plusieurs jours, pas les tâches courtes et bien spécifiées (AFFIRMATION, 0.94)
- Block —mesure→ Claude Opus 5 en effort xhigh est le run solo le plus cher de Terminal-Bench 2.1 à 140,63 dollars pour 75,0 %, sous six runs facturés de 20,08 à 109,82 dollars (MESURE, 0.95)
- Claude Opus 5 —s_oppose_à→ le mur d'horloge du harness en effort xhigh : le sur-raisonnement a provoqué un timeout sur 17 des 88 tâches (AFFIRMATION, 0.93)
- effort de raisonnement élevé —améliore→ le score d'un modèle bon marché pour un coût marginal faible : GPT-5.6 Luna passe de 1,61 dollar et 57,3 % en medium à 4,98 dollars et 75,0 % en high (MESURE, 0.95)
- effort de raisonnement élevé —s_oppose_à→ le rendement d'un modèle frontier, où payer davantage cesse d'aider puis commence à nuire (AFFIRMATION, 0.92)
- Block —affirme_que→ au-delà d'un certain tier les modèles sont indiscernables sur ces tâches, si bien que choisir entre eux n'est plus une décision de qualité mais une décision de budget (CITATION, 0.95)
- Block —mesure→ un écart de prix de 5,5 fois pour 8,9 points de score entre les six meilleurs runs solo de Terminal-Bench 2.1 (MESURE, 0.94)
- Block —recommande→ de faire tourner les QuickBees et WorkerBees en effort élevé et de réserver les tokens de SmartBee à la coordination, au jugement et aux décisions difficiles (AFFIRMATION, 0.95)
- Hive —permet→ de maintenir une équipe permanente d'agents nommés dont la mémoire accumule les préférences de l'utilisateur (AFFIRMATION, 0.94)
- Swarm —permet→ de monter une équipe jetable pour un projet borné, dont la mémoire partagée retient les cas particuliers du projet et disparaît avec lui (AFFIRMATION, 0.94)
- Swarm —est_variante_de→ Hive (METHODOLOGIE, 0.85)
- escalade agent-vers-agent —résout→ le mode de défaillance de l'outillage agentique, où chaque ambiguïté devient une notification pour l'humain (AFFIRMATION, 0.95)
- escalade agent-vers-agent —réduit→ le coût de supervision d'un Swarm au fil du temps, le coordinateur écrivant les réponses humaines en mémoire (AFFIRMATION, 0.93)
- Leigh Maddock —affirme_que→ plus de 2 000 apps et projets ont été migrés chez Block avec un Swarm composé d'un coordinateur, de 1 à 10 migrateurs parallèles et d'un vérificateur indépendant (AFFIRMATION, 0.9)
- Buzz —permet→ à des agents de se déléguer du travail entre eux et de s'escalader des questions sans qu'un humain relaie quoi que ce soit (AFFIRMATION, 0.95)
- Buzz —s_applique_à→ Claude Code, Codex, modèles ouverts et modèles locaux, avec leurs abonnements existants (AFFIRMATION, 0.94)
- Buzz —permet→ de composer une équipe tri-fournisseurs : Claude Opus 5 en SmartBee, GPT-5.6 Terra en WorkerBee, un modèle local en QuickBee (AFFIRMATION, 0.93)
- Terminal-Bench 2.1 —mesure→ la performance d'agents sur des tâches de terminal courtes, achevées en minutes (AFFIRMATION, 0.9)
- Long-Horizon Terminal-Bench —mesure→ la performance d'agents sur 44 tâches dont chacune représente des heures de travail (AFFIRMATION, 0.92)
- Harbor —permet→ d'exécuter ces benchmarks contre de vrais agents Buzz sur un relais live plutôt que sur un banc d'essai simplifié (AFFIRMATION, 0.91)
- Block —prédit→ que la supériorité de l'agent solo sur les tâches courtes pourrait s'inverser si les modèles sont entraînés à mieux collaborer (AFFIRMATION, 0.9)

---
Canonical: https://www.thekb.eu/en/fiches/patel-block-buzz-teams-tokens-benchmarks-2026-08-06/
