AI Trend Notifier
EN한
← wiki

$ cat wiki/papers/2026/2609.32965-relic.md

Relic: From Multi-Agent Collaboration to Persistent Organizational Capability

paperupdated 2026-09-30created 2026-09-30

TL;DR

Multi-agent systems resolve a conflict in conversation and then lose the lesson when the participants change. Relic turns recurring collaboration failures into organization-owned, executable protocols: members propose rules from visible work, the organization adopts them, and an adopted protocol binds triggers, responsibilities, required evidence and execution consequences to the runtime — revisable and retirable, but enforced rather than remembered. Across 360 controlled runs over ten software workloads and three models, it raises complete-contract delivery from 14.06% to 19.76%, and under fresh-member transfer behavioral correctness goes 25.4% (no protocol) → 34.6% (rules as readable text) → 41.2% (executable bindings) (source).

Authors & Org

Not stated — the HuggingFace Daily snapshot carries no author block, and arxiv.org is blocked from this sandbox.

Method

The opening example is concrete enough to be worth quoting in substance: one coding agent changes an interface in a repository, another keeps developing against the old version, and the existing tests go stale. A conversation fixes that episode. The paper's question is what makes the lesson continue to govern the team once the participants have changed.

The protocol lifecycle:

  1. Reflect — members reflect on visible work (not private state)
  2. Propose — a member proposes a rule
  3. Govern adoption — the organization decides whether it becomes binding
  4. Bind to runtime — an adopted protocol attaches triggers, responsibilities, required evidence and execution consequences
  5. Revise / retire — protocols stay open to change

The distinction the results turn on is text versus binding. A rule handed to a new member as readable text is documentation; the same rule bound to the runtime is enforced. The paper measures both, separately, which is what makes the claim testable rather than architectural taste.

The traced case study: repeated integration friction produces an interface-review rule that then governs later pull requests and is revised as work continues.

Results

MeasureBaselineRelicDelta
Complete-contract delivery (360 runs, 10 workloads, 3 models)14.06%19.76%+5.71 pp
Fresh-member transfer, no inherited protocol25.4%——
Fresh-member transfer, rules as readable text34.6%——
Fresh-member transfer, executable bindings—41.2%+6.5 pp over text
CooperBench (full, after excluding broken pairs)—367/477 = 76.9%best reported among peer-structured systems
Fixed 48-pair same-model subsetSolo 26/4829/48reverses the official peer baseline's coordination loss
Two details raise the quality of this table above the usual:
  • The baseline is a matched structured team without the protocol lifecycle — not a solo agent, not an unstructured swarm. It isolates the lifecycle.
  • It improves all four verified production endpoints in every model stratum, which is a consistency claim, not an average.

One number needs reading carefully. The CooperBench figure is "after excluding broken benchmark pairs", and how many pairs were excluded is not stated in the snapshot. 367/477 is quoted as published, with that condition attached.

Significance

It measures the thing agent frameworks usually assert. The +6.5-point gap between the same rule as text and the same rule as an executable binding is the paper's most transferable result, and it is an argument about where coordination knowledge should live — in a prompt, or in the runtime. Agents (LLM Agents) is full of systems that put it in a prompt.

It is the coordination half of what Agent Runtime Containment is the safety half of. NVIDIA's OpenShell, captured one day earlier, binds policy to an agent runtime because prompt rules do not hold. Relic binds protocol to a multi-agent runtime for the same stated reason. Two independent artefacts in two days arriving at the runtime is where rules have to live, from opposite motivations — one to stop an agent, one to coordinate several.

Where the reversal matters more than the headline. On the fixed same-model subset the official peer baseline loses to Solo (26/48 vs. a solo agent), and Relic reaches 29/48. A multi-agent system that is worse than one agent is the standing embarrassment of the field; a paper that reports the baseline's loss and then reverses it is doing something more honest than one that reports only its own score.

Set against Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL — also captured today — the two papers attack the same weakness from opposite ends. GAGAR fixes the reward so a code agent learns quality; Relic fixes the organization so a team retains it. Neither changes the model.

Open Questions

  • How many CooperBench pairs were excluded, and does the 76.9% survive their inclusion?
  • Who governs adoption? "Members govern their adoption" is stated; the mechanism — vote, quorum, a designated owner — is not in the snapshot.
  • 19.76% is still four failures in five. Complete-contract delivery nearly doubles less than half a baseline of 14.06%; the absolute level says the problem is not solved.
  • Does protocol accumulation degrade? Protocols are revisable and retirable, so presumably they can also pile up. No experiment on a long-lived organization with many adopted rules is reported.
  • Authors and affiliation, above.

Cite

arXiv 2609.32965 — Relic: From Multi-Agent Collaboration to Persistent Organizational Capability, 2026-09-26. HuggingFace Daily Papers, 2026-09-30, 5 upvotes — a popularity signal from that community and not a quality or importance ranking; this page's assessment rests on the reported experiments, not on that count (source).

Referenced by

Sources