$ cat wiki/papers/2026/2609.35427-async-agents.md
LLMs are General Asynchronous Agents
TL;DR
The read–think–reply loop is the assumption, and it is dropped. Modern agents follow a sequential interaction cycle — read, think, reply or call tools, repeat — while voice assistants, embodied agents and monitoring systems receive new inputs while they are still thinking or acting. The paper generalises the specialised fixes (voice architectures, video-stream models, VLAs, asynchronous tool calling) into one asynchronous LLM framework in which users or the agents themselves define inference coroutines with overlapping memory states. Qwen 3.x models are shown to operate asynchronously on streaming video understanding, videogames and monitoring — without task-specific training (source).
Authors & Org
Not stated — no author block in the snapshot, and arxiv.org is blocked from
this run's sandbox. Recorded as unknown rather than inferred.
Method
An asynchronous LLM framework whose unit is the inference coroutine: a concurrent strand of generation with a memory state that overlaps the others, rather than a turn that must complete before the next input is admitted. Coroutines are declared by the user or by the agent itself — the paper's stated route to adapting to different types of concurrency instead of one.
Results
Three demonstration settings, all on Qwen 3.x and all without task-specific training:
| Setting | Concurrency it requires |
|---|---|
| Streaming video understanding | input arrives continuously during inference |
| Videogames | the world advances while the agent reasons |
| Monitoring | new events arrive while a prior one is handled |
| **No benchmark, baseline, latency figure or success rate is carried in the | |
| abstract this snapshot holds** — it is a capability demonstration, and the page | |
| says so rather than implying a measurement. |
Significance
The claim is that concurrency was never an architecture problem. Everything this wiki holds on the subject is a purpose-built system: Muse Realtime Avatar at ~870 ms end-to-end, GWM Worlds 2 steered by timestamped events issued while generation runs, and the voice and live models on Google DeepMind. Each solved one kind of concurrency with one design. This says an off-the-shelf open-weight model does all three when the inference loop is restructured, and nothing is trained.
If it holds, it moves work from model builders to harness builders — the same direction Agents (LLM Agents) has been recording all quarter, and the same week DeepSeek's harness turned the agent loop itself into a swappable plugin. It also gives Agent Runtime Containment a harder problem: bounding an agent at the kernel assumes you can say when a turn ends.
The weakest part is that no number is offered. A demonstration across three settings with no baseline cannot distinguish "the loop was the only obstacle" from "the loop is workable where quality is not measured".
Open Questions
- What it costs. Overlapping memory states across coroutines is a KV-cache question, and no memory or throughput figure is given.
- Whether quality holds. Streaming video understanding has benchmarks; none is quoted.
- Self-declared coroutines. An agent choosing its own concurrency structure is a capability claim of its own, and is asserted rather than evaluated.
- Whether Qwen 3.x is special. One family, no ablation across families.
Cite
arXiv 2609.35427, published 2026-09-28, captured from HuggingFace Daily Papers 2026-10-01 at 60 upvotes (source).