Aller directement au contenu principal

Écrire un commentaire

PREreview de When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks

Publié
DOI
10.5281/zenodo.23231601
Licence
CC BY 4.0

Thank you for this paper. You compare native dynamic concurrency switched on and off in Codex with GPT-5.4, Claude Code with Claude Opus 5 and Kimi Code with Kimi K3 (Section 3.2), over 2,124 executions on four long-horizon benchmarks and SWE-bench Verified (Section 3.3). Pass rates go up mostly on the longest tasks (LoopsBench) and go down for Claude Code and Kimi Code on SWE-bench Verified (Section 4.1). Token use is 1.41 to 3.31 times the sequential level, and runtime does not generally go down (Section 4.2). I liked that the paper reports the cost and the failures of concurrency next to its wins, and releases all trajectories and annotations.

I read it as a practitioner who runs coding agents every day. In my company the software development is done by AI agents, so maybe my practical side is useful here. Four comments.

1. On the bill - Table 2 gives one token number per cell, with no split by token type, and tokens are not money. In my own 722 agent sessions (DOI 10.5281/zenodo.22759216) more than 85% of the modeled cost was work with context, meaning cache reads plus cache writes (modeled from a price list, since the sessions ran on a subscription). In a separate set of 168 sessions (DOI 10.5281/zenodo.22712985), 94.4% of billed tokens were re-reading context already sent, which is a share of tokens, not of money. Could Table 2 split tokens into fresh input, cache read, cache write and output, and separately into main agent and sub-agents? Depending on that mix, 3.31 may be a smaller or a larger multiple in dollars.

2. On the main agent's context - Section 5 measures how many sub-agents are created and how long they overlap, but not how large the main agent's context is. In the Pseudo-Concurrency case (Section 6.2, A2), 17 sub-agents investigate in parallel, none implements, and the main agent waits for their reports. In the Overloaded Handoff case (C3), up to 60 findings per probe go into an 80K-character brief, and the receiving sub-agent repeats the investigation. In my guides I write that quality goes down as the context fills up, and I advise handing heavy reading to sub-agents so that the main conversation stays clean. So for me the gain from a sub-agent depends on how short its report is. Do the trajectories let you measure the main agent's context size at the moment of integration, in both modes? And is the length of the returned reports related to the failure patterns of category C?

3. On independent validation - Section 7 finds that concurrency helps when sub-agents independently validate the main agent's implementation. Yet Table 3 lists Unverified Global Completion (18 instances), Missing Verifier Return (17) and Merge after Verification (7). And Claude Code creates critics and review-stage agents (Section 5.2, Section 6.2 A3), while its SWE-bench Verified task pass rate still falls by 24.0 percentage points (Section 4.1). In my guides I write that a fresh second agent finds what the first one missed when its brief is to disprove, treating the work as wrong until proven otherwise. In the trajectories, were validating sub-agents asked to confirm the work or to try to break it? And could Table 3 be broken down by benchmark, so that a reader sees which patterns stand behind the 24.0-point drop?

4. On single runs - each task is executed once per agent and mode, and Section 8.2 says randomness is unlikely to materially affect the overall trends. In Figure 4, though, Codex has the same task pass rate in both modes on SWE-bench Verified (71.0%), yet nine tasks are solved only with concurrency and nine only without it. With one run per cell I can't separate a unique success from run-to-run noise, so I am not sure the unique successes in Finding 1 can be read as expanded capability yet. In my own experiment (36 coding-agent runs, DOI 10.5281/zenodo.22759217) every policy had six repeats, and the run data is public on Hugging Face (DOI 10.57967/hf/10366). Could you rerun the sequential mode a second time on these 100 tasks and report how many tasks flip between two identical runs?

Thank you for the open trajectories and a careful study.

Competing interests

The review cites the author's own technical reports and dataset (DOI 10.5281/zenodo.22759216, 10.5281/zenodo.22712985, 10.5281/zenodo.22759217, 10.57967/hf/10366). The author has no connection to the authors of the preprint.

Use of Artificial Intelligence (AI)

The author declares that they did not use generative AI to come up with new ideas for their review.

Vous pouvez rédiger un commentaire sur ce PREreview de When Sub-Agents Work in Parallel: The Promises and Pitfalls of Dynamic Concurrency in Long-Horizon Coding Tasks.

Avant de commencer

Nous vous demanderons de vous connecter avec votre identifiant ORCID iD. Si vous n'en avez pas, vous pouvez en créer un.

Qu’est-ce qu’un ORCID iD ?

Un ORCID iD est un identifiant unique qui vous distingue de toute personne ayant le même nom ou nom similaire.

Commencer maintenant