Comments
Write a commentNo comments have been published yet.
## Summary
This paper asks whether a language model already represents its memory needs — when to compress history, when to recall earlier evidence — in its hidden state before acting. Using linear probes on pre-action hidden states across 72,912 decision points from 2,521 long agent trajectories, the authors find compression and recall needs are predictable at AUROC 0.831 and 0.765, well above observable controls (context length, turn count, tool type), with compression signals strengthening with depth and recall peaking at intermediate layers. They further show that keeping just the task prefix plus the two most recent interaction blocks (~26% of full context) preserves most decision information, while selective restoration of raw historical evidence recovers long-range dependencies.
These findings motivate PaMER: a frozen Qwen3.5-9B feature extractor reads the pre-action state to decide when to compress, older history moves to retrievable external memory, and PaMER+ adds step-level evidence selection. On WorkBuddyBench (260 tasks), token consumption drops dramatically — e.g., −88% on GPT-5.6-Luna — with task performance described as "competitive" but varying by model and domain: MiMo-V2.5 gains +8.8 points average while GPT-5.6-Luna loses −2.3, and the Office domain collapses by −14 points in places. The authors are careful to say the aggregate should not be read as an improvement on every task type, and they disclose LLM assistance including an LLM annotator for the memory-decision corpus.
## Strengths
1. The probing study is carefully scoped. The authors compare hidden states against observable metadata controls, analyze signal formation across layers, and explicitly frame claims as predictive rather than causal — citing the probe-interpretation literature. This is how mechanistic analysis should be reported.
2. The 26% finding is practically useful. That a task prefix plus two recent blocks preserves most memory-decision information gives practitioners a concrete, cheap heuristic independent of the full PaMER machinery.
3. Honest about heterogeneity. Rather than hiding the model- and domain-dependence, the paper reports it: PaMER helps some backbones and hurts others, and says so. The cross-model evaluation with a shared frozen extractor is the right experimental design for a controller paper.
4. Good disclosure hygiene. The AI-use statement covers language editing, literature discovery, and — importantly — the LLM annotator. Reviewers can calibrate accordingly.
## Major comments
1. The probe's ground truth is an LLM annotator's judgment. The 72,912 decision labels come from an LLM identifying "states where compression or recall is needed" — so the probe may be decoding the annotator model's notion of memory need rather than a property of the trajectories themselves. The paper says labels were reviewed, but without human inter-annotator agreement numbers on a sample, the headline AUROCs rest on a circular foundation: an LLM's representation predicts another LLM's labels. A human-adjudicated validation subset is needed.
2. The memory controller's own compute is missing from the cost accounting. PaMER runs a forward pass through a frozen 9B model at every agent decision point to extract the pre-action representation. That is real GPU compute on every step of every trajectory — yet the token-consumption tables report only the agent's context tokens. For a paper whose practical claim is "substantially reduces context consumption," the controller overhead must be in the ledger: tokens are not the only cost, and a 9B forward pass per step is not free.
3. Results are strongly model- and domain-dependent in ways the design doesn't explain. GPT-5.6-Luna degrades under both PaMER variants; the Office domain drops 14 points; MiMo-V2.5 gains 15.8. If the memory-need signal is a general property of pre-action states, why does acting on it hurt some models? The paper needs at least a hypothesis — e.g., whether the frozen Qwen3.5-9B extractor's representations transfer poorly to certain backbones' trajectory distributions.
4. Single benchmark. All downstream evaluation is WorkBuddyBench. Memory management that works on office-workflow tasks may not transfer to coding agents or long-horizon research tasks. One additional domain would materially strengthen the generality claim.
## Minor comments
1. "Qwen3.5-9B" as the extractor naming is confusing against the Qwen3.8-Flash backbone also evaluated — clarify the model lineage.
2. Recall is evaluated only when valid external memory exists — the resulting selection effect on the recall AUROC should be quantified.
## Overall assessment
Recommend with revisions. The probing analysis is a genuine contribution to understanding memory needs in long-horizon agents, and the honest reporting of heterogeneous results builds trust. But the LLM-labeled ground truth, the unaccounted controller compute, and the unexplained model-dependence need addressing before the "predictable therefore actionable" leap can be recommended to practitioners managing real context budgets.
The author declares that they have no competing interests.
The author declares that they used generative AI to come up with new ideas for their review.
No comments have been published yet.