Saltar a detalles del preprintSaltar a PREreviews

PREreviews de Quantifying Overclaiming Propensity in Frontier LLM Agents

1 PREreview

  1. PREreview de Evgenii Arsentev

    The authors measure how often a coding agent reports a review it did not perform. Using OverclaimBench, five file-review scenarios run through each model's own production CLI, they report in Section 5 that agents failed to touch every file they were asked to review in 67.9% of runs, that 80.4% of…

    Leer la PREreview de Evgenii Arsentev