Ir para o conteúdo principal

Escrever uma avaliação PREreview

Adaptive Context Pruning for Multi-Turn LLM Conversations: An Implemented and Evaluated Alternative to Naive Full-History Retention

Publicado
Servidor
Preprints.org
DOI
10.20944/preprints202608.1172.v1

Large Language Models (LLMs) deployed in multi-turn conversational settings are typically served under a naive full-history pattern: the entire prior transcript is re-submitted as input context at every turn. This has two known consequences, usually studied separately: growing computational and billed cost, driven by the quadratic cost of self-attention over a linearly growing context, and degrading answer quality at long context lengths, documented in the “lost-in-the-middle” and “context rot” literature. We implement and empirically evaluate Adaptive Context Pruning (ACP), a contextmanagement strategy combining a sliding window, TF-IDF relevance-budgeted retention of older turns, and periodic extractive summarization, against naive full-history retention, on a controlled synthetic multi-turn benchmark with probe questions requiring recall of specific facts planted earlier in the conversation. Using a real LLM (Groq-hosted Llama-3.1-8B-Instruct) to answer all probes and grading answers against ground truth, we find that at a conversation length of 183 turns, naive retention consumes approximately 14.9× more context tokens than ACP, and its real measured API cost grows correspondingly; more importantly, naive retention’s probe accuracy collapses under this setup, from 100% up to 78 turns to 0% at 183 turns, where all of its requests are rejected outright by the provider for exceeding a free-tier tokens-per-minute limit for this model (verified in §6.4 to be a tier-specific rate limit, not the model’s much larger advertised context window), whereas ACP sustains 100% probe accuracy at every conversation length we tested (up to 183 turns), using a bounded context of under 700 tokens throughout, with sub-millisecond-to-low-single-digit-millisecond own bookkeeping overhead per turn. We report these results with full transparency about the experiment’s scope: a synthetic dataset, TF-IDF rather than neural embeddings for relevance scoring, extractive rather than LLM-based summarization, and a single small open-weight model behind a single inference provider. We present a theoretical cost model motivating why this failure mode is expected under naive retention, and discuss what would be required to validate these findings at larger scale and with production-grade components.

Você pode escrever uma avaliação PREreview de Adaptive Context Pruning for Multi-Turn LLM Conversations: An Implemented and Evaluated Alternative to Naive Full-History Retention. Uma avaliação PREreview é uma avaliação de um preprint e pode variar de algumas frases a um parecer extenso, semelhante a um parecer de revisão por pares realizado por periódicos.

Antes de começar

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Começar agora