Aller directement au contenu principal

Rédiger un PREreview

Adaptive Context Pruning for Multi-Turn LLM Conversations: An Implemented and Evaluated Alternative to Naive Full-History Retention

Publié
Serveur de preprints
Preprints.org
DOI
10.20944/preprints202608.1172.v1

Large Language Models (LLMs) deployed in multi-turn conversational settings are typically served under a naive full-history pattern: the entire prior transcript is re-submitted as input context at every turn. This has two known consequences, usually studied separately: growing computational and billed cost, driven by the quadratic cost of self-attention over a linearly growing context, and degrading answer quality at long context lengths, documented in the “lost-in-the-middle” and “context rot” literature. We implement and empirically evaluate Adaptive Context Pruning (ACP), a contextmanagement strategy combining a sliding window, TF-IDF relevance-budgeted retention of older turns, and periodic extractive summarization, against naive full-history retention, on a controlled synthetic multi-turn benchmark with probe questions requiring recall of specific facts planted earlier in the conversation. Using a real LLM (Groq-hosted Llama-3.1-8B-Instruct) to answer all probes and grading answers against ground truth, we find that at a conversation length of 183 turns, naive retention consumes approximately 14.9× more context tokens than ACP, and its real measured API cost grows correspondingly; more importantly, naive retention’s probe accuracy collapses under this setup, from 100% up to 78 turns to 0% at 183 turns, where all of its requests are rejected outright by the provider for exceeding a free-tier tokens-per-minute limit for this model (verified in §6.4 to be a tier-specific rate limit, not the model’s much larger advertised context window), whereas ACP sustains 100% probe accuracy at every conversation length we tested (up to 183 turns), using a bounded context of under 700 tokens throughout, with sub-millisecond-to-low-single-digit-millisecond own bookkeeping overhead per turn. We report these results with full transparency about the experiment’s scope: a synthetic dataset, TF-IDF rather than neural embeddings for relevance scoring, extractive rather than LLM-based summarization, and a single small open-weight model behind a single inference provider. We present a theoretical cost model motivating why this failure mode is expected under naive retention, and discuss what would be required to validate these findings at larger scale and with production-grade components.

Vous pouvez rédiger un PREreview de Adaptive Context Pruning for Multi-Turn LLM Conversations: An Implemented and Evaluated Alternative to Naive Full-History Retention. Un PREreview est une évaluation d'un preprint et peut varier de quelques phrases à un rapport détaillé, semblable à un rapport d'évaluation par les pairs organisé par une revue.

Avant de commencer

Nous vous demanderons de vous connecter avec votre identifiant ORCID iD. Si vous n'en avez pas, vous pouvez en créer un.

Qu’est-ce qu’un ORCID iD ?

Un ORCID iD est un identifiant unique qui vous distingue de toute personne ayant le même nom ou nom similaire.

Commencer maintenant