Saltar al contenido principal

Escribe una PREreview

Adaptive Context Pruning for Multi-Turn LLM Conversations: An Implemented and Evaluated Alternative to Naive Full-History Retention

Publicada
Servidor
Preprints.org
DOI
10.20944/preprints202608.1172.v1

Large Language Models (LLMs) deployed in multi-turn conversational settings are typically served under a naive full-history pattern: the entire prior transcript is re-submitted as input context at every turn. This has two known consequences, usually studied separately: growing computational and billed cost, driven by the quadratic cost of self-attention over a linearly growing context, and degrading answer quality at long context lengths, documented in the “lost-in-the-middle” and “context rot” literature. We implement and empirically evaluate Adaptive Context Pruning (ACP), a contextmanagement strategy combining a sliding window, TF-IDF relevance-budgeted retention of older turns, and periodic extractive summarization, against naive full-history retention, on a controlled synthetic multi-turn benchmark with probe questions requiring recall of specific facts planted earlier in the conversation. Using a real LLM (Groq-hosted Llama-3.1-8B-Instruct) to answer all probes and grading answers against ground truth, we find that at a conversation length of 183 turns, naive retention consumes approximately 14.9× more context tokens than ACP, and its real measured API cost grows correspondingly; more importantly, naive retention’s probe accuracy collapses under this setup, from 100% up to 78 turns to 0% at 183 turns, where all of its requests are rejected outright by the provider for exceeding a free-tier tokens-per-minute limit for this model (verified in §6.4 to be a tier-specific rate limit, not the model’s much larger advertised context window), whereas ACP sustains 100% probe accuracy at every conversation length we tested (up to 183 turns), using a bounded context of under 700 tokens throughout, with sub-millisecond-to-low-single-digit-millisecond own bookkeeping overhead per turn. We report these results with full transparency about the experiment’s scope: a synthetic dataset, TF-IDF rather than neural embeddings for relevance scoring, extractive rather than LLM-based summarization, and a single small open-weight model behind a single inference provider. We present a theoretical cost model motivating why this failure mode is expected under naive retention, and discuss what would be required to validate these findings at larger scale and with production-grade components.

Puedes escribir una PREreview de Adaptive Context Pruning for Multi-Turn LLM Conversations: An Implemented and Evaluated Alternative to Naive Full-History Retention. Una PREreview es una revisión de un preprint y puede variar desde unas pocas oraciones hasta un extenso informe, similar a un informe de revisión por pares organizado por una revista.

Antes de comenzar

We will ask you to log in with your ORCID iD. If you don’t have an iD, you can create one.

What is an ORCID iD?

An ORCID iD is a unique identifier that distinguishes you from everyone with the same or similar name.

Comenzar ahora