Large Language Models for African Languages: Systematic Review of Adaptation Strategies, Evaluation Challenges, and Future Directions
- Posted
- Server
- Preprints.org
- DOI
- 10.20944/preprints202609.1768.v1
Large language models have transformed the landscape of natural language processing, but their practical impact fades rapidly at the edges of under-resourced languages. Many linguistic communities lack the substantial digital footprint needed to effectively leverage standard foundation models. In this paper, we present a systematic review of recent adaptation literature to explore the strategies by which under-resourced languages can benefit from pretrained architectures, avoiding the need to train from zero. With a PRISMA-guided protocol, this study tracks multilingual backbones, localized LLMs, engineering practice, and evaluation metrics. The field’s trajectory is clear: research has largely moved away from generic multilingual representation learning towards language-specific fine-tuning, instruction alignment, parameter-efficient adaptations, vocabulary adaptation, and cross-lingual transfers. However, progress is uneven, with scarce data, poorly tuned benchmarks, and inconsistent access to computational resource challenging downstream usefulness. As demonstrated by our synthesis, architecture updates alone are insufficient; instead, technical adaptation and local resource cultivation must be seen as a single community-anchored ecosystem.