The Evolution of Deep Learning: From Perceptrons to Generative Foundation Models
- Publicado
- Servidor
- Preprints.org
- DOI
- 10.20944/preprints202607.1764.v1
Deep learning has evolved from early perceptron models to modern foundation models through a series of breakthroughs that addressed fundamental limitations in representation learning, optimization, computational efficiency, and scalability. While numerous surveys review specific architectures, learning paradigms, or application domains, few examine the field from a historical and evolutionary perspective. This paper traces the evolution of deep learning by identifying the key discoveries that shaped its development, the bottlenecks that motivated each advance, and the innovations that enabled subsequent progress. Without an understanding of the historical motivations behind major breakthroughs, researchers may inadvertently revisit previously resolved challenges or overlook design principles that have shaped modern deep learning architectures. Understanding why particular architectures emerged and how they addressed existing limitations can help researchers avoid revisiting previously solved problems and make more informed design decisions for future models. Beginning with the origins of artificial neural networks and the introduction of backpropagation, we examine the resurgence of deep learning driven by large-scale datasets, GPU computing, and architectural innovations, followed by developments in convolutional, recurrent, attention-based, graph, generative, and foundation models. Rather than providing an exhaustive review, this paper focuses on the seminal methods that established new research directions and highlights representative applications in which these advances have been successfully adopted. By connecting major breakthroughs with the challenges they addressed, this survey equips new researchers with a principled understanding of deep learning’s evolution and a foundation for developing novel approaches to emerging problems.