Synthetic Data Autophagy: The Ecological Limit of Generative Artificial Intelligence Evolution
Abstract
The article proposes a conceptual reconstruction of synthetic data autophagy as a systemic ecological limit to the development of generative artificial intelligence (GenAI). The analysis begins with two interrelated problems: the finite availability of high-quality human-generated data and the increasing contamination of training corpora by texts, images, and other artefacts produced by generative models themselves. It is argued that recursive training on synthetic outputs results not only in a measurable decline in model performance but also in a qualitative deterioration of representational capacity. This deterioration is evident in the disappearance of rare semantic features, the contraction of data distributions, the homogenisation of outputs, and a reduced capacity to generate novel or non-trivial content. The article distinguishes three dimensions of the ecological limits confronting GenAI. The first is a resource limit arising from the scarcity of sufficiently diverse and reliable human-generated data. The second is a representational limit, manifested in the progressive loss of rare elements within the data distribution. The third is an evolutionary limit, expressed in the disruption of model scaling laws associated with increasing parameter counts. The article demonstrates that synthetic data do not constitute a neutral extension of the training environment, since they reproduce and amplify the biases, errors, and limitations of previous generations of models. Synthetic data autophagy is therefore interpreted as a form of epistemic closure in which an artificial system increasingly reproduces not the external world but its own statistical traces. From a philosophical perspective, autophagy is examined as a symptom of the exhaustion of the self-training paradigm and as a basis for challenging the metaphor of the autonomous evolution of GenAI. The author considers several strategies for preserving the openness of training environments, including active data curation, grounding in interactive environments, and the co-evolution of human and artificial intelligence. The article concludes by proposing an epistemic ecology of GenAI as a field concerned with the conditions necessary for the sustainable coexistence of human and artificial cognitive systems.
About the Author
Sergey F. SergeevRussian Federation
Sergey F. Sergeev – D.Sc. in Psychology, Professor at the Department of Information Systems in Arts and Humanities, Faculty of Arts, St. Petersburg State University; Head of the Research Laboratory “Ergonomics of Complex Systems,” Peter the Great St. Petersburg Polytechnic University.
Saint Petersburg
References
1. Bender E.M. & Koller A. (2020) Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data. In: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185–5198). Retrieved from https://aclanthology.org/2020.acl-main.463/.
2. Dohmatob E., Feng Y., Subramonian A., & Kempe J. (2025) Strong Model Collapse. In: International Conference on Learning Representations (pp. 15656–15691). arXiv:2410.04840.
3. Dubrovsky D.I. & Sergeev S.F. (2023) Methodological Problems of Evaluating Generative Artificial Intelligence. Artificial Intelligence. Theory and Practice = Iskusstvennyy intellekt. Teoriya i praktika. No. 3 (3), pp. 2–10 (in Russian).
4. Fodor J.A. & Pylyshyn Z.W. (1988) Connectionism and Cognitive Architecture: A Critical Analysis. Cognition. Vol. 28, no. 1–2, pp. 3–71.
5. Gambetta D., Gezici G., Giannotti F., Pedreschi D., Knott A., & Pappalardo L. (2025) Characterizing Model Collapse in Large Language Models Using Semantic Networks and Next-Token Probability. arXiv preprint. arXiv:2410.12341.
6. Haenlein M. & Kaplan A. (2019) A Brief History of Artificial Intelligence: On the Past, Present, and Future of Artificial Intelligence. California Management Review. Vol. 61, no. 4, pp. 5–14.
7. Hoffmann J., Borgeaud S., Mensch A., Buchatskaya E., Cai T., Rutherford E., …, & Sifre L. (2022) Training Compute-Optimal Large Language Models. Advances in Neural Information Processing Systems. Vol. 35, pp. 30016–30030.
8. Holtzman A., Buys J., Du L., Forbes M., Choi Y., & Get P. (2020) The Curious Case of Neural Text Degeneration. In: Proceedings of International Conference on Learning Representations. arXiv:1904.09751.
9. Kolomeytsev I.V. (2020) Ecosystem Models of the Global World in the Reports of the Club of Rome. Sociology = Sotsiologiya. No. 2, pp. 351–361 (in Russian).
10. Meadows D.H., Meadows D.L., Randers J., & Behrens W.W. III (1972) The Limits to Growth: A Report for the Club of Rome’s Project on the Predicament of Mankind. New York: Universe Books.
11. Qin Z., Dong Q., Zhang X., Dong L., Huang X., Yang Z., …, & Wei F. (2025) Scaling Laws of Synthetic Data for Language Models. In: Conference on Language Modeling. arXiv:2503.19551.
12. Satharasi T. & Iyengar S.S. (2025) Future of AI Models: A Computational Perspective on Model Collapse. arXiv preprint. arXiv:2511.05535.
13. Seddik M.E.A., Chen S.-W., Hayou S., Youssef P., & Debbah M. (2024) How Bad Is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse. In: First Conference on Language Modeling. arXiv:2404.05090.
14. Searle J.R. (1980) Minds, Brains, and Programs. Behavioral and Brain Sciences. Vol. 3, no. 3, pp. 417–424.
15. Sergeev S.F. (2012) The Problem of Intelligent Symbionts in Technogenic Educational Environments. Educational Technologies = Obrazovatelnye tekhnologii. No. 3, pp. 36–50 (in Russian).
16. Sergeev S.F. (2013a) Intelligent Symbionts in Ergatic Systems. Scientific and Technical Journal of Information Technologies, Mechanics and Optics. No. 2 (84), pp. 149–154 (in Russian).
17. Sergeev S.F. (2013b) Intelligent Symbionts of Organized Technogenic Control Means for Mobile Objects. Mechatronics, Automation, Control = Mekhatronika, avtomatizatsiya, upravlenie. No. 9, pp. 30–36 (in Russian).
18. Sergeev S.F. (2021) Intelligent Technosymbiosis in Complex Human-Machine Systems. Ergodesign = Ergodizayn. No. 1 (11), pp. 70–76 (in Russian).
19. Shumailov I., Shumaylov Z., Zhao Y., Papernot N., Anderson R., & Gal Y. (2024) AI Models Collapse When Trained on Recursively Generated Data. Nature. Vol. 631, pp. 755–759.
20. Villalobos P., Ho A., Sevilla J., Besiroglu T., Heim L., & Hobbhahn M. (2024) Position: Will We Run Out of Data? Limits of LLM Scaling Based on Human-Generated Data. In: Proceedings of the 41st International Conference on Machine Learning (Vol. 235, pp. 49523–49544). arXiv:2211.04325.
21. Wu R. & Papyan V. (2024) Linguistic Collapse: Neural Collapse in (Large) Language Models. Advances in Neural Information Processing Systems. Vol. 37, pp. 137432–137473.
22. Zhu X., Cheng D., Li H., Zhang K., Hua E., Lv X., …, & Zhou B. (2025) How to Synthesize Text Data without Model Collapse? In: Proceedings of the 42nd International Conference on Machine Learning (Vol. 267, pp. 79746–79771). arXiv:2412.14689.
Review
For citations:
Sergeev S.F. Synthetic Data Autophagy: The Ecological Limit of Generative Artificial Intelligence Evolution. Russian Journal of Philosophical Sciences. 2026;69(1). (In Russ.)
































