Preview

Russian Journal of Philosophical Sciences

Advanced search
Open Access Open Access  Restricted Access Subscription Access

Do Large Language Models Think? A Technical Answer to a Philosophical Question

Abstract

The article investigates whether contemporary large language models instantiate processes that are functionally equivalent to thinking and, if they do not, where exactly the gap between machine and biological intelligence lies. Its central claim is that consciousness is neither an ontologically necessary condition for thought nor an epistemologically sufficient criterion for establishing the presence of thought. This claim, which can be traced back to Leibniz and has been repeatedly reinforced in the history of cognitive science, from Helmholtz to Friston, is extended here to artificial systems. In a functional sense, a language model may be said to possess understanding, which makes both the denial of such understanding on introspective grounds and its uncritical apologetic defense inadequate. The architectural paradox of the transformer – recursive productivity in the absence of a recursive architecture – empirically settles the Chomsky–Skinner dispute without resolving the theoretical question of the nature of recursive competence. The typological basis for this claim lies in the structural correspondence between two operations: the minimization of variational free energy in predictive coding and the computation of weighted similarity in the attention mechanism. In both cases, representations are iteratively refined through the reduction of the discrepancy between the model and the input data. The possibility of substantive intellectual exchange between humans and language models – joint reasoning, mutual correction, and productive disagreement – would be difficult to account for in the absence of this typological kinship. At the same time, the gap between biological and artificial intelligence can be traced across four levels: energetic, computational, semantic, and temporal. At each of these levels, what is at stake is not a merely quantitative lag but a qualitatively distinct mode of organizing informational processes. The article concludes that the question of machine thinking must therefore be reformulated: transformer-based and biological processes may be understood as different instantiations of a single computational type – the hierarchical reduction of uncertainty – which constitutes a typological feature of intelligence as such. The gap between these two forms of intelligence is real, but it lies within a common genus rather than between fundamentally different genera.

About the Author

Igor F. Mikhailov
Institute of Philosophy, Russian Academy of Sciences
Russian Federation

Igor F. Mikhailov – D.Sc. in Philosophy, Leading Research Fellow at the Department of Methodology of the Interdisciplinary Study of Man, Institute of Philosophy, Russian Academy of Sciences. 

Moscow



References

1. Chomsky N. (1959) A Review of B.F. Skinner’s Verbal Behavior. Language. Vol. 35, no. 1, pp. 26–58.

2. Dehghani M., Gouws S., Vinyals O., Uszkoreit J., & Kaiser Ł. (2019) Universal Transformers. In: Proceedings of the 7th International Conference on Learning Representations. arXiv:1807.03819.

3. Friston K., Parr T., & de Vries B. (2017) The Graphical Brain: Belief Propagation and Active Inference. Network Neuroscience. Vol. 1, no. 4, pp. 381–414.

4. Helmholtz H. von (1867) Handbuch der physiologischen Optik. Leipzig: Leopold Voss (in German).

5. Juschkewitsch A.P. & Kopelewitsch Ju.Ch. (1988) La correspondance de Leibniz avec Goldbach. Studia Leibnitiana. Vol. 20, no. 2, pp. 175–189 (in French and Latin).

6. Kim N. & Linzen T. (2020) COGS: A Compositional Generalization Challenge Based on Semantic Interpretation. In: Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (pp. 9087–9105). arXiv:2010.05465.

7. Lake B.M. & Baroni M. (2018) Generalization without Systematicity: On the Compositional Skills of Sequence-to-Sequence Recurrent Networks. In: Proceedings of the 35th International Conference on Machine Learning. Vol. 80, pp. 2873–2882. arXiv:1711.00350.

8. Leibniz G.W. (1982) Monadology (E.A. Bobrov, Trans.). In: Leibniz G.W. Works in 4 Vols. (Vol. 1, pp. 413–429). Moscow: Mysl’ (Russian translation).

9. Marcus G.F., Vijayan S., Bandi Rao S., & Vishton P.M. (1999) Rule Learning by Seven-Month-Old Infants. Science. Vol. 283, no. 5398, pp. 77–80.

10. Mikhailov I.F. (2024a) The Concept of Recursion in Cognitive Studies. Part I: From Mathematics to Cognition. Philosophical Problems of IT and Cyberspace. No. 1, pp. 58–76.

11. Mikhailov I.F. (2024b) The Concept of Recursion in Cognitive Studies. Part II: From Turing to Bayes to Consciousness. Philosophical Problems of IT and Cyberspace. No. 2, pp. 4–22.

12. Parr T., Pezzulo G., & Friston K. (2025a) Beyond Markov: Transformers, Memory, and Attention. Cognitive Neuroscience. Vol. 16, no. 1–4, pp. 5–23.

13. Parr T., Pezzulo G., & Friston K. (2025b) Closing the Box. Cognitive Neuroscience. Vol. 16, no. 1–4, pp. 43–48.

14. Rao R.P.N. & Ballard D.H. (1999) Predictive Coding in the Visual Cortex: A Functional Interpretation of Some Extra-Classical Receptive-Field Effects. Nature Neuroscience. Vol. 2, no. 1, pp. 79–87.

15. Searle J.R. (1980) Minds, Brains, and Programs. Behavioral and Brain Sciences. Vol. 3, no. 3, pp. 417–424.

16. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser Ł., & Polosukhin I. (2017) Attention Is All You Need. Advances in Neural Information Processing Systems. Vol. 30, pp. 5998–6008.


Review

For citations:


Mikhailov I.F. Do Large Language Models Think? A Technical Answer to a Philosophical Question. Russian Journal of Philosophical Sciences. 2026;69(1). (In Russ.)



Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 License.


ISSN 0235-1188 (Print)
ISSN 2618-8961 (Online)