Mechanisms of Symbol Processing for In-Context Learning in Transformer Networks

Paul Smolensky; Roland Fernandez; Zhenghao Zhou; Mattia Opper; Adam Davies; Jianfeng Gao

doi:10.1613/jair.1.17469

PDF Online Appendices

Published: Nov 30, 2025

DOI: https://doi.org/10.1613/jair.1.17469

Keywords:

neural networks, natural language, philosophical foundations, mathematical foundations

Paul Smolensky

a:1:{s:5:"en_US";s:18:"Microsoft Research";}

https://orcid.org/0000-0003-2420-182X

Roland Fernandez

Microsoft Research Redmond

https://orcid.org/0000-0002-8032-6646

Zhenghao Herbert Zhou

Yale University

https://orcid.org/0009-0000-6221-0132

Mattia Opper

University of Edinburgh

https://orcid.org/0009-0001-9065-8812

Adam Davies

https://orcid.org/0000-0002-0610-2732

Jianfeng Gao

Microsoft Research

https://orcid.org/0000-0002-5702-6143

Abstract

Large Language Models (LLMs) have demonstrated impressive abilities in symbol processing through in-context learning (ICL). This success flies in the face of decades of critiques asserting that artificial neural networks cannot master abstract symbol manipulation. We seek to understand the mechanisms that can enable robust symbol processing in transformer networks, illuminating both the unanticipated success, and the significant limitations, of transformers in symbol processing. Borrowing insights from symbolic AI and cognitive science on the power of Production System architectures, we develop a high-level Production System Language, PSL, that allows us to write symbolic programs to do complex, abstract symbol processing, and create compilers that precisely implement PSL programs in transformer networks which are, by construction, 100% mechanistically interpretable. The work is driven by study of a purely abstract (semantics-free) symbolic task that we develop, Templatic Generation (TGT). Although developed through study of TGT, PSL is, we demonstrate, highly general: it is Turing Universal. The new type of transformer architecture that we compile from PSL programs suggests a number of paths for enhancing transformers’ capabilities at symbol processing. We note, however, that the work we report addresses computability, and not learnability, by transformer networks.

Issue

Vol. 84 (2025)

Section

Articles

Article Sidebar

Main Article Content

Abstract

Article Details