The application of artificial intelligence to quantum information science has emerged as a frontier of research. This Technical Review summarizes how AI techniques, including machine learning, deep learning and language models, are establishing a new way to describe complex quantum systems.
Efficient characterization of large-scale quantum systems, especially those produced by quantum analog simulators and megaquop quantum computers, poses a central challenge in quantum science owing to the exponential scaling of the Hilbert space with respect to system size. Recent advances in artificial intelligence (AI), with its aptitude for high-dimensional pattern recognition and function approximation, have emerged as a powerful tool to address this challenge. A growing body of research has leveraged AI to represent and characterize scalable quantum systems, spanning from theoretical foundations to experimental realizations. Depending on how previous knowledge and learning architectures are incorporated, the integration of AI into quantum system characterization can be categorized into three synergistic paradigms: machine learning, deep learning and language models. This Technical Review discusses how each of these AI paradigms contributes to two core tasks in representing and characterizing quantum systems: quantum property prediction and quantum system reconstruction. These tasks underlie a range of applications, from quantum certification and benchmarking to enhancing quantum algorithms and identifying critical quantum phenomena. We also discuss key challenges and open questions, together with future prospects at the interface of AI and quantum science.
Artificial intelligence models can be leveraged to represent and characterize scalable quantum systems in a data-driven manner, enabling quantum property prediction and approximate quantum system reconstruction. Provably efficient machine learning models have been designed to characterize linear properties of scalable quantum systems and to classify quantum phases. Deep learning models offer powerful tools for predicting a wide range of quantum properties through representation learning, as well as for implicitly reconstructing quantum systems using generative modelling approaches. Language models, building on the generative pre-trained transformer architecture, provide a flexible framework for auto-regressively representing large families of quantum states, paving the way towards foundation models for quantum systems and enabling new directions for research and application.
Y.D. acknowledges the funding from A*STAR (H25-MRO3488) and NTU SUG (025257-00001). Y.-D.W. acknowledges funding from the National Natural Science Foundation of China through grant no. 12405022. J.E. acknowledges funding from the German BMFTR, Berlin Quantum, the Munich Quantum Valley, the Quantum Flagship (Millenion and PasQuans2), and the European Research Council. G.C. acknowledges support from the Hong Kong Research Grant Council through grant nos. SRFS2021-7S02, R7035-21F and 17310725. D.T. acknowledges the funding from NRF-P2024-001. Research at the Perimeter Institute is supported by the Government of Canada through the Department of Innovation, Science and Economic Development Canada and by the Province of Ontario through the Ministry of Research, Innovation and Science.
This section details the affiliations and contributions of the twelve authors of the article. Authors Yuxuan Du and Dacheng Tao are affiliated with the College of Computing and Data Science, Nanyang Technological University, Singapore, with Yuxuan Du also associated with the School of Physical and Mathematical Sciences. Yan Zhu and Giulio Chiribella are from the QICI Quantum Information and Computation Initiative, School of Computing and Data Science, The University of Hong Kong, with Giulio Chiribella also linked to the Department of Computer Science, University of Oxford, and Perimeter Institute for Theoretical Physics. Yuan-Hang Zhang and Yi-Zhuang You are from the Department of Physics, University of California, San Diego. Min-Hsiu Hsieh is from Hon Hai (Foxconn) Research Institute, Taipei. Patrick Rebentrost and Weibo Gao are from the Centre for Quantum Technologies, National University of Singapore, with Patrick Rebentrost also at the Department of Computer Science, National University of Singapore, and Weibo Gao at the School of Electrical & Electronic Engineering, Nanyang Technological University. Jens Eisert is from the Dahlem Center for Complex Quantum Systems, Freie Universitat Berlin, and Helmholtz-Zentrum Berlin für Materialien und Energie. Barry C. Sanders is from the Institute for Quantum Science and Technology, University of Calgary. Ya-Dong Wu, from the John Hopcroft Center for Computer Science, Shanghai Jiao Tong University, is the corresponding author. Y.D. and Y.-D.W. initiated the writing and drafted the first version of this Technical Review. M.-H.H., P.R., and W.G. contributed to the 'Machine learning paradigm' section, Y.Z. and Y.-Z.Y. to the 'Deep learning paradigm' section, and Y.-H.Z. to the 'Language model paradigm' section. J.E., G.C., D.T., and B.C.S. provided overall editing.
The authors declare no competing interests related to this work.
Nature Reviews Physics acknowledges Hans J. Briegel and Uman Khalid for their contributions to the peer review process of this work.
Publisherās note: Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. Information regarding reprints and permissions is also available.
The glossary defines key technical terms used in the article. **Attention mechanisms** are neural-network operations that assign different weights to parts of the input, crucial for transformer architectures to focus on relevant information for a given task. **Clifford circuits** are quantum circuits composed solely of Clifford gates, which can be efficiently simulated classically using the GottesmanāKnill theorem. A **Convolutional neural network (CNN)** is a deep learning architecture using convolutional layers to extract spatial and hierarchical features from grid-structured data like images. A **Fully connected neural network (FCNN)** connects every neuron in one layer to all neurons in the adjacent layer. **Graph neural networks (GNNs)** are designed for graph-structured data, aggregating information from node neighbors to learn representations capturing both node features and topological structure. A **Kernel machine** uses a kernel function to quantify similarity between data points in a feature space for non-linear learning with a simple linear structure. **Kraus operators** offer a mathematical framework describing quantum channels, representing any completely positive trace-preserving map as a sum of operators (Ī£i KiĻKiā ) with (Ī£i Kiā Ki = I). **Language models**, typically based on transformer architectures, learn statistical structures of sequences by modeling token probabilities, applicable beyond natural language to quantum systems. **Long short-term memory (LSTM)** is a specialized recurrent neural network (RNN) that addresses the vanishing gradient problem, enabling learning of long-range dependencies over extended sequences. A **Megaquop quantum computer** is an error-corrected quantum computer capable of executing approximately 1 million coherent quantum operations. **Non-Clifford gates** are quantum gates (e.g., T gate, Toffoli gate) essential for universal quantum computing and not generatable by Clifford circuits alone. A **Positive operator-valued measure (POVM)** is the most general mathematical description of a quantum measurement, being a set of positive semidefinite operators that sum to the identity. **Recurrent neural networks (RNNs)** process sequential or time-dependent data using recurrent connections and hidden states. The **SuāSchriefferāHeeger system** is a 1D tight-binding model with alternating hopping amplitudes, serving as a prototype for topological insulators. **Transformers** are neural-network architectures relying on self-attention mechanisms to model long-range dependencies in structured data, forming the basis for modern large language models. The **Transverse-field Ising model** describes a spin system with Ising interactions and a transverse field, used for studying quantum phase transitions, criticality, and non-equilibrium quantum dynamics.
This section provides bibliographic information about the article, including its acceptance date (June 1, 2026) and publication date (July 29, 2026). The Digital Object Identifier (DOI) is https://doi.org/10.1038/s42254-026-00962-5. Options to download citations and a shareable link for accessing the content, provided by the Springer Nature SharedIt content-sharing initiative, are also mentioned.