TR2026-130

Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes


    •  Nourozi, V., Koike-Akino, T., Mitchell, D., "Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes", IEEE International Conference on Quantum Computing and Engineering (QCE), September 2026.
      BibTeX TR2026-130 PDF
      • @inproceedings{Nourozi2026sep2,
      • author = {Nourozi, Vahid and Koike-Akino, Toshiaki and Mitchell, David},
      • title = {{Reinforcement-Learning-Guided Multi-Branch Decoding of Quantum LDPC Codes}},
      • booktitle = {IEEE International Conference on Quantum Computing and Engineering (QCE)},
      • year = 2026,
      • month = sep,
      • url = {https://www.merl.com/publications/TR2026-130}
      • }
  • MERL Contact:
  • Research Area:

    Signal Processing

Abstract:

Belief propagation (BP) is attractive for quantum lowdensity parity-check (QLDPC) codes, yet short cycles, degeneracy, and trapping configurations can cause oscillation, nonconvergence, or convergence to an incorrect logical class. We propose RL-MBOSD, a reinforcement-learning-guided multi-branch decoder for QLDPC codes. For each code matrix, a separately trained 28- parameter linear action-value policy is shared across its variable nodes and ranks sequential message updates using degreenormalized local and global features. Deterministic score perturbations generate B complementary trajectories, each executed for at most T outer sweeps. Bounded component-wise OSD repairs selected stalled branches; the resulting X- and Z-component lists are paired and ranked using the joint negative log-likelihood under the Pauli channel. Syndrome-valid candidates are then grouped by logical equivalence class and selected using an aggregate logicalclass score. On the [[144, 12, 12]] bivariate-bicycle code and the A5 code, the proposed decoder gives lower error-rate point estimates than the re-plotted prior baselines.