MemoryATHENA: Adaptive Routing over Latent and Generated Memories

AuthorsMingyuan Li, Guangsheng Yu, Juyuan Zhang, Xu Wang, Zhibo Man, Haonan Zhang, Shaoxiong Ji
VenuePreprint
DateSeptember 2026
Read Paper Code Models

Abstract

Learned-memory methods store information in an explicit table and consume it through a separate reader, allowing addressing, storage, and reading to be modified independently. Prior work on cross-model memory transfer exploits this separation to reuse learned memory across frozen backbones through an adapted reader, while the representation consumed by the model still originates from stored memory. This raises a natural question: must useful memory always be retrieved from storage, or can it also be generated? We investigate this question with MemoryAthena, a memory interface with three pathways: direct Engram retrieval (E), generation from retrieved Engram cues (GE), and generation from causal backbone states without consulting the memory table (GH). Generated memory is not uniformly better than direct retrieval: it can complement E in one context but interfere with it in another. MemoryAthena therefore treats E as an explicit anchor and learns when a generated representation should intervene. With the backbone, memory, generators, and readers frozen, a lightweight causal routing head is trained from counterfactual future-token likelihood advantages of GE and GH relative to E. At inference time, an admitted candidate modifies the E residual through bounded interpolation, while rejection recovers the direct pathway exactly. On question answering, MemoryAthena raises the five-task average from 37.65 to 39.28 over the direct pathway of the same checkpoint, while the six-task general-NLP average increases from 76.73 to 79.13. The complete memory-side system contains approximately 201M parameters, excluding the frozen backbone. Further analyses show that the utility of E, GE, and GH varies across tasks and inputs, while gold-label oracles reveal additional complementarity among the three pathways. These results support generated memory as a selective correction to direct retrieval rather than a universal replacement, and highlight routing when, which, and how strongly to intervene as the central challenge.

MemoryATHENA overview showing prior memory decomposition, three memory pathways, the E-anchored router, and training and inference
Figure 1: Overview of MemoryATHENA. Direct Engram retrieval (E) remains the reference pathway while generated-memory candidates (GE and GH) are admitted conditionally by an E-anchored router. [Download vector graphics]

Direct and generated memory pathways

MemoryATHENA keeps direct Engram retrieval as the reference while adding two generated-memory pathways. The generators construct target-side memory proposals from different causal sources, and the router learns when either proposal should modify the E pathway.

E

Direct Engram retrieval

The stable reference pathway. Retrieved memory is projected into the target backbone without requiring a generated candidate.

GE

Generated from Engram cues

A generator conditions on retrieved Engram representations to produce a target-side memory proposal.

GH

Generated from backbone states

A generator conditions on clean causal hidden states, providing a complementary proposal that does not directly consult the memory table.

Bounded admission + exact fallback.

A proposal is admitted only when its predicted advantage and confidence pass the configured rule; otherwise inference returns exactly to the direct E reference.

Experiments organized around three questions

Results across QA and general NLP

Headline aggregates from the current paper-facing audit. Values are percentages and use one training seed (42). The tables keep protocol caveats visible instead of collapsing unlike baselines into one claim.

39.28Five-task QA average · paper configuration
79.13Six-task NLP average · Yahoo τ=1 display
3 pathsE reference plus GE/GH generated candidates

Five-task open-domain QA

Open-domain QA cells are EM/F1; TruthfulQA cells are MC1/MC2/MC3/mean.

SettingNQWebQATriviaQATruthfulQAHotpotQA
Engram-only20.28/28.1814.86/33.2863.92/69.1827.42/44.32/22.69/31.4717.95/25.92
E-only path20.20/28.2814.96/33.3564.05/69.3426.93/44.18/22.64/31.2518.19/26.04
Three-source router22.72/33.0217.86/34.6062.78/70.6828.15/44.04/23.02/31.7415.62/26.34
Mistral → Llama transfer20.37/29.9818.06/36.4060.35/67.3927.78/41.57/21.87/30.4116.00/25.17

The standalone Engram-only row and joint-checkpoint rows have different training histories; the transfer row is reported separately.

Six-task general NLP

Accuracy (%). E/router rows use the next-token synonym-sum dCPMI protocol; the historical Vanilla row uses a different full-choice scorer.

MethodSST2MRCRRTAGNYahooAverage
Vanilla Mistral · historical scorer81.0875.6074.0074.6773.2455.0372.27
Engram-only84.1781.0082.4082.3672.9357.5176.73
Three-source router · Yahoo τ=1.088.0784.7084.1083.8676.6457.4379.13

Yahoo τ=1.0 is a test-set threshold sweep; the selected threshold is shown explicitly for transparency.

What the experiments show

Read, reproduce, inspect

The public release separates code, model artifacts, and compact experiment metadata. Large scratch paths and raw task data are intentionally not part of the web page.

BibTeX

@misc{li2026memoryathenaadaptiveroutinglatent,
  title = {MemoryAthena: Adaptive Routing over Latent and Generated Memories},
  author = {Mingyuan Li and Guangsheng Yu and Juyuan Zhang and Xu Wang and Zhibo Man and Haonan Zhang and Shaoxiong Ji},
  year = {2026},
  eprint = {2609.25853},
  archivePrefix = {arXiv},
  primaryClass = {cs.CL},
  url = {https://arxiv.org/abs/2609.25853}
}