Library · 11-persistent-memory-continual-adaptation

Persistent memory

TitlePeerLink
It's All Connected (Miras)◦ preprintarxiv.org/abs/2504.13173
Atlas: Optimally Memorize the Context at Test Time◦ preprintarxiv.org/abs/2505.23735
Gated Delta Networks (Mamba2 + delta rule)◦ preprintarxiv.org/abs/2412.06464
Test-Time Training with Next-Token Prediction (TTT-NTP)◦ preprintarxiv.org/abs/2606.21803
Lattice: Learning to Efficiently Compress the Memory◦ preprintarxiv.org/abs/2504.05646
SR-TTT: Surprisal-Aware Residual Test-Time Training◦ preprintarxiv.org/abs/2603.06642
Selective Memory: Write-Time Gating w/ Hierarchical Archiving◦ preprintarxiv.org/abs/2603.15994
Loss of Plasticity in Deep Continual Learning✓ peernature.com/articles/s41586-024-07711-7
loss-of-plasticity (official repo, CBP + benchmarks)github.com/shibhansh/loss-of-plasticity
AdaLin: Preserving Plasticity via Adaptive Linearity Injection◦ preprintarxiv.org/abs/2505.09486
L2 Init: Maintaining Plasticity via Regenerative Regularization◦ preprintarxiv.org/abs/2308.11958
Plasticity Loss in Deep RL: A Survey◦ preprintarxiv.org/abs/2411.04832
Plasticine: Plasticity-Motivated Deep RL benchmark◦ preprintarxiv.org/abs/2504.17490
AltNet: Plasticity-Stability Dilemma in RL— unrefarxiv.org/abs/2512.01034
Mem0: Scalable Long-Term Memory for AI Agents◦ preprintarxiv.org/abs/2504.19413
A-MEM: Agentic Memory for LLM Agents◦ preprintarxiv.org/abs/2502.12110
LoCoMo: Evaluating Very-Long-Term Conversational Memoryaclanthology.org/2024.acl-long.747/
HebbGate: Local Reward-Modulated Gating for Continual Learning
Computational principles of synaptic memory consolidation (Benna & Fusi)✓ peerdoi:10.1038/nn.4401
Enabling Continual Learning with Differentiable Hebbian Plasticity (Thangarasa, Miconi)— unrefarxiv.org/abs/2006.16558
AGMP — Astrocyte-Gated Multi-Timescale Plasticity✓ peerdoi.org/10.3389/fnins.2025.1768235
Hippocampus supports multi-task RL under partial observability✓ peernature.com/articles/s41467-025-64591-9
MESU — Bayesian continual learning & forgetting (Bonnet et al.)◦ preprintarxiv.org/abs/2504.13569
Hybrid NNs for continual learning inspired by corticohippocampal circuits
Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention (Katharopoulos et al., ICML 2020)✓ peerarxiv.org/abs/2006.16236
Linear Transformers Are Secretly Fast Weight Programmers (Schlag, Irie, Schmidhuber, ICML 2021)✓ peerarxiv.org/abs/2102.11174
Learning to Control Fast-Weight Memories (Schmidhuber)✓ peerdoi:10.1162/neco.1992.4.1.131
Parallelizing Linear Transformers with the Delta Rule over Sequence Length — DeltaNet (Yang et al., NeurIPS 2024)◦ preprintarxiv.org/abs/2406.06484
Mean teachers are better role models (Tarvainen & Valpola, NeurIPS 2017)✓ peerarxiv.org/abs/1703.01780
Memory-R1: Enhancing LLM Agents to Manage & Utilize Memories via RL (Yan et al.)◦ preprintarxiv.org/abs/2508.19828
DeltaMem: Towards Agentic Memory Management via RL (Zhang et al.)◦ preprintarxiv.org/abs/2604.01560
Just-In-Time RL: Continual Learning in LLM Agents Without Gradient Updates (Li et al.)◦ preprintarxiv.org/abs/2601.18510
Reinforced Fast Weights with Next-Sequence Prediction (REFINE)◦ preprintarxiv.org/abs/2602.16704
Reliability-Adjusted Prioritized Experience Replay (ReaPER)◦ preprintarxiv.org/abs/2506.18482
Governing Evolving Memory in LLM Agents (SSGM Framework)◦ preprintarxiv.org/abs/2603.11768
AURA: Action-Gated Memory for Robot Policies at Constant VRAM (Chen et al.)◦ preprintarxiv.org/abs/2606.02775
Gated DeltaNet-2: Decoupling Erase and Write in Linear Attention (Hatamizadeh et al.)◦ preprintarxiv.org/abs/2605.22791
Worth Remembering: Surprise-Gated Robot Episodic Memory (Gorlo et al.)◦ preprintarxiv.org/abs/2606.03787
Context-Gated Associative Retrieval: From Theory to Transformers (Choraria et al.)◦ preprintarxiv.org/abs/2605.10970
Short-Term Synaptic Plasticity Stabilizes Goal-Conditioned Dynamics (PFC reservoir)◦ preprintarxiv.org/abs/2606.03481
Continual Learning Through Synaptic Intelligence◦ preprintarxiv.org/abs/1703.04200
Gradient Episodic Memory for Continual Learning (GEM)✓ peerarxiv.org/abs/1706.08840
Efficient Lifelong Learning with A-GEM◦ preprintarxiv.org/abs/1812.00420
Prioritized Experience Replay (PER)◦ preprintarxiv.org/abs/1511.05952
Continual lifelong learning with neural networks: a review✓ peerdoi.org/10.1016/j.neunet.2019.01.012
Overcoming catastrophic forgetting in neural networks (EWC)✓ peerdoi.org/10.1073/pnas.1611835114
Catastrophic Interference in Connectionist Networks✓ peerdoi.org/10.1016/S0079-7421(08)60536-8
Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training◦ preprintarxiv.org/abs/2607.00368
Can Scale Save Us From Plasticity Loss in Large Language Models?◦ preprintarxiv.org/abs/2606.24752
Fast and Slow Variational Continual Learning◦ preprintarxiv.org/abs/2606.24007
Can a Language Model Learn Facts Continually in Its Weights?◦ preprintarxiv.org/abs/2607.11020
Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents◦ preprintarxiv.org/abs/2607.22157
No Time Like the Present: Agentic Test-Time Training for LLM Agents◦ preprintarxiv.org/abs/2607.03441
Is Our Benchmark Enough? An Analysis of Continual Learning for MLLMs◦ preprintarxiv.org/abs/2606.20961
Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments◦ preprintarxiv.org/abs/2606.05661
LatentGym: A Testbed For Cross-Task Experiential Learning With Controllable Latent Structure◦ preprintarxiv.org/abs/2606.15306
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents◦ preprintarxiv.org/abs/2606.04703
A Decision-Theoretic View of Test-Time Training: When, How Far, and Which Directions to Adapt◦ preprintarxiv.org/abs/2606.15569
Scaling Self-Evolving Agents via Parametric Memory (TMEM)◦ preprintarxiv.org/abs/2606.04536
Retrievable Gradients: Continual Post-Training Without Cumulative Weight Drift◦ preprintarxiv.org/abs/2606.15734
The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoningarxiv.org/abs/2607.20952