arXiv:2610.00232v1 Announce Type: new Abstract: Fixed-size recurrent memory limits storage growth during inference, but successful recall depends on the task and training. We study a small DeltaNet variant with fixed token-specific key biases, trained to remember 32 new key-value…
arXiv:2512.05990v2 Announce Type: replace Abstract: Memory consolidation determines both what a learner can do now and which changes remain implementable later. We develop a finite-model synthesis of operational state abstraction and optimal control under the stability-evidence-r…
arXiv:2604.13460v2 Announce Type: replace-cross Abstract: A central challenge in continual learning is forgetting: the loss of performance on previously learned tasks after learning new ones. Prior theory has analyzed forgetting under random orderings of fixed task collections in…
arXiv cs.LG
TIER_1English(EN)·Sima Hashemi, Daniel Durstewitz, Georgia Koppe·
arXiv:2609.38356v1 Announce Type: new Abstract: Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continual DSR (cDSR) requires learning new systems while preserving previously learned d…
arXiv:2609.38833v1 Announce Type: new Abstract: Federated continual learning must integrate new tasks over time without losing earlier-task knowledge. Most existing methods attach an anti-forgetting mechanism to the client-trained, server-aggregated loop of federated learning, wh…
arXiv cs.LG
TIER_1English(EN)·Lukas Thede, Shengzhuang Chen, Stefan Winzeck, Matthias Bethge, Zeynep Akata, Jonathan Richard Schwarz·
arXiv:2609.40089v1 Announce Type: new Abstract: Continued pretraining enables language models to adapt to new domains and knowledge, but often at the cost of forgetting previously acquired capabilities. Replay can mitigate this trade-off, but fixed replay mixtures allocate traini…
arXiv cs.AI
TIER_1English(EN)·Fujie Gao, Zuyue Zhang, Gang Sun·
arXiv:2609.38768v1 Announce Type: cross Abstract: Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and compression-like effects during training, but a unified dynamical account of selec…
arXiv:2609.36375v1 Announce Type: new Abstract: Continual learning is usually studied through mechanisms that preserve old knowledge. We develop Successional Learning Theory (SLT), a mesoscopic account in which the current representation is a resident community, the incoming task…
arXiv:2609.37836v1 Announce Type: new Abstract: Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by earlier training history. We study this effect using controlled sequential-training …
Prior work has shown that neural networks exhibit implicit biases toward low-complexity structure (e.g., spectral bias), memorization dynamics, and compression-like effects during training, but a unified dynamical account of selective retention remains incomplete. We propose Repe…
Modern language models are likely to be updated throughout their lifetime rather than trained once and frozen. Each update therefore participates in a recurring cycle: decide which experience to learn from, understand what that update changes, and remain capable of learning from …
arXiv:2609.39367v1 Announce Type: new Abstract: Despite the widespread use of Low-Rank Adaptation (LoRA), little is known about its dynamics in continual learning and the mechanisms by which low-rank updates affect catastrophic forgetting. We provide an asymptotically exact dynam…