Multi-hop reasoning collapses by 1M tokens; with a learning system it doesn't

lvrzhn1 pts0 comments

Sapience: A Hybrid Architecture for Long-Context Reasoning and Continual Learning | Zenodo

Skip to main

You are using an outdated browser. Please upgrade your browser to improve your experience.

Published August 14, 2026

| Version v4.3

Preprint

Open

Sapience: A Hybrid Architecture for Long-Context Reasoning and Continual Learning

Authors/Creators

Sapience Labs

Contributors

Contact person:

Zahn, Oliver

Description

Frontier models approach long-context reasoning by scaling the transformer context window: one of the several memory systems the brain evolved, scaled alone. We present Sapience, a hybrid architecture that builds the missing ones: a substitutable transformer cortex coupled to hippocampal episodic retrieval (L1), cortical-replay consolidation (L2), and sleep-stage weight adaptation (L3) — the complementary systems that Complementary Learning Systems theory (McClelland et al., 1995; Kumaran et al., 2016) says a complete brain pairs with a cortex. Four results carry the claim. (1) The architecture carries the long-context selection and scaling load; the reader sets much of the absolute accuracy level: holding reader weights fixed, a frontier reader answers 3-hop questions at 34.4% from a raw ∼959K-token window and at 66.7% from ∼700 tokens retrieved by the store (n = 90 item-paired; McNemar p = 1.5 × 10−5 ; judge-invariant across five judges from three vendors). (2) The system operates at 10M-token history scale where raw-window configurations fail: on one continuous benchmark the reader’s working set stays bounded while history grows 2,500-fold (flat ∼82% at 2M–10M, 3 seeds); on a paired forward-growth construction the same reader scores 0/20 with the answer verified inside its delivered window while Sapience answers 45.0% on identical items (b = 9, c = 0, p = 3.9 × 10−3 ), at over 1000× lower cost per query; and on knowledge that changes, supersession structure beats relevance retrieval by +40.4pp on real Wikipedia revision histories — retrieval saw both values and could not tell which was current. (3) The consolidation tier has one measured system-level function, and a broader one still open: the write-time aggregation gate materializes exact distinct-counts at write time and beats a compute-matched LLM counter emitting the identical sentence by +22.8pp (95% CI [+8.1, +38.2]) on covered aggregation questions, attributable to deterministic count correctness; the same contrast runs −3.3pp against the gate on trend questions and is exactly neutral on single-episode controls; broader semantic and schema consolidation is implemented but unestablished. (4) Replay drawn from the same store makes adaptation durable: in an internally pre-registered matched-acquisition design (L3), replay generated from the episodic store cuts catastrophic forgetting by 47.0pp (95% CI [39.7, 54.0]; acquisition formally non-inferior), generalizes across model families, domain pairs, and a real git-history corpus (75.5/51.8/+52.9pp), outperforms a tuned regularization baseline, and survives dose- and step-matched ablations that isolate the replayed content, consistent with the CLS prediction that interleaved replay reduces interference. Under genuine conflicting updates, generic replay wins retention and composite current-world accuracy while paying a measured stale-value tax; restricting replay to the memories the store marks current removed every observed stale response (0/81 superseded probes; 95% upper bound 12.5% at n = 27) at non-inferior retention and the programme’s best composite current-world accuracy (90.7%); a volume-matched exclusion control does not reproduce the effect (13.6% tax, −7.9pp retention), so it follows from which content is excluded, not from replaying less. The architecture also predicts where memory should not help (injection where the cortex suffices costs accuracy; within the window a strong reader ties retrieval), and both predictions are confirmed by measurement. Every claim is reported with its scope: the primary mechanism, transfer, real-corpus, replay-format and conflict instruments are ≤8B adaptation instruments, while a fixed-load scale ladder and a three-seed learning-rate-matched experiment extend measurement to 72B under a stated protocol asymmetry and a load-to-capacity confound; per-layer accounting separating measured positives from open bets, and an adversarial audit protocol behind every headline number (Table 10). Throughout, cortex names the architectural slot holding a standard LLM and reader that same LLM in its answering role; in every matched comparison both arms use the identical reader, and only what it reads changes.

Files

sapience-architecture-v4.3.pdf

Files<br>(3.1 MB)

Name<br>Size

Download all

sapience-architecture-v4.3.pdf

md5:f2e4c67b324820e880f27eece1dfe583

3.1 MB

Preview

Download

Additional details

Dates

Issued

2026-08-14

38

Views

Downloads

Show more details

All...

reader from sapience architecture replay matched

Related Articles