Sapience: A Hybrid Architecture for Long-Context Reasoning and Continual Learning | Zenodo
Skip to main
You are using an outdated browser. Please upgrade your browser to improve your experience.
Published August 14, 2026
| Version v4.3
Preprint
Open
Sapience: A Hybrid Architecture for Long-Context Reasoning and Continual Learning
Authors/Creators
Sapience Labs
Contributors
Contact person:
Zahn, Oliver
Description
Frontier models approach long-context reasoning by scaling the transformer context window: one of the several memory systems the brain evolved, scaled alone. We present Sapience, a hybrid architecture that builds the missing ones: a substitutable transformer cortex coupled to hippocampal episodic retrieval (L1), cortical-replay consolidation (L2), and sleep-stage weight adaptation (L3) — the complementary systems that Complementary Learning Systems theory (McClelland et al., 1995; Kumaran et al., 2016) says a complete brain pairs with a cortex. Four results carry the claim. (1) The architecture carries the long-context selection and scaling load; the reader sets much of the absolute accuracy level: holding reader weights fixed, a frontier reader answers 3-hop questions at 34.4% from a raw ∼959K-token window and at 66.7% from ∼700 tokens retrieved by the store (n = 90 item-paired; McNemar p = 1.5 × 10−5 ; judge-invariant across five judges from three vendors). (2) The system operates at 10M-token history scale where raw-window configurations fail: on one continuous benchmark the reader’s working set stays bounded while history grows 2,500-fold (flat ∼82% at 2M–10M, 3 seeds); on a paired forward-growth construction the same reader scores 0/20 with the answer verified inside its delivered window while Sapience answers 45.0% on identical items (b = 9, c = 0, p = 3.9 × 10−3 ), at over 1000× lower cost per query; and on knowledge that changes, supersession structure beats relevance retrieval by +40.4pp on real Wikipedia revision histories — retrieval saw both values and could not tell which was current. (3) The consolidation tier has one measured system-level function, and a broader one still open: the write-time aggregation gate materializes exact distinct-counts at write time and beats a compute-matched LLM counter emitting the identical sentence by +22.8pp (95% CI [+8.1, +38.2]) on covered aggregation questions, attributable to deterministic count correctness; the same contrast runs −3.3pp against the gate on trend questions and is exactly neutral on single-episode controls; broader semantic and schema consolidation is implemented but unestablished. (4) Replay drawn from the same store makes adaptation durable: in an internally pre-registered matched-acquisition design (L3), replay generated from the episodic store cuts catastrophic forgetting by 47.0pp (95% CI [39.7, 54.0]; acquisition formally non-inferior), generalizes across model families, domain pairs, and a real git-history corpus (75.5/51.8/+52.9pp), outperforms a tuned regularization baseline, and survives dose- and step-matched ablations that isolate the replayed content, consistent with the CLS prediction that interleaved replay reduces interference. Under genuine conflicting updates, generic replay wins retention and composite current-world accuracy while paying a measured stale-value tax; restricting replay to the memories the store marks current removed every observed stale response (0/81 superseded probes; 95% upper bound 12.5% at n = 27) at non-inferior retention and the programme’s best composite current-world accuracy (90.7%); a volume-matched exclusion control does not reproduce the effect (13.6% tax, −7.9pp retention), so it follows from which content is excluded, not from replaying less. The architecture also predicts where memory should not help (injection where the cortex suffices costs accuracy; within the window a strong reader ties retrieval), and both predictions are confirmed by measurement. Every claim is reported with its scope: the primary mechanism, transfer, real-corpus, replay-format and conflict instruments are ≤8B adaptation instruments, while a fixed-load scale ladder and a three-seed learning-rate-matched experiment extend measurement to 72B under a stated protocol asymmetry and a load-to-capacity confound; per-layer accounting separating measured positives from open bets, and an adversarial audit protocol behind every headline number (Table 10). Throughout, cortex names the architectural slot holding a standard LLM and reader that same LLM in its answering role; in every matched comparison both arms use the identical reader, and only what it reads changes.
Files
sapience-architecture-v4.3.pdf
Files<br>(3.1 MB)
Name<br>Size
Download all
sapience-architecture-v4.3.pdf
md5:f2e4c67b324820e880f27eece1dfe583
3.1 MB
Preview
Download
Additional details
Dates
Issued
2026-08-14
38
Views
Downloads
Show more details
All...