MemoReason

(Memorization in Reasoning)

Evaluating the Effect of Parametric Memory on Contextual Reasoning in LLMs

Keep the task and logic fixed. Change only whether the entities are familiar. Measure what happens to reasoning.

1 BNP Paribas 2 Sorbonne Université 3 Criteo AI Lab

MemoReason leaderboard

Ranked by accuracy on fully fictitious documents. Higher is better; robustness metrics should be read jointly.

Nine-model MemoReason report card. Official rank uses fully fictitious accuracy; metric columns are sortable.
Rank
Rank 1Claude Sonnet 4.692.30%3.37 pp1.78×2.01%
Rank 2GPT-OSS 120B90.28%3.38 pp1.53×1.33%
Rank 3Qwen3.5 27B88.80%2.95 pp1.36×1.18%
04GPT-OSS 20B87.65%3.10 pp1.34×2.18%
05Qwen3.5 35B-A3B86.34%2.74 pp1.25×2.33%
06Gemma-4 26B-A4B-it85.91%1.34 pp1.11×1.10%
07Llama-3.1 8B-Instruct76.88%1.37 pp1.06×3.19%
08OLMo-3 7B-Think75.61%5.31 pp1.28×0.98%
09OLMo-3 7B-Instruct66.42%2.83 pp1.09×2.53%

Ranking metric: overall accuracy after full entity replacement.

A smaller gap is not necessarily better when absolute accuracy is low.

Shortcut rate uses failed fictitious variant-answer reasoning items as its denominator.

Ranks are point estimates; adjacent differences are not claimed to be statistically significant.

One document. The same reasoning. A different world.

Choose a world, then follow one pair of values from the document to the preserved constraints and finally to the recomputed answer.

Showing the factual source.

Step 1 of 3 · Compare the document. The factual share values are 53.5% and 46.5%.

01

Document

What the model reads

Typed names, places, dates, and numbers occupy the same slots in both worlds.

AstraZeneca plc (AZ) is a Swedish-British multinational pharmaceutical and biotechnology company with a portfolio of products in oncology, cardiovascular, gastrointestinal, infection, neuroscience, respiratory, and inflammation. The company was founded in 1999 through the merger of Astra AB and Zeneca Group (itself formed … in 1993). Zeneca shareholders received 53.5% of the shares, while Astra shareholders received the remaining 46.5%. … it has made numerous corporate acquisitions, including MedImmune (in 2007) and Definiens (by MedImmune in 2014). AstraZeneca traces its earliest corporate history to 1913. Throughout the twentieth century, it grew into the largest pharmaceutical company in Sweden. … Astra AB and Zeneca PLC merged six years later.

Organization Place Number Time
02

Rules

What must stay true

Replacement values are chosen so the original relationships still hold.

53.5 + 46.5 = 100 ✓ share split = 100%
1999 − 1993 = 6 ✓ merger − formation = stated lag
century(1913) = 20th ✓ founding year = stated century
03

Questions

What is re-evaluated

The prompts stay fixed; answers that depend on replacements are recomputed.

Arithmetic · Variant

By how many percentage points did Zeneca shareholders’ share exceed Astra shareholders’ share?

Answer 53.5 − 46.5 = 7 pp
Inference · Variant

Which acquired company was itself responsible for acquiring Definiens?

Answer MedImmune
Extractive · Invariant

Which therapeutic area is listed besides oncology, cardiovascular, gastrointestinal, infection, neuroscience, and respiratory?

Answerinflammation
Extractive · Refusal

Who is the current CEO of AstraZeneca?

AnswerCannot be determined

The identities and values change, but the executable constraints and question structure do not. Variant answers are recomputed; invariant and refusal answers remain fixed.

9 / 9
models became less accurate when familiar names and values were replaced
6.5–8.1 pp
drop in overall reasoning accuracy across the nine models
15.7 pp
largest accuracy drop, on questions that require inference
1.0–3.2%
of failures simply repeated the answer from the original real-world version

The task stays fixed.
Accuracy does not.

Every evaluated model is significantly worse on variant-answer reasoning when familiar entities are replaced by coherent fictitious counterparts.

Main effect · Variant reasoning

Fictitious minus factual accuracy

Negative values mean lower reasoning accuracy after entity replacement.

Bootstrap density 95% CI Mean
Smoothed paired document-bootstrap distributions of variant-answer reasoning performance change.
Model Accuracy change (percentage points) 95% CI Mean Δ
OLMo-3 7B-Think [−12.4, −3.8] −8.1 pp
OLMo-3 7B-Instruct [−12.6, −2.9] −7.8 pp
GPT-OSS 20B [−10.2, −3.8] −7.0 pp
Gemma-4 26B-A4B-it [−10.2, −2.9] −6.5 pp
GPT-OSS 120B [−11.3, −4.7] −8.0 pp
Qwen3.5 27B [−10.1, −3.1] −6.6 pp
Qwen3.5 35B-A3B [−9.9, −3.2] −6.5 pp
Llama-3.1 8B-Instruct [−10.4, −2.8] −6.6 pp
Claude Sonnet 4.6 [−10.8, −5.0] −7.8 pp
All 9 models decline, and every 95% confidence interval excludes zero. Curves are Gaussian KDEs of 100,000 paired document-bootstrap draws per model (shared 0.35 pp bandwidth and density scale); bars mark the 95% intervals and dots the observed means. 100 documents · 300 paired factual reasoning questions · 3,000 fictitious instances.
Table 2: Mean performance change (fictional minus factual, percentage points) by answer and question type. Brackets are 95% CIs; significant cells are framed and bold. Reason. aggregates arithmetic, temporal, and inference questions. Purple and blue denote decreases and increases, respectively.

MemoReason at a glance.

One balanced, human-curated benchmark: 100 source templates spanning nine themes, fourteen entity types, and a fixed question–answer contract.

a) Controlled scale

100source
documents
1,200factual
QA templates
12,000fictitious
instances

Each source supports 12 question templates and 10 paired fictitious worlds.

b) Source themes

  • Award winners12
  • Biographies10
  • Places10
  • Companies12
  • Natural disasters12
  • Public attacks10
  • Retail banking11
  • Space missions12
  • Sport events11

c) Annotation density

Entities Rules
Awards58.0 · 57.2
Biographies51.4 · 35.5
Places48.7 · 31.4
Companies41.4 · 21.7
Disasters39.8 · 38.5
Attacks22.3 · 17.4
Banking14.3 · 8.4
Space26.2 · 23.5
Sport39.4 · 44.9

38.0 entities and 31.2 rules per document on average.

d) Question–answer design

4 × 3
Variant
Invariant
Refusal
Extractive
Arithmetic
Temporal
Inference

Every source contains all 12 combinations: four reasoning classes crossed with three answer behaviors.

Figure 2 · Benchmark composition. Document counts and annotation densities are computed over the final human-reviewed set; bars show theme-level means.

Mixed worlds are often hardest.

Accuracy often bottoms out before every entity is replaced, suggesting that familiar cues can conflict with counterfactual relations inside the same context.

91.75%
Qwen3.5-27B factual endpoint
88.09%
minimum accuracy at 80% replacement
88.80%
fully fictitious endpoint
6 / 7
models recover after an intermediate minimum

Figure 3 Accuracy along the replacement path.

All seven curves fall below their factual baseline by 50% replacement; six recover after minima between 50% and 80%.

Seven model accuracy curves across entity replacement proportions from zero to one hundred percent
Accuracy across partial entity replacement. Bands show 95% document-cluster bootstrap intervals; dashed lines mark factual baselines. Overlapping intervals do not establish a causal mechanism by themselves.

One source. Ten controlled worlds.

Human annotations identify replaceable entities, encode consistency constraints, and define questions whose logical structure stays fixed across factual and fictitious settings.

  1. 01SourceSelect an entity-rich document
  2. 02AnnotateMark typed entities
  3. 03ConstrainEncode semantic rules
  4. 04ReplaceGenerate coherent worlds
  5. 05EvaluateAsk paired questions
MemoReason annotation interface showing a source document, entity annotations, rules, and questions
Figure 4 · The annotation interface keeps source text, typed entities, executable constraints, and the 12-question contract visible in one review surface.

Memory changes more than recall.

MemoReason pairs factual reasoning tasks with structurally identical fictitious versions in which real entities are replaced by same-type counterparts under explicit constraints. Across recent LLMs, contextual reasoning is consistently weaker in the fictitious condition. Yet models seldom answer with the memorized factual response, suggesting a memory influence more complex than direct factual recall.

Evidence checks

  • Document-cluster bootstrap inference
  • 200 / 200 observed judge agreement
  • Temperature robustness on two models
  • Weak frequency correlation
  • Reasoning effort does not remove the gap

BibTeX

@misc{tighidet2026memoreason,
  title  = {MemoReason: Evaluating the Effect of
            Parametric Memory on Contextual Reasoning in LLMs},
  author = {Tighidet, Zineddine and Mogini, Andrea and
            Mei, Jiali and Gallinari, Patrick and
            Piwowarski, Benjamin},
  year   = {2026}
}

Accuracy by replacement proportion

Seven model accuracy curves across entity replacement proportions from zero to one hundred percent

Shaded bands are 95% document-cluster bootstrap intervals.