Ranked by accuracy on fully fictitious documents. Higher is better; robustness metrics should be read jointly.
Nine-model MemoReason report card. Official rank uses fully fictitious accuracy; metric columns are sortable.
Rank
🥇Rank 1
Claude Sonnet 4.6
92.30%
3.37 pp
1.78×
2.01%
🥈Rank 2
GPT-OSS 120B
90.28%
3.38 pp
1.53×
1.33%
🥉Rank 3
Qwen3.5 27B
88.80%
2.95 pp
1.36×
1.18%
04
GPT-OSS 20B
87.65%
3.10 pp
1.34×
2.18%
05
Qwen3.5 35B-A3B
86.34%
2.74 pp
1.25×
2.33%
06
Gemma-4 26B-A4B-it
85.91%
1.34 pp
1.11×
1.10%
07
Llama-3.1 8B-Instruct
76.88%
1.37 pp
1.06×
3.19%
08
OLMo-3 7B-Think
75.61%
5.31 pp
1.28×
0.98%
09
OLMo-3 7B-Instruct
66.42%
2.83 pp
1.09×
2.53%
Ranking metric: overall accuracy after full entity replacement.
A smaller gap is not necessarily better when absolute accuracy is low.
Shortcut rate uses failed fictitious variant-answer reasoning items as its denominator.
Ranks are point estimates; adjacent differences are not claimed to be statistically significant.
Interactive benchmark example
One document. The same reasoning. A different world.
Choose a world, then follow one pair of values from the document to the preserved constraints and finally
to the recomputed answer.
Showing the factual source.
Step 1 of 3 · Compare the document.The factual share values are 53.5% and 46.5%.
01
Document
What the model reads
Typed names, places, dates, and numbers occupy the same slots in both worlds.
AstraZeneca plc (AZ)
is a
Swedish-British
multinational pharmaceutical and biotechnology company with a portfolio of products in oncology,
cardiovascular, gastrointestinal, infection, neuroscience, respiratory, and inflammation. The company
was founded in
1999
through the merger of
Astra AB
and
Zeneca Group
(itself formed … in
1993).
Zeneca
shareholders received
53.5%
of the shares, while
Astra
shareholders received the remaining
46.5%.
… it has made numerous corporate acquisitions, including
MedImmune
(in
2007)
and
Definiens
(by
MedImmune
in
2014).
AstraZeneca
traces its earliest corporate history to
1913.
Throughout the
twentieth
century, it grew into the largest pharmaceutical company in
Sweden.
…
Astra AB
and
Zeneca PLC
merged
six
years later.
Organization Place Number Time
02
Rules
What must stay true
Replacement values are chosen so the original relationships still hold.
53.5
+
46.5
= 100
✓
share split = 100%
1999
−
1993
=
6✓
merger − formation = stated lag
century(1913)
=
20th✓
founding year = stated century
03
Questions
What is re-evaluated
The prompts stay fixed; answers that depend on replacements are recomputed.
Arithmetic · Variant
By how many percentage points did
Zeneca
shareholders’ share exceed
Astra
shareholders’ share?
Answer53.5 − 46.5 = 7 pp
Inference · Variant
Which acquired company was itself responsible for acquiring
Definiens?
AnswerMedImmune
Extractive · Invariant
Which therapeutic area is listed besides oncology, cardiovascular, gastrointestinal, infection, neuroscience, and respiratory?
Answerinflammation
Extractive · Refusal
Who is the current CEO of
AstraZeneca?
AnswerCannot be determined
The identities and values change, but the executable constraints and question structure do not. Variant
answers are recomputed; invariant and refusal answers remain fixed.
9 / 9
models became less accurate when familiar names and values were replaced
6.5–8.1 pp
drop in overall reasoning accuracy across the nine models
15.7 pp
largest accuracy drop, on questions that require inference
1.0–3.2%
of failures simply repeated the answer from the original real-world version
Table 2, unfolded
The task stays fixed. Accuracy does not.
Every evaluated model is significantly worse on variant-answer reasoning when familiar entities are
replaced by coherent fictitious counterparts.
Main effect · Variant reasoning
Fictitious minus factual accuracy
Negative values mean lower reasoning accuracy after entity replacement.
All 9 models decline, and every 95% confidence interval excludes zero.
Curves are Gaussian KDEs of 100,000 paired document-bootstrap draws per model (shared 0.35 pp bandwidth
and density scale); bars mark the 95% intervals and dots the observed means. 100 documents · 300 paired
factual reasoning questions · 3,000 fictitious instances.
Table 2: Mean performance change (fictional minus factual, percentage points) by answer
and question type. Brackets are 95% CIs; significant cells are framed and bold. Reason.
aggregates arithmetic, temporal, and inference questions. Purple and blue denote decreases and increases,
respectively.
Dataset statistics
MemoReason at a glance.
One balanced, human-curated benchmark: 100 source templates spanning nine themes, fourteen entity types,
and a fixed question–answer contract.
a) Controlled scale
100source documents
→
1,200factual QA templates
→
12,000fictitious instances
Each source supports 12 question templates and 10 paired fictitious worlds.
b) Source themes
9 themes
Award winners12
Biographies10
Places10
Companies12
Natural disasters12
Public attacks10
Retail banking11
Space missions12
Sport events11
c) Annotation density
Entities Rules
Awards58.0 · 57.2
Biographies51.4 · 35.5
Places48.7 · 31.4
Companies41.4 · 21.7
Disasters39.8 · 38.5
Attacks22.3 · 17.4
Banking14.3 · 8.4
Space26.2 · 23.5
Sport39.4 · 44.9
38.0 entities and 31.2 rules per document on average.
d) Question–answer design
4 × 3
Variant
Invariant
Refusal
Extractive
Arithmetic
Temporal
Inference
Every source contains all 12 combinations: four reasoning classes crossed with three answer behaviors.
Figure 2 · Benchmark composition. Document counts and annotation densities are computed over the final
human-reviewed set; bars show theme-level means.
Beyond the endpoints
Mixed worlds are often hardest.
Accuracy often bottoms out before every entity is replaced, suggesting that familiar cues can conflict
with counterfactual relations inside the same context.
91.75%
Qwen3.5-27B factual endpoint
88.09%
minimum accuracy at 80% replacement
88.80%
fully fictitious endpoint
6 / 7
models recover after an intermediate minimum
Figure 3 Accuracy along the replacement path.
All seven curves fall below their factual baseline by 50% replacement; six recover after minima
between 50% and 80%.
Accuracy across partial entity replacement. Bands show 95% document-cluster bootstrap intervals; dashed
lines mark factual baselines. Overlapping intervals do not establish a causal mechanism by themselves.
Method
One source. Ten controlled worlds.
Human annotations identify replaceable entities, encode consistency constraints, and define questions
whose logical structure stays fixed across factual and fictitious settings.
01SourceSelect an entity-rich document
02AnnotateMark typed entities
03ConstrainEncode semantic rules
04ReplaceGenerate coherent worlds
05EvaluateAsk paired questions
Figure 4 · The annotation interface keeps source text, typed entities, executable constraints, and the 12-question
contract visible in one review surface.
Abstract
Memory changes more than recall.
MemoReason pairs factual reasoning tasks with structurally identical fictitious versions in which real
entities are replaced by same-type counterparts under explicit constraints. Across recent LLMs,
contextual reasoning is consistently weaker in the fictitious condition. Yet models seldom answer with
the memorized factual response, suggesting a memory influence more complex than direct factual recall.
Evidence checks
Document-cluster bootstrap inference
200 / 200 observed judge agreement
Temperature robustness on two models
Weak frequency correlation
Reasoning effort does not remove the gap
Citation
BibTeX
@misc{tighidet2026memoreason,
title = {MemoReason: Evaluating the Effect of
Parametric Memory on Contextual Reasoning in LLMs},
author = {Tighidet, Zineddine and Mogini, Andrea and
Mei, Jiali and Gallinari, Patrick and
Piwowarski, Benjamin},
year = {2026}
}