Why LongMemEval Scores Are Not Comparable Across Memory Products
Same-stack LongMemEval-500 with a local 7B reader is a weak-baseline floor (naive RAG 0.252). Published Mem0 and Graphiti numbers use different harnesses and models. A sealed qwen3.8:27b n=500 run scored 0.626 versus 0.702 on a gpt-4o-mini stack.
Leaderboards hide the harness
LongMemEval is often quoted as if it were a single number. It is not. Published competitor scores (for example Mem0 around 0.94, Graphiti around 0.64) use their own models, judges, and packing. They are not apples-to-apples with a local 7B run.
On our locked same-stack protocol (local qwen2.5-coder:7B, nomic-embed, per-item isolation, n=500), naive RAG scored 0.252 and an LLM reranker scored 0.240 at about 13 times the latency. That floor is the honest baseline for that stack, not a claim that competitors are worse.
What the Qwen 3.8 27B 500-run actually showed
A sealed n=500 reader swap on the R230 stack replaced gpt-4o-mini with local qwen3.8:27b (think off) and scored 0.626 versus 0.702. The gap was mostly temporal and preference abstention when the reasoning channel was off, not a hidden product win. Think-high pilots are research, not a shipping SLA.
MemStrata’s published papers argue a different axis: temporal validity (stale-fact error near 0 on marker-free supersession tasks), not “highest LME-500 with any model.”
How to read competitor grids
Same-stack grids in the Paper 5 draft show MemStrata at 1.000 on marker-free code mutation, MemArch supersession, poisoning, and TEMPO as-of-T, while flat RAG and several agent-memory systems still serve stale values. Cue-rich synthetic exports can be “read” from raw passages and should not be over-interpreted.
Paper 5 is still a draft. Until it is on arXiv, treat LME-500 numbers as evaluation notes, not marketing medians.