AI memory performance proven to be top-tier

Memory.Inc Achieves 94.8% on LongMemEval-S

Memory.Inc scored 94.8% on LongMemEval-S, a key AI memory benchmark.
This pre-launch result, recorded via GPT-5.5 output and GPT-4o internal processing, is state-of-the-art (SOTA) compared to top public AI memory systems.

LongMemEval-S evaluates how accurately AI recalls key details in long chats and answers using context scattered across multi-turn sessions.

Memory isn't about finding answers in one doc like standard search. It requires synthesising past user input, assistant responses, updated facts, user habits, timelines, and relative dates across multi-session data.

Simply put, it tests the core memory capacity an AI needs to chat with users long-term, not just if it can "match similar sentences."

Type

Memory.Inc

Mastra OM

Supermemory

Zep

Full Context

single-session-user

95.7%

98.6%

97.1%

92.9%

81.4%

single-session-assistant

100.0%

82.1%

96.4%

80.4%

94.6%

single-session-preference

96.7%

73.3%

70.0%

56.7%

20.0%

knowledge-update

97.4%

85.9%

88.5%

83.3%

78.2%

temporal-reasoning

95.5%

85.7%

76.7%

62.4%

45.1%

multi-session

83.5%

79.7%

71.4%

57.9%

44.3%

Total

94.8%

84.23%

81.6%

71.2%

60.2%

  • Scroll the table left/right to view all content.

  • Supermemory's public 95% score is based on Recall@15 with aggregation (combining multiple search results). The table above uses standard LongMemEval-S QA accuracy for accurate comparison.

  • MemKraft, MemPalace, etc., are omitted as they use subset scores or different evaluation metrics instead of the full 500-question LongMemEval-S.


The hard part of AI memory isn't "storing a lot of data."

It is recalling accurately.
And, more importantly, keeping changing information up to date.

If a user says A first, then changes to B later, the AI shouldn't stick to A. It must answer based on the latest update, B.
Even as budgets, tastes, schedules, and projects shift, the AI must keep pace with those changes.

Memory.Inc is not just a chat storage; it is an AI memory system built to keep user and team context current.

After launch, we will open-source our evaluation code so anyone can run and verify the benchmark.

Memory.Inc is building memory infra for AI to remember more accurately and retain context longer.