Harny / Benchmarks

77.6 on BEAM. Memory. With receipts.

The best published score at the 128K tier.

BEAM asks what a memory system still knows 128,000 tokens into one conversation: what changed, in what order, and what was never said.

Next-highest published at this tier: 76.9.
We got there with budget-tier models.

Scored on BEAM’s own protocol
One response per question · temperature 0 · all 400 questions

Three benchmarks.
Different tests of memory.

BEAM128K tier · a long conversation
77.6
LongMemEvalAcross earlier sessions
96.9
LoCoMoWeeks of conversation
99.3Standard
88.3With traps
Different tests. Different scoring rules.
See how each score works ↓

Memory is part of the stack. Harny brings connected accounts, per-user context, and agent-reported records into one integration.

What this means
for your product.

These runs used budget-tier models. Here’s why the tasks matter.

96.9 · LONGMEMEVAL

Don’t make them explain it twice.

The order number. The plan. What they tried last week. LongMemEval tests recall across earlier sessions: the memory a support agent needs to pick up the conversation.

88.3 · LOCOMO WITH TRAPS

A confident guess is still a guess.

A personal agent needs to know when the answer isn’t there. LoCoMo’s trap questions test whether it invents one anyway.

77.6 · BEAM 128K

Keep up when life changes.

People move cities, switch teams, and change their minds. BEAM tests whether memory keeps track of what changed and when.

Where we sit among
published results.

BEAM · 100K / 128K tier
Comparison snapshot · 6 September 2026.

Published BEAM scores at the 100K/128K tier, with each publisher’s reported setup. Models and scoring rules vary. Rows are ordered by how much methodology was disclosed, not by score.
SystemFigureTierAnswererJudgeAggregation rule stated
Harny77.6128KBudget-tier, routedgemini-3.5-flash-liteYes
LIGHT — the BEAM paper’s own method35.8100KPer the paperPer the paperYes
RAG baseline32.3100KPer the paperPer the paperYes
SelfMem50.4100KGPT-5.4-nanoGPT-5.4-miniYes
Exabase M-176.9100KGemini 3 FlashGemini 3 FlashNo
Mnemosyne v3.0.065.2100KLlama 3.3 70BDeepSeek V4 FlashNo
Hindsight73.4100KNot statedNot statedNo
Honcho63.0100KNot statedNot statedNo
Comparison scope

Eywa reports 81.45 without a tier label in the source reviewed for this comparison. We have therefore excluded it from this tier’s ranking.

BEAM’s published “100K” tier is the 128K tier. The 500K, 1M and 10M tiers are separate boards; their figures are not comparable to these.

Weeks of conversation.
Two ways to lose the user.

Memory fails when it forgets what the user told you, or invents something they didn’t. Different failures deserve different numbers.

99.3

Standard questions

Four standard question categories, in their natural mix. This score focuses on recalling what the user actually said.

88.3

Including trap questions

All five categories, weighted equally. Questions designed to provoke invented answers account for one fifth of the score.

Both scores use a corrected answer key. We identified errors in 99 of LoCoMo’s 1,540 published answers. Ask us for the corrections and the evidence behind each one. Request the corrections ↗︎

How we measure.

The setup matters as much as the score. Here are the samples, response counts, and judges behind ours. Run artifacts are available on request.

BEAM

Benchmark’s own protocol · 400 questions · one response each · judge gemini-3.5-flash-lite

LongMemEval

150 questions · three responses each · judge deepseek-v4-pro

LoCoMo · four standard categories

Natural mix · 150 questions · corrected key · judge deepseek-v4-pro

LoCoMo · all five categories

Equal weight · 150 questions · five responses each · corrected key · judge deepseek-v4-pro

Turn your agent on.

Bring us your use case. Let’s work out the integration.

Get early access