Harny / Benchmarks
77.6 on BEAM. Memory. With receipts.
The best published score at the 128K tier.
BEAM asks what a memory system still knows 128,000 tokens into one conversation: what changed, in what order, and what was never said.
Next-highest published at this tier: 76.9.
We got there with budget-tier models.
One response per question · temperature 0 · all 400 questions
Three benchmarks.
Different tests of memory.
See how each score works ↓
Memory is part of the stack. Harny brings connected accounts, per-user context, and agent-reported records into one integration.
What this means
for your product.
These runs used budget-tier models. Here’s why the tasks matter.
Where we sit among
published results.
BEAM · 100K / 128K tier
Comparison snapshot · 6 September 2026.
| System | Figure | Tier | Answerer | Judge | Aggregation rule stated |
|---|---|---|---|---|---|
| Harny | 77.6 | 128K | Budget-tier, routed | gemini-3.5-flash-lite | Yes |
| LIGHT — the BEAM paper’s own method | 35.8 | 100K | Per the paper | Per the paper | Yes |
| RAG baseline | 32.3 | 100K | Per the paper | Per the paper | Yes |
| SelfMem | 50.4 | 100K | GPT-5.4-nano | GPT-5.4-mini | Yes |
| Exabase M-1 | 76.9 | 100K | Gemini 3 Flash | Gemini 3 Flash | No |
| Mnemosyne v3.0.0 | 65.2 | 100K | Llama 3.3 70B | DeepSeek V4 Flash | No |
| Hindsight | 73.4 | 100K | Not stated | Not stated | No |
| Honcho | 63.0 | 100K | Not stated | Not stated | No |
Comparison scope
Eywa reports 81.45 without a tier label in the source reviewed for this comparison. We have therefore excluded it from this tier’s ranking.
BEAM’s published “100K” tier is the 128K tier. The 500K, 1M and 10M tiers are separate boards; their figures are not comparable to these.
Weeks of conversation.
Two ways to lose the user.
Memory fails when it forgets what the user told you, or invents something they didn’t. Different failures deserve different numbers.
Standard questions
Four standard question categories, in their natural mix. This score focuses on recalling what the user actually said.
Including trap questions
All five categories, weighted equally. Questions designed to provoke invented answers account for one fifth of the score.
Both scores use a corrected answer key. We identified errors in 99 of LoCoMo’s 1,540 published answers. Ask us for the corrections and the evidence behind each one. Request the corrections ↗︎
How we measure.
The setup matters as much as the score. Here are the samples, response counts, and judges behind ours. Run artifacts are available on request.
Benchmark’s own protocol · 400 questions · one response each · judge gemini-3.5-flash-lite
150 questions · three responses each · judge deepseek-v4-pro
Natural mix · 150 questions · corrected key · judge deepseek-v4-pro
Equal weight · 150 questions · five responses each · corrected key · judge deepseek-v4-pro
Turn your agent on.
Bring us your use case. Let’s work out the integration.