KiokuBench
KiokuBench measures what a memory system remembered, replaced, and forgot by comparing the stored memories themselves against a ledger of true facts. Kioku is Japanese for memory.
Read accuracy alone, as in LoCoMo or LongMemEval, rewards systems that store everything. KiokuBench also checks the contents of memory: stale versions, leftovers from forget requests, noise, language, and size.
Results as of September 28, 2026. Supermemory was measured through its API on September 27 and 28, 2026.
Results
Totals for all 10 personas. Bold marks the better value within the same route.
| Route | Recall | Update applied | Old version kept | Forget violation | Noise | Language mismatch | Check questions | Compression |
|---|---|---|---|---|---|---|---|---|
| Store-everything baseline | 100% | 100% | 100% | 100% | 83% | 0% | 97% | 12.30 |
| sepiace, raw conversation | 99% | 99% | 1% | 0% | 13% | 0% | 97% | 1.29 |
| Supermemory, raw conversation | 98% | 100% | 5% | 78% | 20% | 42% | 88% | 2.30 |
| sepiace, requests | 99% | 99% | 13% | 0% | 0% | 0% | 95% | 1.03 |
| Supermemory, requests, without routing | 83% | 78% | 35% | 65% | 18% | 42% | 82% | 2.06 |
| Supermemory, requests, with routing | 90% | 83% | 46% | 2% | 8% | 42% | 77% | 1.80 |
| sepiace, change signaled | 99% | 99% | 22% | 0% | 0% | 0% | 94% | 1.14 |
| Supermemory, change signaled, without routing | 92% | 90% | 11% | 40% | 12% | 44% | 89% | 1.95 |
The store-everything baseline keeps every utterance as is. For check questions, it passes the utterances closest in embedding up to the context limit.
By language
| Route | Language | Recall | Update applied | Old version kept | Forget violation | Noise | Language mismatch | Check questions |
|---|---|---|---|---|---|---|---|---|
| sepiace, raw conversation | English | 100% | 100% | 0% | 0% | 10% | 0% | 97% |
| Supermemory, raw conversation | English | 100% | 100% | 4% | 70% | 15% | 0% | 89% |
| sepiace, raw conversation | Japanese | 98% | 98% | 2% | 0% | 16% | 0% | 96% |
| Supermemory, raw conversation | Japanese | 97% | 100% | 6% | 85% | 25% | 79% | 87% |
| sepiace, requests | English | 99% | 100% | 16% | 0% | 0% | 0% | 96% |
| Supermemory, requests, without routing | English | 68% | 58% | 30% | 60% | 30% | 0% | 86% |
| sepiace, requests | Japanese | 99% | 98% | 10% | 0% | 0% | 0% | 95% |
| Supermemory, requests, without routing | Japanese | 98% | 98% | 40% | 70% | 7% | 81% | 77% |
| sepiace, change signaled | English | 99% | 100% | 8% | 0% | 0% | 0% | 95% |
| Supermemory, change signaled, without routing | English | 87% | 84% | 2% | 15% | 20% | 0% | 95% |
| sepiace, change signaled | Japanese | 99% | 98% | 36% | 0% | 0% | 0% | 93% |
| Supermemory, change signaled, without routing | Japanese | 98% | 96% | 20% | 65% | 5% | 84% | 83% |
What the results show
- sepiace matched or beat Supermemory on every metric, except for old version kept on the change-signaled route.
- On that route, sepiace kept the old version 22% of the time, and 36% in Japanese. 17 of the 22 cases were facts where only part changed. For example, “I swim at the city pool every Tuesday at 6 p.m.” followed by “I moved my swim from 6 to 6:30.” The new sentence mentions only the time, so removing the old one would lose the pool. Since sepiace never rewrites memory text, it keeps both.
- Supermemory stored about 80% of Japanese input as English memories.
Metrics
| Metric | Meaning | Better |
|---|---|---|
| Recall | Share of facts true at the end that remain in memory | Higher |
| Update applied | Share of changed facts whose new version is in memory | Higher |
| Old version kept | Share of changed facts whose old version also remains in memory | Lower |
| Forget violation | Share of content the user asked to forget that still remains | Lower |
| Noise | Share of memory sentences that match no fact in the ledger | Lower |
| Language mismatch | Share of memory sentences written in a different language from the conversation | Lower |
| Check questions | Share of check questions answered correctly from about 1,000 tokens of context built from memory | Higher |
| Compression | Characters in memory divided by characters in the ledger facts. Closer to 1 means less waste | Closer to 1 |
Routes
There are three ways to feed data in.
- Raw conversation. The conversation is passed in as is, and the memory system decides what to keep. sepiace receives one utterance at a time. Supermemory receives each session as one document.
- Requests. “Remember” and “forget” requests built from the ledger are passed in. For a changed fact, only the new fact is sent as “remember”, without saying that it changed.
- Change signaled. The same as Requests, but a change request says that the fact changed from A to B.
For Supermemory on the request routes, two setups were tried. Without routing, forget requests are sent as documents and Supermemory decides by itself, which is the same condition as sepiace. With routing, only forget requests are sent to its forget-matching endpoint, which tells it the kind of operation.
Dataset v0.1
- 10 personas: 5 in English and 5 in Japanese.
- Each persona has a ledger of 40 facts, or 39 for two personas, with 30 conversation sessions built from the ledger and one check question per fact.
- 10 facts change along the way, and the user asks to forget 4 of them.
- The ledgers, conversations, and questions were generated with openai/gpt-6-luna on OpenRouter.
- A grader model checked that the conversations follow the ledger. It found one fact not in the ledger and two requests to forget only part of a fact. No human review has been done yet.
Conditions
- The grader and the answerer are both openai/gpt-6-luna on OpenRouter.
- sepiace was measured with its code as of September 28, 2026.
- Supermemory ingested with dreaming set to instant. The default, dynamic, builds memories about 10 minutes after ingestion pauses, and waiting for that after every conversation would take more than 5 hours per persona.
- Supermemory's forget-matching was used with its default settings.
- Supermemory's search used the settings Supermemory itself uses in MemoryBench: hybrid, threshold 0.3, 30 results.
Limitations
- This is a single run on 10 personas. Judgments vary, so the same code can shift by a few cases from run to run.
- The dataset is small and has not been reviewed by people yet. A larger v1 is planned.