KiokuBench

KiokuBench measures what a memory system remembered, replaced, and forgot by comparing the stored memories themselves against a ledger of true facts. Kioku is Japanese for memory.

Read accuracy alone, as in LoCoMo or LongMemEval, rewards systems that store everything. KiokuBench also checks the contents of memory: stale versions, leftovers from forget requests, noise, language, and size.

Results as of September 28, 2026. Supermemory was measured through its API on September 27 and 28, 2026.

Results

Totals for all 10 personas. Bold marks the better value within the same route.

KiokuBench v0.1 results, totals for 10 personas
RouteRecallUpdate appliedOld version keptForget violationNoiseLanguage mismatchCheck questionsCompression
Store-everything baseline100%100%100%100%83%0%97%12.30
sepiace, raw conversation99%99%1%0%13%0%97%1.29
Supermemory, raw conversation98%100%5%78%20%42%88%2.30
sepiace, requests99%99%13%0%0%0%95%1.03
Supermemory, requests, without routing83%78%35%65%18%42%82%2.06
Supermemory, requests, with routing90%83%46%2%8%42%77%1.80
sepiace, change signaled99%99%22%0%0%0%94%1.14
Supermemory, change signaled, without routing92%90%11%40%12%44%89%1.95

The store-everything baseline keeps every utterance as is. For check questions, it passes the utterances closest in embedding up to the context limit.

By language

KiokuBench v0.1 results by language
RouteLanguageRecallUpdate appliedOld version keptForget violationNoiseLanguage mismatchCheck questions
sepiace, raw conversationEnglish100%100%0%0%10%0%97%
Supermemory, raw conversationEnglish100%100%4%70%15%0%89%
sepiace, raw conversationJapanese98%98%2%0%16%0%96%
Supermemory, raw conversationJapanese97%100%6%85%25%79%87%
sepiace, requestsEnglish99%100%16%0%0%0%96%
Supermemory, requests, without routingEnglish68%58%30%60%30%0%86%
sepiace, requestsJapanese99%98%10%0%0%0%95%
Supermemory, requests, without routingJapanese98%98%40%70%7%81%77%
sepiace, change signaledEnglish99%100%8%0%0%0%95%
Supermemory, change signaled, without routingEnglish87%84%2%15%20%0%95%
sepiace, change signaledJapanese99%98%36%0%0%0%93%
Supermemory, change signaled, without routingJapanese98%96%20%65%5%84%83%

What the results show

  • sepiace matched or beat Supermemory on every metric, except for old version kept on the change-signaled route.
  • On that route, sepiace kept the old version 22% of the time, and 36% in Japanese. 17 of the 22 cases were facts where only part changed. For example, “I swim at the city pool every Tuesday at 6 p.m.” followed by “I moved my swim from 6 to 6:30.” The new sentence mentions only the time, so removing the old one would lose the pool. Since sepiace never rewrites memory text, it keeps both.
  • Supermemory stored about 80% of Japanese input as English memories.

Metrics

MetricMeaningBetter
RecallShare of facts true at the end that remain in memoryHigher
Update appliedShare of changed facts whose new version is in memoryHigher
Old version keptShare of changed facts whose old version also remains in memoryLower
Forget violationShare of content the user asked to forget that still remainsLower
NoiseShare of memory sentences that match no fact in the ledgerLower
Language mismatchShare of memory sentences written in a different language from the conversationLower
Check questionsShare of check questions answered correctly from about 1,000 tokens of context built from memoryHigher
CompressionCharacters in memory divided by characters in the ledger facts. Closer to 1 means less wasteCloser to 1

Routes

There are three ways to feed data in.

  • Raw conversation. The conversation is passed in as is, and the memory system decides what to keep. sepiace receives one utterance at a time. Supermemory receives each session as one document.
  • Requests. “Remember” and “forget” requests built from the ledger are passed in. For a changed fact, only the new fact is sent as “remember”, without saying that it changed.
  • Change signaled. The same as Requests, but a change request says that the fact changed from A to B.

For Supermemory on the request routes, two setups were tried. Without routing, forget requests are sent as documents and Supermemory decides by itself, which is the same condition as sepiace. With routing, only forget requests are sent to its forget-matching endpoint, which tells it the kind of operation.

Dataset v0.1

  • 10 personas: 5 in English and 5 in Japanese.
  • Each persona has a ledger of 40 facts, or 39 for two personas, with 30 conversation sessions built from the ledger and one check question per fact.
  • 10 facts change along the way, and the user asks to forget 4 of them.
  • The ledgers, conversations, and questions were generated with openai/gpt-6-luna on OpenRouter.
  • A grader model checked that the conversations follow the ledger. It found one fact not in the ledger and two requests to forget only part of a fact. No human review has been done yet.

Conditions

  • The grader and the answerer are both openai/gpt-6-luna on OpenRouter.
  • sepiace was measured with its code as of September 28, 2026.
  • Supermemory ingested with dreaming set to instant. The default, dynamic, builds memories about 10 minutes after ingestion pauses, and waiting for that after every conversation would take more than 5 hours per persona.
  • Supermemory's forget-matching was used with its default settings.
  • Supermemory's search used the settings Supermemory itself uses in MemoryBench: hybrid, threshold 0.3, 30 results.

Limitations

  • This is a single run on 10 personas. Judgments vary, so the same code can shift by a few cases from run to run.
  • The dataset is small and has not been reviewed by people yet. A larger v1 is planned.