| Category | What It Tests | Example |
|---|---|---|
| Single-hop | Recall a specific fact from one conversation | "When did Caroline go to the LGBTQ support group?" |
| Multi-hop | Connect information across multiple sessions | "What activities did both Caroline and Melanie discuss?" |
| Temporal | Reason about dates, durations, and sequences | "How many days between the charity race and the concert?" |
| Open-domain | Answer broad questions about preferences and opinions | "What are Caroline's plans for the summer?" |
| Adversarial | Correctly refuse to answer trick questions | "Did Caroline make the bowl in the photo?" (she didn't) |
| Category | Score | Questions |
|---|---|---|
| Single-hop | 100% | 37 |
| Temporal | 100% | 13 |
| Adversarial | 100% | 2 |
| Open-domain | 94.3% | 70 |
| Multi-hop | 93.8% | 32 |
| Overall | 96.1% | 154 |
| System | Overall Score | Funding |
|---|---|---|
| ZeroMemory | 96.1% | Bootstrapped |
| Mem0 (April 2026) | 92.5% | $23.6M Series A |
| MemMachine | 84.9% | Open source |
| Memobase | 75.8% | Open source |
| Zep | 75.1% | $3.5M Seed |
| Mem0 (2025 algorithm) | 66.9% | — |
| LangMem | 58.1% | — |
| OpenAI (baseline) | 52.9% | — |
| Iteration | Score | What Changed |
|---|---|---|
| V1 | 10.3% | Raw turns, basic vector search |
| V4 | 33.9% | Enriched content with timestamps and speaker identity |
| V6 | 37.0% | Query splitting for multi-hop questions |
| V10 | 66.9% | Backend upgrades: 1024-dim embeddings, graph recall, context expansion |
| V11 | 94.2% | Claude Sonnet 4.5 for answer generation |
| V18 | 96.1% | Optimized retrieval prompting + adversarial scoring |