| Metric | V1 (March 2026, GPT-4o) | V2 (May 2026, Llama 70B) |
|---|---|---|
| Recall@1 | 96% | 100% |
| Recall@3 | 96% | 100% |
| Recall@10 | 96% | 100% |
| QA Accuracy | 76% | 94% |
| Store Failures | 6 / 50 | 0 / 50 |
| Answer Model | GPT-4o ($2.50/M tokens) | Llama 3.3 70B (free) |
| Recall Latency p50 | 399ms | 1,236ms |
| System | Oracle QA Accuracy | Recall@1 |
|---|---|---|
| GPT-4o (full 128K context) | ~70% | N/A |
| MemGPT / Letta | ~45% | ~40% |
| Mem0 | ~35% | ~25% |
| ZeroMemory V2 | 94% | 100% |