[ICLR 2026] RIVER: A Real-Time Interaction Benchmark for Video LLMs
-
Updated
Apr 20, 2026 - Python
[ICLR 2026] RIVER: A Real-Time Interaction Benchmark for Video LLMs
The first open evaluation framework for AI continuity. 250 narrative tests, 1835 verification questions, 10 checkpoints. Benchmark for AI memory systems, stateful agents, and long-term context persistence. No LLM in the evaluation loop.
Reproducible LoCoMo, LongMemEval, PersonaMem, FAMA, recall and latency benchmarks for long-term AI memory systems
A cross-platform Python tool for benchmarking system memory performance. Easily compare RAM speed and efficiency across different hardware, operating systems, and configurations. Includes robust logging, CSV export, and built-in graphing for visual analysis.
Add a description, image, and links to the memory-benchmark topic page so that developers can more easily learn about it.
To associate your repository with the memory-benchmark topic, visit your repo's landing page and select "manage topics."