Skip to content

Results by release

One row per release, per benchmark. Rows are never edited after a release ships — a regression is shown by the next row being worse, not by the previous row changing.

Each row states its item count and standard deviation across repeated runs. A mean without a spread is not a result: two systems whose intervals overlap have not been separated by the measurement.

LongMemEval — retrieval (tier 2, judge-free)

How these rows were produced — kept as the record of what was run, not as a reproduction recipe. The harness now clears the store before every run, so this command yields a cold result against a table of warm ones. What that costs you depends on your deployment: with the reranker off it reproduces these figures exactly (measured 2026-08-27, sd 0.0000 on all three metrics); with the reranker on the MRR varies run to run. The script cannot control that — it is memory.reranker.enabled on the daemon. See the axis note below.

curl -sLO https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_oracle.json
scripts/bench-reproduce.sh --dataset-path ./longmemeval_oracle.json \
    --category temporal-reasoning --items 6 --runs 3 --database <your-bench-db>

Item selection is deterministic: the first N items of the named category, in dataset file order. The rows below scored gpt4_2655b836, gpt4_2487a7cb, gpt4_76048e76, gpt4_2312f94c, 0bb5a684, 08f4fc43 — listed so a reproduction can confirm it measured the same questions rather than inferring it from a count.

Dataset: LongMemEval v1-cleaned (xiaowu0162/longmemeval-cleaned, Sept 2025), longmemeval_oracle.json, sha256 821a2034d219ab45846873dd14c14f12cfe7776e73527a483f9dac095d38620c. Budget: max_tokens=4096 both arms. Judge: none — tier 2 is judge-free.

This axis is CLOSED. It measures a warm corpus — the run repeats against an already-ingested store. Since 98037e34 (2026-08-21) the harness clears the store before every run, unconditionally, so a warm repeat can no longer be produced. The table is kept because a shipped row is history; it is not extended. The open axis is all six abilities, cold.

Release Items Ability Context recall Context precision MRR Runs Key
2026.8.3-13-g13476b8f (baseline, pre-tracking) 6 temporal-reasoning 1.0000 ±0.0000 0.9444 ±0.0000 1.0000 ±0.0000 n=3 d0fa4f0f8389
2026.8.7-48-g0e450f9c (shipped as 2026.8.8) 6 temporal-reasoning 1.0000 ±0.0000 0.9444 ±0.0000 1.0000 ±0.0000 n=3 e865104e9959

Both rows were produced by the scripts/bench-reproduce.sh invocation above, on the binary named in the Release column. All three metrics are deterministic across repeated runs — sd exactly 0.0000 — so these rows can be gated on exact equality rather than a tolerance.

The two rows carry different comparability keys, and that is stated rather than smoothed over. The 2026.8.8 run left our_extraction_model, answer_model and judge_model empty, which marks its key PARTIAL; the baseline run named them. Every field that determines what was measured matches — harness version, dataset name and digest, item selection, max_tokens=4096, observed_recall_method=context-assembly. So the relationship is "not provably comparable" rather than "not comparable", and the exact metric equality is offered as the evidence. A reader who wants the strict reading should treat them as two tables of one row each.

Why 2026.8.8's row is dated a day before its tag. The run was taken at 0e450f9c, four commits before 2026.8.8; all four are release-notes and installer plumbing, touching no code the benchmark exercises. This is recorded rather than rounded off, because the alternative — writing the tag and hoping — is how a track record stops being believable.

Re-ingesting the corpus moves the numbers slightly

The rows above repeat the run against an already-ingested corpus, which is what the reproduction script did at the time. Clearing the database and re-ingesting before every run instead gives MRR 0.9444 ±0.0481 (n=3) on the same six items, with recall and precision unchanged.

The cause was recorded at the time as ingest non-determinism: an LLM titler runs over each chunk and does not produce identical titles every time, so retrieval is deterministic given a corpus while the corpus is not identical across re-ingests.

That attribution turned out to be wrong, and is corrected below. Re-measured 2026-08-27 with the reranker disabled, the same six items are deterministic cold — MRR 1.0000, sd exactly 0.0000 across three runs. The ±0.0481 came from the LLM reranker reordering results between runs, not from the corpus changing. The original figure was taken with each system as it ships, reranker active, and the reranker was never named as the source. The advice to compare warm-to-warm or cold-to-cold still holds on general grounds; it is just not what produced this particular spread.

This distinction stopped being a caveat and became the axis break. Until 2026-08-21 the harness never actually performed the clear that its own flag (--i-know-this-wipes), its guard text and its design all promised — the authorisation was built, the action was not (98037e34). Every run before that date is warm; every run after it is cold, and there is no flag to opt out. The fix also established that the missing clear moved the score: on three runs of an identical 120 items, admitted deposits fell 426 → 426 → 209 as the store filled, and judged accuracy moved 0.692 → 0.750 on a manual wipe with nothing else changed — in the direction that understates the system, because a deduped item loses the haystack it is scored against. Warm numbers are therefore not merely incomparable to cold ones; the older ones are pessimistic by an amount nobody has bounded.

The same six items, measured cold

Taken 2026-08-27 on 2026.8.9-50-g4b343821, reranker disabled, store cleared before every run. It is not appended to the table above, because that table is warm and this is cold — appending it is precisely the silent axis change rule 3 forbids.

Release Items Ability Context recall Context precision MRR Runs Key
2026.8.9-50-g4b343821 6 temporal-reasoning 1.0000 ±0.0000 0.9444 ±0.0000 1.0000 ±0.0000 n=3 e865104e9959

On this subset, warm and cold are identical — the same three figures, all with sd exactly 0.0000, as the warm 2026.8.7-48-g0e450f9c row. The corpus-warmth effect the section above warns about does not appear here at all once the reranker is out of the retrieval path.

That reframes the earlier caveat rather than contradicting it. The published cold figure of MRR 0.9444 ±0.0481 was measured with each system as it ships — LLM reranker active — so its variance was a model reordering results between runs, not the corpus differing. Take the model out and the RRF path is deterministic cold, exactly as it is warm.

A gap this exposed — FIXED 2026-09-03. The warm row above and the cold row here carry the same comparability key e865104e9959, because ComparabilityFields did not encode whether the corpus was warm or cold. Two runs on opposite sides of the 2026-08-21 clear fix compared clean, and the tooling would have merged them without complaint. It went unnoticed because on these six items the two regimes give identical numbers — but a key that cannot express an axis this documentation calls decisive is a guard that looks protective and is not.

The key now carries corpus_regime, observed rather than declared, alongside daemon_revision (two releases with the same config also keyed identically). HarnessVersion is 4 → 5, so every key changes and the incomparability is explicit rather than left to a field older runs never carried. The rows above keep their v4 keys as history — they are not recomputed, because a key is a record of what a run could say about itself at the time. Until a v5 run exists for each row, warm-versus-cold across these published figures still has to be checked by hand against the run date.

LongMemEval — retrieval, all six abilities (tier 2, cold corpus)

This is the open axis. It replaces the six-item warm axis above, which the 2026-08-21 clear fix closed. It is a better measurement on both counts that matter: 120 items rather than 6, and all six abilities rather than one — a single ability is a statement about that ability, not about retrieval.

Reproduce:

The harness reads its target and credentials from the ENVIRONMENT, and the three variables below are not optional — a run without them either refuses or, worse, aims somewhere you did not intend. VORNIK_URL and VORNIK_COMPANION_TOKEN are what the harness reads; vornikctl's own VORNIK_API_URL / VORNIK_API_KEY are a different pair and setting only those points the run at http://localhost:8080 — which on a single-host deployment is production. The database guard is what catches that, and it is the last line of defence, not the first.

# The daemon under test. NOT the production one: this run bulk-writes and
# clears the store. Mint the token against that same daemon:
#   VORNIK_API_URL=<daemon> vornikctl companion grant \
#       --project bench --client claude-code --memory-all
export VORNIK_URL=<your-bench-daemon>
export VORNIK_COMPANION_TOKEN=<companion key with memory_read + memory_write>
export VORNIK_BENCH_DSN=postgres://<user>:<pass>@<host>:5432/<your-bench-db>?sslmode=disable

curl -sLO https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_oracle.json
vornikctl bench memory run --system vornik --dataset longmemeval \
    --dataset-path ./longmemeval_oracle.json \
    --dataset-sha256 821a2034d219ab45846873dd14c14f12cfe7776e73527a483f9dac095d38620c \
    --database <your-bench-db> --i-know-this-wipes <your-bench-db> \
    --tier2-only --accept-unverified-path \
    --our-extraction-model "<the embedder the daemon reports>" \
    --recall-method context-assembly \
    --max-items-per-category 20 --max-tokens 4096 \
    --run-dir ./run-$i          # repeat for i in 1 2 3
vornikctl bench memory aggregate ./run-*

--our-extraction-model must match what the system reports it is embedding with, or the run is refused: "a number labelled with a model that did not produce it is worse than no number". Read the value from the daemon rather than typing one — whoami reports it as embedder, and the id the key uses is <provider>/<model>@<dimensions>d, e.g. openai/qwen3-embedding:0.6b@1024d. Leaving it unset, or naming a model the daemon is not using, is what made the 2026-09-08 run's key partial.

Dataset: as above. Budget: max_tokens=4096. Judge: none — tier 2 is judge-free. Corpus: cold; the harness clears the store before each run.

The deployment under test must have memory.reranker.enabled: false. --tier2-only stops requesting the reranked path but cannot disable it, and it refuses the run if recall reports a rerank happened. An LLM reranker is billed per call and reorders between identical runs, so leaving it on is the difference between a deterministic gate and one that fires on noise.

Release Items Abilities Context recall Context precision MRR Runs Key
2026.8.8-40-g2698eb67 120 all 6 0.9900 ±0.0036 0.9854 ±0.0000 0.9958 ±0.0000 n=3 93a6e7a0729b
2026.8.9-50-g4b343821 120 all 6 0.9891 ±0.0036 0.9854 ±0.0000 0.9958 ±0.0000 n=3 93a6e7a0729b
2026.9.4-205-g97fb8b36e 120 all 6 0.9851 ±0.0000 0.9854 ±0.0000 0.9958 ±0.0000 n=3 28fd464fb163
2026.9.4-207-g61634a063 120 all 6 0.9851 ±0.0000 0.9854 ±0.0000 0.9958 ±0.0000 n=3 ed7ea84021e0
2026.9.6-117-gee40525c8 (2026.9.7) 120 all 6 0.9851 ±0.0000 0.9854 ±0.0000 0.9958 ±0.0000 n=3 4a95cccf5f7a
2026.9.8-52-g86d3af802 (2026.10.1) 120 all 6 0.9851 ±0.0000 0.9854 ±0.0000 0.9958 ±0.0000 n=3 0c4971cc78ae
2026.10.1-55-gc800c23be (2026.10.2) 120 all 6 0.9851 ±0.0000 0.9854 ±0.0000 0.9958 ±0.0000 n=3 dcfeaf709dc5
2026.10.1-77-gbb5238eb2 (2026.10.3) 120 all 6 0.9851 ±0.0000 0.9854 ±0.0000 0.9958 ±0.0000 n=3 16a79559dd84

The eighth row is 2026.10.3's release candidate (the tagged commit differs from it only in documentation, this row among them), measured on 2026-10-02 on the memory-benchmark deployment, same embedder, reranker off, cold corpus: each run cleared the store, re-ingested the pinned dataset and recorded the daemon's own revision. Every metric is identical to 2026.10.2 to four decimal places across three deterministic runs; the key is new (16a79559dd84) because daemon_revision moved. The cycle did not change the retrieval or ingest path (it shipped the macOS release check, the Hermes plugin's catalog readiness and EULA version 3), so this row is the no-regression check, not a measurement of a change. The harness reports retrieval_path_unverified: it trusts the daemon's reported recall method (context-assembly) rather than observing it.

The seventh row is 2026.10.2's release candidate (the tagged commit differs from it only in tests and documentation), measured on 2026-10-02 on the memory-benchmark deployment, same embedder, reranker off, cold corpus. Every metric is identical to 2026.10.1 to four decimal places across three deterministic runs; the key is new (dcfeaf709dc5) because daemon_revision moved. The cycle did not change the retrieval or ingest path (it added agent administration, update-path fixes and release packaging), so this row is the no-regression check, not a measurement of a change.

The sixth row is 2026.10.1's release build, measured on 2026-10-01 on the memory-benchmark deployment, same embedder, reranker off, cold corpus. Every metric is identical to 2026.9.7 to four decimal places across three deterministic runs. The key is new (0c4971cc78ae) because daemon_revision moved. The run was required because the cycle lowered the minimum size of a deliberate deposit from a companion client, and this benchmark ingests through that path. What it shows about that change: the stored corpus is unchanged, 1,897 chunks before each run, exactly as in 2026.9.7, so the lower floor admitted nothing the old one refused on this corpus. It is not a test of the floor itself; that is covered by the daemon's own tests.

The fifth row is 2026.9.7's release build, measured on 2026-09-27 on the memory-benchmark deployment. Every metric is identical to the 2026.9.4 build that shipped, to four decimal places, across three deterministic runs. Its key is new (4a95cccf5f7a) because daemon_revision moved.

The cycle between the two rows changed the retrieval path twice: - 2026.9.5 added recency ranking and series supersession. - 2026.9.7 added document supersession, and stores each document version whole under chunk hashes salted by artifact.

The row shows that neither moved tier-2 retrieval on this corpus. What it does NOT show: this benchmark ingests each item once into a cold store, so re-ingest supersession, the behaviour 2026.9.7 exists for, is never exercised. It is covered by the daemon's own tests and was verified live on the reference deployment, not by this table.

The fourth row is the build that shipped, re-measured on the binary actually installed in production (2026.9.4-207-g61634a063) rather than on a development build. Every metric is identical to the row above it, to four decimal places — and the key nevertheless CHANGED, from 28fd464fb163 to ed7ea84021e0, because daemon_revision moved from -205-…-dirty to -207-…. That is the repaired key doing its job: two runs of the same code on different builds are no longer allowed to look like one measurement. Before 2026-09-17 these two rows would have shared a key.

The third row is the first with a COMPLETE comparability key, and it is deliberately not comparable to the two above it. Until 2026-09-17 two of the key's fields — observed_embedder and daemon_revision — were empty on every run ever made (harness design §13.25), so the key could not distinguish one build, or one embedding model, from another. That is why rows one and two share 93a6e7a0729b despite being different releases. Both fields are now populated, which necessarily changes the key: 28fd464fb163 starts a new comparable set rather than extending the old one. compare will refuse to diff across that boundary, correctly.

What the numbers say anyway. Precision and MRR are bit-identical to both earlier rows. Recall is 0.9851, −0.0040 against the last published row — about a third of that row's own 3σ threshold of ±0.0107, so nothing here resembles a regression.

Two caveats that stop this being a clean release-over-release claim, stated rather than left for a reader to infer:

  • The ingest regime differed. The self-hosted vLLM arm is down, so this run had no LLM extraction model and the stored corpus is produced deterministically by the embedder alone. That is visible in the spread: every metric has sd = 0.0000 here against ±0.0036 for both earlier rows, whose variance the 2026-08-21 note attributes to LLM-driven ingest differing per run. A −0.0040 recall difference and a collapsed spread have a common candidate cause, and this measurement cannot separate "the corpus was built differently" from "the code changed".
  • retrieval_path_unverified is set, because the run passed --accept-unverified-path. The reranker was disabled on the deployment under test (memory.reranker.enabled: false) as --tier2-only requires — its model is the same unreachable vLLM — so the path exercised was plain context-assembly.

Re-measuring against a live extraction model is what would turn this into a comparable row; it is blocked on the benchmark arm's model endpoint, not on anything in the harness.

The second row is the first release-over-release comparison this table can actually support, and it shows no regression. Precision and MRR are identical to four decimal places; recall moved −0.0009 against a per-release sd of 0.0036 and a narrowest-defensible (3σ) threshold of ±0.0107 — about a twelfth of the smallest move that could fire without being noise. The sd itself reproduced exactly (0.0036 both times), which is the more reassuring number: it says the measurement is stable, not merely that two point estimates happened to agree.

Figures from bench memory aggregate over the three run directories, not hand-computed. Standard deviations are sample (n−1), which is what the tool reports.

Per ability, same three runs:

Ability Context recall Context precision MRR Items
knowledge-update 1.0000 ±0.0000 1.0000 ±0.0000 1.0000 ±0.0000 20
multi-session 0.9564 ±0.0250 0.9667 ±0.0000 0.9750 ±0.0000 20
single-session-assistant 1.0000 ±0.0000 1.0000 ±0.0000 1.0000 ±0.0000 20
single-session-preference 1.0000 ±0.0000 1.0000 ±0.0000 1.0000 ±0.0000 20
single-session-user 1.0000 ±0.0000 1.0000 ±0.0000 1.0000 ±0.0000 20
temporal-reasoning 0.9833 ±0.0144 0.9458 ±0.0000 1.0000 ±0.0000 20

Where it loses. multi-session is the only ability that misses recall, and it is also the only one whose recall varies run to run (sd 0.0250 against 0.0000 on four of six). Questions needing evidence spread across sessions are where this system is weakest, and the variance says the weakness is not a fixed set of items — the boundary moves. temporal-reasoning carries the lowest precision (0.9458): it retrieves the right documents and pads the set.

Caveats specific to this row.

  • The key is PARTIAL. our_extraction_model, answer_model and judge_model were left empty, so the run cannot be proven comparable to another — only observed to match on every field it does record.
  • retrieval_path_unverified is set. The reranker was enabled on the deployment, so the observed path is context-assembly|context-assembly+rerank — a model is in the retrieval path and the run is therefore not deterministic by construction. This is the correct setting for reporting the system as shipped and the wrong one for a CI gate, which wants the reranker off and the RRF path proven.
  • n=3 on 120 items. Enough for a spread, not enough to resolve a sub-point move.

Agent quality — 2026.9.0 (first scored arm)

A different benchmark from the LongMemEval rows above, measuring a different thing: not retrieval, but the decisions the control logic makes — what the lead granted, whether roles followed their output schemas, and whether agents called tools correctly. It runs 30 software tasks through a multi-agent dev-pipeline swarm against an operator-reviewed answer key.

Release Tasks Task success Schema conformance Tool-call validity Steps with no output Cost/task
2026.8.9-70-g5d247f72 (shipped as 2026.9.0) 30 100.0% 0.985 1.000 9.7% $0.29

Efficiency, same arm: 667,801 tokens and 62.4 tool calls per task, 0 escalations, 0 schema retries. Total spend $8.79.

Cost corrected downwards, 2026-09-19. This row first published $0.74 per task and $25.23 total. Those figures were wrong and we are the ones who found it: the observed model was absent from our pricing table, so every call was billed at the table's default rate of $1.00/$3.00 per million tokens instead of the model's real $0.35/$2.75. The corrected figures are recomputed from the arm's own recorded token counts — 22,353,287 prompt and 351,933 completion — and nothing else about the run changed.

The true cost may be lower still. This arm reached the model over a path that did not report prompt-cache reads, and a later pass on the same model measured 79.5% of prompt tokens served from cache at a tenth of the input rate. We do not know that share for this arm, so we have not applied it: the number above is an upper bound, stated as one.

Quality figures in this row are unaffected — pricing enters no scoring path.

Read both layers, because the first one alone flatters us. Task success is 100%, and underneath it 14 of 144 terminal steps (9.7%) produced no output at all — 8 hit the iteration cap, 4 entered a degenerate tool loop, 2 failed outright. Recovery absorbs those, which is the system working as designed, but a headline "100%" without the step figure would describe a smoother product than exists.

The model is part of the result, not a footnote

Model: Qwen/Qwen3.8-27B-FP8, self-hosted, one box. Every figure above is that model's behaviour as much as the control logic's, and the difference is large enough to change how the numbers read.

Measured over ~244,000 tool calls across two deployments: this model enters identical-repeat tool loops 26x more often than the mix of larger hosted models (0.52% of calls against 0.02%), and once nudged out of one it changes approach 36% of the time against 82%.

That is deliberate and it is the point. A self-hosted 27B on a single machine is what someone running this at home actually has. An organisation putting a large hosted model behind the same control logic sees the lower rate; these figures describe the harder case rather than the flattering one. The comparability key pins the model identity, so a future arm on a different model refuses to compare against this row rather than quietly superseding it.

What this row does NOT establish

  • No pass or fail. bench agent gate refused a verdict, correctly: resolving a 5-point effect at the inherited σ=0.0604 needs 12 paired tasks and this arm has 5, which can only resolve 7.6 points. A smaller movement must be reported as inconclusive with that floor, never as "no change".
  • No trend. It is the first scored arm; there is nothing to compare it against. The releases before it have no agent row at all.
  • No noise floor for this task set. The σ above is inherited from an earlier 3-task measurement and describes a different set. The honest σ for these 30 tasks does not exist yet, which is why the gate refuses.
  • Not independently reproducible. Unlike the LongMemEval rows, which run against a public dataset anyone can fetch, the agent task set and its answer key are not published. An external reader can see the method and the figures and cannot re-run them. That is a real limitation of this row and is stated rather than left to be discovered.

Provenance

Both the daemon binary and the agent image are recorded by content digest in the run's arm key, and the run refuses to merge batches that disagree on any axis. Arms recorded before 2026-08-29 carry no agent-image identity at all — the executor discarded every image ID it observed, so those runs are marked untrustworthy and cannot serve as a baseline. This is the first arm whose image provenance is real.

Releases with no row, and why

The policy in RELEASE.md requires a row per release from 2026.8.4. It was not followed. Rather than leave the gaps silent — which reads as "nothing regressed" — each is stated:

Release Row Why
2026.8.4 none No benchmark run was taken. Not backfillable: it would need the tagged binary rebuilt and the bench deployment's models restored to what they were, and those were changed during 2026.8.7.
2026.8.5 none As above.
2026.8.6 none As above.
2026.8.7 none No run. The reason was recorded at the time in the 2026.8.7 release notes: every outward-facing provider was disabled and all roles moved to a locally served model during that cycle, so a run could not be pinned to the baseline row's axes.
2026.8.8 yes Both axes above. The warm row was measured at 0e450f9c; the cold row at 2698eb67.
2026.8.9 none at the tag No run was taken at 2026.8.9 itself. The tree past it is measured on both memory axes above.
2026.9.0 agent row only The agent arm above was taken on the release candidate (2026.8.9-70-g5d247f72). No LongMemEval row: the memory axes were measured 50 commits earlier in the same cycle and nothing in the intervening work touches the retrieval path. Stated rather than left as a silent gap.
2026.9.1 none No run was taken. The cycle is bug fixes and the forge re-review feature; nothing in it touches ingestion, embedding, retrieval or the agent harness, so the axes stand where 2026.8.9-50-g4b343821 (memory, both axes) and 2026.8.9-70-g5d247f72 (agent arm) left them. A row measured on an unchanged path would add a data point without adding evidence, and the cost of a 120-item n=3 run is not free. Stated rather than left as a silent gap.
2026.9.2 none No run was taken. The cycle moves the agent loop's eleven filesystem/git tools from bash into a Go helper with behaviour pinned by a 64-case golden, and adds persistence and replay instruments; it does not touch ingestion, embedding, retrieval or the harness's scoring path, so the axes stand where the rows above left them. The agent arm should be re-measured on the Go helper before the next loop slice moves more of it — that is filed as work, not claimed here. Stated rather than left as a silent gap.
2026.9.3 none No run was taken. The cycle is the update path and the pulled agent image; it does not touch ingestion, embedding, retrieval or the harness. Stated late — this row was added with 2026.9.4's — rather than left as a silent gap.
2026.9.4 none No run could be taken: the benchmark arm's model endpoint has been unreachable since 2026-09-04 (it now accepts and immediately resets, which a port check reads as up). The cycle does not touch ingestion, embedding, retrieval or the scoring path; it does change what a multi-system-step workflow's agent receives, which is not measured here. Stated rather than left as a silent gap.
2026.9.5 none — a gap, not a waiver No run was taken, and this cycle DID change the retrieval path: recall began ranking by recency and letting a recurring series supersede its earlier members. It shipped unmeasured on the axes above. Stated late — this row was added with 2026.9.7's, whose run measures the tree that includes it — rather than left as a silent gap.
2026.9.6 none No run was taken. The cycle bounds agent containers and fixes the adoption board; the board reads the retrieval and ingest audit ledgers but changes neither path. Stated late, with 2026.9.7's, rather than left as a silent gap.
2026.9.7 yes The memory open axis above, on the release build (2026.9.6-117-gee40525c8, n=3). The cycle changes ingestion and retrieval, so a waiver was not open to it. The v9 agent-harness arms run on the slow-hardware track and are reported in the release notes, not in this table.
2026.9.8 none No run was taken. The patch changes only the companion endpoint's reply to a GET stream request, plus test fixtures. It touches no ingestion, embedding, retrieval or scoring path, so the axes stand where 2026.9.7's measured row left them. Stated rather than left as a silent gap.
2026.10.1 memory row only The memory open axis above, on the release build (n=3). The cycle lowers the minimum size of a deliberate memory deposit from a companion client, and this benchmark ingests through that path, so a waiver was not open to it. No agent row. The cycle changes the agent harness substantially (the router step, the retry ladder, the prompt-token budget and the tool-result cap), so this is a gap, not a waiver. The agent arm's model is the self-hosted Qwen3.8-27B on the endpoint that has been unreachable since 2026-09-04, and the comparability key pins the model, so a run on any other model would start a new table that cannot show a regression against the 2026.9.0 row. A substitute arm on a locally served 20B model was considered and not run: one arm is about 20M prompt tokens, which on the reference host's integrated GPU is days, not hours. The changes are covered instead by regression tests that replay each incident, and by an end-to-end lane that drives a front-end agent through the broker on a locally served model.

A missing row stated as missing is honest. An absent row is not — which is the whole reason this section exists.

LongMemEval — judged accuracy (tier 1) and the judge's noise floor

A judge with an unmeasured noise floor is not a standard. Before any judged number can be gated or quoted as a win, the variance of the judge itself has to be known — otherwise a "win" is indistinguishable from the judge disagreeing with itself.

Ten identical runs, same items, same models, warm corpus:

Metric Mean sd Min Max Narrowest defensible gate (3σ)
judged accuracy 0.8241 0.0340 0.7500 0.8750 ±0.1020
context recall 0.9764 0.0059 0.9722 0.9861 ±0.0176
context precision 0.9410 0.0000 0.9410 0.9410 exact
MRR 0.9917 0.0108 0.9792 1.0000 ±0.0323

n=10 runs × 24 items, temporal-reasoning, comparability key a2fd4d245b21. Judge: gemma4:31b. Answerer: gpt-oss:120b — a different model family, so the judge is not grading its own output. Both open-weight and version-pinned, served from a single provider (Ollama Cloud). A multi-provider router was deliberately avoided: switching backends between calls would inject variance into the very number being measured.

What this means in practice. Judged accuracy moved between 0.750 and 0.875 with nothing changing but the models' own sampling — 3 of 24 items flipping verdict. A tier-1 gate would therefore need a ±10.2 point threshold to avoid firing on noise, which is wider than most real improvements. That is the entire argument for leading with tier 2: on the same runs, context precision was exactly constant.

Note that context recall is not deterministic here (sd 0.0059) although it is on the tier-2 rows above. These judged runs measure the system as shipped, with the LLM reranker in the retrieval path; the tier-2 rows measure the RRF path, which has no model in it. Same product, two paths, and only one of them can carry an exact-equality gate.

Reproducing the judge-variance figure

export VORNIK_BENCH_LLM_URL=https://ollama.com/v1
export VORNIK_BENCH_LLM_KEY=<your key>
for i in $(seq 1 10); do
  vornikctl bench memory run --system vornik --dataset longmemeval \
    --dataset-path ./longmemeval_oracle.json \
    --database <your-bench-db> --i-know-this-wipes <your-bench-db> \
    --answer-model gpt-oss:120b --judge-model gemma4:31b \
    --category temporal-reasoning --max-items 24 --max-tokens 4096 \
    --run-dir ./jv-$i
done
vornikctl bench memory aggregate ./jv-*

Roughly 4.5 minutes per run on the reference deployment; the first run is much slower because it ingests the corpus.

Comparison: hindsight 0.9.0, same items, same budget

Measured under identical conditions on the same machine, each system as it ships — both with their own rerankers active. Recorded because a comparison that only flatters us is worth less than one that is obviously honest.

System Context recall Context precision MRR Runs
Vornik 1.0000 ±0.0000 0.9444 ±0.0000 0.9444 ±0.0481 n=3
hindsight 0.9.0 1.0000 ±0.0000 0.6806 ±0.0833 0.7292 ±0.1250 n=4

Both arms are cold here — database truncated / bank deleted before every run — so the two are measured the same way. That is why Vornik's MRR reads 0.9444 ±0.0481 in this table and 1.0000 ±0.0000 in the release row above: the release row repeats against a warm corpus, as the reproduction script does. Comparing a warm number with a cold one would flatter whichever arm was warm.

Recall is tied at 1.000. Both systems retrieved every gold document on all six items, so this subset does not separate them on recall — it is the oracle haystack, which carries about three sessions per item where the full longmemeval_s carries 38–62. Expect recall to separate the systems on the larger haystack, and treat a tie on an easy subset as a statement about the subset.

Vornik's retrieved set is tighter (precision 0.944 vs 0.681) and better ordered (MRR 0.944 vs 0.729). hindsight reaches the same recall by returning more.

One difference matters more than the scores. Vornik's tier-2 figures are deterministic: recall and precision have standard deviation exactly 0.0000 across repeated runs, because its ingest path is chunk-and-embed with no model deciding what gets stored. hindsight's precision and MRR vary run to run (sd 0.083 and 0.125, with MRR ranging 0.667–0.917 across four runs) because its ingest consolidates memories through an LLM, so the corpus itself differs each time.

That has a practical consequence: a single hindsight run cannot be quoted, and any comparison against it needs repeated runs and a spread. Vornik can be gated on exact equality; hindsight would need a ±0.375 MRR tolerance to avoid firing on its own noise.

Method and honest caveats

  • n=6, one of six abilities. A smoke-scale result. It exists to prove the pipeline end to end and to establish a baseline, not to rank products.
  • Both arms are marked retrieval_path_unverified in their comparability keys: each system was measured with its own reranker in the path, so neither run establishes a deterministic retrieval path. This is the correct setting for a product comparison and the wrong one for a CI gate.
  • hindsight's internal extraction model was pinned to an open-weight hosted model (google/gemma-4-26b-a4b-it). Its embedder and reranker were its own bundled local defaults. Vornik used its configured local embedder.
  • Hindsight banks are deleted between runs. Without that its corpus accumulates and precision falls run over run — the first measurement taken before this was wired read 0.806 precision against 0.639 on an identical repeat.