Results by release¶
One row per release, per benchmark. Rows are never edited after a release ships — a regression is shown by the next row being worse, not by the previous row changing.
Each row states its item count and standard deviation across repeated runs. A mean without a spread is not a result: two systems whose intervals overlap have not been separated by the measurement.
LongMemEval — retrieval (tier 2, judge-free)¶
How these rows were produced — kept as the record of what was run, not as a
reproduction recipe. The harness now clears the store before every run, so this command
yields a cold result against a table of warm ones. What that costs you depends on
your deployment: with the reranker off it reproduces these figures exactly
(measured 2026-08-27, sd 0.0000 on all three metrics); with the reranker on the
MRR varies run to run. The script cannot control that — it is
memory.reranker.enabled on the daemon. See the axis note below.
curl -sLO https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_oracle.json
scripts/bench-reproduce.sh --dataset-path ./longmemeval_oracle.json \
--category temporal-reasoning --items 6 --runs 3 --database <your-bench-db>
Item selection is deterministic: the first N items of the named category, in dataset
file order. The rows below scored gpt4_2655b836, gpt4_2487a7cb, gpt4_76048e76,
gpt4_2312f94c, 0bb5a684, 08f4fc43 — listed so a reproduction can confirm it
measured the same questions rather than inferring it from a count.
Dataset: LongMemEval v1-cleaned (xiaowu0162/longmemeval-cleaned, Sept 2025),
longmemeval_oracle.json, sha256 821a2034d219ab45846873dd14c14f12cfe7776e73527a483f9dac095d38620c.
Budget: max_tokens=4096 both arms. Judge: none — tier 2 is judge-free.
This axis is CLOSED. It measures a warm corpus — the run repeats against an
already-ingested store. Since 98037e34 (2026-08-21) the harness clears the store
before every run, unconditionally, so a warm repeat can no longer be produced. The
table is kept because a shipped row is history; it is not extended. The open axis is
all six abilities, cold.
| Release | Items | Ability | Context recall | Context precision | MRR | Runs | Key |
|---|---|---|---|---|---|---|---|
2026.8.3-13-g13476b8f (baseline, pre-tracking) |
6 | temporal-reasoning | 1.0000 ±0.0000 | 0.9444 ±0.0000 | 1.0000 ±0.0000 | n=3 | d0fa4f0f8389 |
2026.8.7-48-g0e450f9c (shipped as 2026.8.8) |
6 | temporal-reasoning | 1.0000 ±0.0000 | 0.9444 ±0.0000 | 1.0000 ±0.0000 | n=3 | e865104e9959 |
Both rows were produced by the scripts/bench-reproduce.sh invocation above, on the
binary named in the Release column. All three metrics are deterministic across
repeated runs — sd exactly 0.0000 — so these rows can be gated on exact equality
rather than a tolerance.
The two rows carry different comparability keys, and that is stated rather than
smoothed over. The 2026.8.8 run left our_extraction_model, answer_model and
judge_model empty, which marks its key PARTIAL; the baseline run named them. Every
field that determines what was measured matches — harness version, dataset name and
digest, item selection, max_tokens=4096, observed_recall_method=context-assembly.
So the relationship is "not provably comparable" rather than "not comparable", and the
exact metric equality is offered as the evidence. A reader who wants the strict
reading should treat them as two tables of one row each.
Why 2026.8.8's row is dated a day before its tag. The run was taken at
0e450f9c, four commits before 2026.8.8; all four are release-notes and installer
plumbing, touching no code the benchmark exercises. This is recorded rather than
rounded off, because the alternative — writing the tag and hoping — is how a track
record stops being believable.
Re-ingesting the corpus moves the numbers slightly¶
The rows above repeat the run against an already-ingested corpus, which is what the reproduction script did at the time. Clearing the database and re-ingesting before every run instead gives MRR 0.9444 ±0.0481 (n=3) on the same six items, with recall and precision unchanged.
The cause was recorded at the time as ingest non-determinism: an LLM titler runs over each chunk and does not produce identical titles every time, so retrieval is deterministic given a corpus while the corpus is not identical across re-ingests.
That attribution turned out to be wrong, and is corrected below. Re-measured 2026-08-27 with the reranker disabled, the same six items are deterministic cold — MRR 1.0000, sd exactly 0.0000 across three runs. The ±0.0481 came from the LLM reranker reordering results between runs, not from the corpus changing. The original figure was taken with each system as it ships, reranker active, and the reranker was never named as the source. The advice to compare warm-to-warm or cold-to-cold still holds on general grounds; it is just not what produced this particular spread.
This distinction stopped being a caveat and became the axis break. Until
2026-08-21 the harness never actually performed the clear that its own flag
(--i-know-this-wipes), its guard text and its design all promised — the
authorisation was built, the action was not (98037e34). Every run before that date
is warm; every run after it is cold, and there is no flag to opt out. The fix also
established that the missing clear moved the score: on three runs of an identical
120 items, admitted deposits fell 426 → 426 → 209 as the store filled, and judged
accuracy moved 0.692 → 0.750 on a manual wipe with nothing else changed — in the
direction that understates the system, because a deduped item loses the haystack
it is scored against. Warm numbers are therefore not merely incomparable to cold ones;
the older ones are pessimistic by an amount nobody has bounded.
The same six items, measured cold¶
Taken 2026-08-27 on 2026.8.9-50-g4b343821, reranker disabled, store cleared before
every run. It is not appended to the table above, because that table is warm and
this is cold — appending it is precisely the silent axis change rule 3 forbids.
| Release | Items | Ability | Context recall | Context precision | MRR | Runs | Key |
|---|---|---|---|---|---|---|---|
2026.8.9-50-g4b343821 |
6 | temporal-reasoning | 1.0000 ±0.0000 | 0.9444 ±0.0000 | 1.0000 ±0.0000 | n=3 | e865104e9959 |
On this subset, warm and cold are identical — the same three figures, all with sd
exactly 0.0000, as the warm 2026.8.7-48-g0e450f9c row. The corpus-warmth effect the
section above warns about does not appear here at all once the reranker is out of the
retrieval path.
That reframes the earlier caveat rather than contradicting it. The published cold figure of MRR 0.9444 ±0.0481 was measured with each system as it ships — LLM reranker active — so its variance was a model reordering results between runs, not the corpus differing. Take the model out and the RRF path is deterministic cold, exactly as it is warm.
A gap this exposed — FIXED 2026-09-03. The warm row above and the cold row here carry the same comparability key
e865104e9959, becauseComparabilityFieldsdid not encode whether the corpus was warm or cold. Two runs on opposite sides of the 2026-08-21 clear fix compared clean, and the tooling would have merged them without complaint. It went unnoticed because on these six items the two regimes give identical numbers — but a key that cannot express an axis this documentation calls decisive is a guard that looks protective and is not.The key now carries
corpus_regime, observed rather than declared, alongsidedaemon_revision(two releases with the same config also keyed identically). HarnessVersion is 4 → 5, so every key changes and the incomparability is explicit rather than left to a field older runs never carried. The rows above keep their v4 keys as history — they are not recomputed, because a key is a record of what a run could say about itself at the time. Until a v5 run exists for each row, warm-versus-cold across these published figures still has to be checked by hand against the run date.
LongMemEval — retrieval, all six abilities (tier 2, cold corpus)¶
This is the open axis. It replaces the six-item warm axis above, which the 2026-08-21 clear fix closed. It is a better measurement on both counts that matter: 120 items rather than 6, and all six abilities rather than one — a single ability is a statement about that ability, not about retrieval.
Reproduce:
The harness reads its target and credentials from the ENVIRONMENT, and the three
variables below are not optional — a run without them either refuses or, worse,
aims somewhere you did not intend. VORNIK_URL and VORNIK_COMPANION_TOKEN are
what the harness reads; vornikctl's own VORNIK_API_URL / VORNIK_API_KEY
are a different pair and setting only those points the run at
http://localhost:8080 — which on a single-host deployment is production.
The database guard is what catches that, and it is the last line of defence, not
the first.
# The daemon under test. NOT the production one: this run bulk-writes and
# clears the store. Mint the token against that same daemon:
# VORNIK_API_URL=<daemon> vornikctl companion grant \
# --project bench --client claude-code --memory-all
export VORNIK_URL=<your-bench-daemon>
export VORNIK_COMPANION_TOKEN=<companion key with memory_read + memory_write>
export VORNIK_BENCH_DSN=postgres://<user>:<pass>@<host>:5432/<your-bench-db>?sslmode=disable
curl -sLO https://huggingface.co/datasets/xiaowu0162/longmemeval-cleaned/resolve/main/longmemeval_oracle.json
vornikctl bench memory run --system vornik --dataset longmemeval \
--dataset-path ./longmemeval_oracle.json \
--dataset-sha256 821a2034d219ab45846873dd14c14f12cfe7776e73527a483f9dac095d38620c \
--database <your-bench-db> --i-know-this-wipes <your-bench-db> \
--tier2-only --accept-unverified-path \
--our-extraction-model "<the embedder the daemon reports>" \
--recall-method context-assembly \
--max-items-per-category 20 --max-tokens 4096 \
--run-dir ./run-$i # repeat for i in 1 2 3
vornikctl bench memory aggregate ./run-*
--our-extraction-model must match what the system reports it is embedding
with, or the run is refused: "a number labelled with a model that did not
produce it is worse than no number". Read the value from the daemon rather than
typing one — whoami reports it as embedder, and the id the key uses is
<provider>/<model>@<dimensions>d, e.g. openai/qwen3-embedding:0.6b@1024d.
Leaving it unset, or naming a model the daemon is not using, is what made the
2026-09-08 run's key partial.
Dataset: as above. Budget: max_tokens=4096. Judge: none — tier 2 is
judge-free. Corpus: cold; the harness clears the store before each run.
The deployment under test must have memory.reranker.enabled: false.
--tier2-only stops requesting the reranked path but cannot disable it, and it
refuses the run if recall reports a rerank happened. An LLM reranker is billed
per call and reorders between identical runs, so leaving it on is the difference
between a deterministic gate and one that fires on noise.
| Release | Items | Abilities | Context recall | Context precision | MRR | Runs | Key |
|---|---|---|---|---|---|---|---|
2026.8.8-40-g2698eb67 |
120 | all 6 | 0.9900 ±0.0036 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | 93a6e7a0729b |
2026.8.9-50-g4b343821 |
120 | all 6 | 0.9891 ±0.0036 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | 93a6e7a0729b |
2026.9.4-205-g97fb8b36e |
120 | all 6 | 0.9851 ±0.0000 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | 28fd464fb163 |
2026.9.4-207-g61634a063 |
120 | all 6 | 0.9851 ±0.0000 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | ed7ea84021e0 |
2026.9.6-117-gee40525c8 (2026.9.7) |
120 | all 6 | 0.9851 ±0.0000 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | 4a95cccf5f7a |
2026.9.8-52-g86d3af802 (2026.10.1) |
120 | all 6 | 0.9851 ±0.0000 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | 0c4971cc78ae |
2026.10.1-55-gc800c23be (2026.10.2) |
120 | all 6 | 0.9851 ±0.0000 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | dcfeaf709dc5 |
2026.10.1-77-gbb5238eb2 (2026.10.3) |
120 | all 6 | 0.9851 ±0.0000 | 0.9854 ±0.0000 | 0.9958 ±0.0000 | n=3 | 16a79559dd84 |
The eighth row is 2026.10.3's release candidate (the tagged commit
differs from it only in documentation, this row among them), measured on
2026-10-02 on the memory-benchmark deployment, same embedder, reranker off,
cold corpus: each run cleared the store, re-ingested the pinned dataset and
recorded the daemon's own revision. Every metric is identical to 2026.10.2 to
four decimal places across three deterministic runs; the key is new
(16a79559dd84) because daemon_revision moved. The cycle did not change the
retrieval or ingest path (it shipped the macOS release check, the Hermes
plugin's catalog readiness and EULA version 3), so this row is the
no-regression check, not a measurement of a change. The harness reports
retrieval_path_unverified: it trusts the daemon's reported recall method
(context-assembly) rather than observing it.
The seventh row is 2026.10.2's release candidate (the tagged commit
differs from it only in tests and documentation), measured on 2026-10-02 on
the memory-benchmark deployment, same embedder, reranker off, cold corpus.
Every metric is identical to 2026.10.1 to four decimal places across three
deterministic runs; the key is new (dcfeaf709dc5) because daemon_revision
moved. The cycle did not change the retrieval or ingest path (it added agent
administration, update-path fixes and release packaging), so this row is the
no-regression check, not a measurement of a change.
The sixth row is 2026.10.1's release build, measured on 2026-10-01 on the
memory-benchmark deployment, same embedder, reranker off, cold corpus. Every
metric is identical to 2026.9.7 to four decimal places across three
deterministic runs. The key is new (0c4971cc78ae) because
daemon_revision moved. The run was required because the cycle lowered the
minimum size of a deliberate deposit from a companion client, and this
benchmark ingests through that path. What it shows about that change: the
stored corpus is unchanged, 1,897 chunks before each run, exactly as in
2026.9.7, so the lower floor admitted nothing the old one refused on this
corpus. It is not a test of the floor itself; that is covered by the
daemon's own tests.
The fifth row is 2026.9.7's release build, measured on 2026-09-27 on the
memory-benchmark deployment. Every metric is identical to the 2026.9.4 build
that shipped, to four decimal places, across three deterministic runs. Its key
is new (4a95cccf5f7a) because daemon_revision moved.
The cycle between the two rows changed the retrieval path twice: - 2026.9.5 added recency ranking and series supersession. - 2026.9.7 added document supersession, and stores each document version whole under chunk hashes salted by artifact.
The row shows that neither moved tier-2 retrieval on this corpus. What it does NOT show: this benchmark ingests each item once into a cold store, so re-ingest supersession, the behaviour 2026.9.7 exists for, is never exercised. It is covered by the daemon's own tests and was verified live on the reference deployment, not by this table.
The fourth row is the build that shipped, re-measured on the binary actually
installed in production (2026.9.4-207-g61634a063) rather than on a
development build. Every metric is identical to the row above it, to four
decimal places — and the key nevertheless CHANGED, from 28fd464fb163 to
ed7ea84021e0, because daemon_revision moved from -205-…-dirty to -207-….
That is the repaired key doing its job: two runs of the same code on different
builds are no longer allowed to look like one measurement. Before 2026-09-17
these two rows would have shared a key.
The third row is the first with a COMPLETE comparability key, and it is
deliberately not comparable to the two above it. Until 2026-09-17 two of the
key's fields — observed_embedder and daemon_revision — were empty on every
run ever made (harness design §13.25), so the key could not distinguish one
build, or one embedding model, from another. That is why rows one and two share
93a6e7a0729b despite being different releases. Both fields are now populated,
which necessarily changes the key: 28fd464fb163 starts a new comparable set
rather than extending the old one. compare will refuse to diff across that
boundary, correctly.
What the numbers say anyway. Precision and MRR are bit-identical to both earlier rows. Recall is 0.9851, −0.0040 against the last published row — about a third of that row's own 3σ threshold of ±0.0107, so nothing here resembles a regression.
Two caveats that stop this being a clean release-over-release claim, stated rather than left for a reader to infer:
- The ingest regime differed. The self-hosted vLLM arm is down, so this run had no LLM extraction model and the stored corpus is produced deterministically by the embedder alone. That is visible in the spread: every metric has sd = 0.0000 here against ±0.0036 for both earlier rows, whose variance the 2026-08-21 note attributes to LLM-driven ingest differing per run. A −0.0040 recall difference and a collapsed spread have a common candidate cause, and this measurement cannot separate "the corpus was built differently" from "the code changed".
retrieval_path_unverifiedis set, because the run passed--accept-unverified-path. The reranker was disabled on the deployment under test (memory.reranker.enabled: false) as--tier2-onlyrequires — its model is the same unreachable vLLM — so the path exercised was plaincontext-assembly.
Re-measuring against a live extraction model is what would turn this into a comparable row; it is blocked on the benchmark arm's model endpoint, not on anything in the harness.
The second row is the first release-over-release comparison this table can actually support, and it shows no regression. Precision and MRR are identical to four decimal places; recall moved −0.0009 against a per-release sd of 0.0036 and a narrowest-defensible (3σ) threshold of ±0.0107 — about a twelfth of the smallest move that could fire without being noise. The sd itself reproduced exactly (0.0036 both times), which is the more reassuring number: it says the measurement is stable, not merely that two point estimates happened to agree.
Figures from bench memory aggregate over the three run directories, not
hand-computed. Standard deviations are sample (n−1), which is what the tool reports.
Per ability, same three runs:
| Ability | Context recall | Context precision | MRR | Items |
|---|---|---|---|---|
knowledge-update |
1.0000 ±0.0000 | 1.0000 ±0.0000 | 1.0000 ±0.0000 | 20 |
multi-session |
0.9564 ±0.0250 | 0.9667 ±0.0000 | 0.9750 ±0.0000 | 20 |
single-session-assistant |
1.0000 ±0.0000 | 1.0000 ±0.0000 | 1.0000 ±0.0000 | 20 |
single-session-preference |
1.0000 ±0.0000 | 1.0000 ±0.0000 | 1.0000 ±0.0000 | 20 |
single-session-user |
1.0000 ±0.0000 | 1.0000 ±0.0000 | 1.0000 ±0.0000 | 20 |
temporal-reasoning |
0.9833 ±0.0144 | 0.9458 ±0.0000 | 1.0000 ±0.0000 | 20 |
Where it loses. multi-session is the only ability that misses recall, and it is
also the only one whose recall varies run to run (sd 0.0250 against 0.0000 on four of
six). Questions needing evidence spread across sessions are where this system is
weakest, and the variance says the weakness is not a fixed set of items — the boundary
moves. temporal-reasoning carries the lowest precision (0.9458): it retrieves the
right documents and pads the set.
Caveats specific to this row.
- The key is PARTIAL.
our_extraction_model,answer_modelandjudge_modelwere left empty, so the run cannot be proven comparable to another — only observed to match on every field it does record. retrieval_path_unverifiedis set. The reranker was enabled on the deployment, so the observed path iscontext-assembly|context-assembly+rerank— a model is in the retrieval path and the run is therefore not deterministic by construction. This is the correct setting for reporting the system as shipped and the wrong one for a CI gate, which wants the reranker off and the RRF path proven.- n=3 on 120 items. Enough for a spread, not enough to resolve a sub-point move.
Agent quality — 2026.9.0 (first scored arm)¶
A different benchmark from the LongMemEval rows above, measuring a different
thing: not retrieval, but the decisions the control logic makes — what the
lead granted, whether roles followed their output schemas, and whether agents
called tools correctly. It runs 30 software tasks through a multi-agent
dev-pipeline swarm against an operator-reviewed answer key.
| Release | Tasks | Task success | Schema conformance | Tool-call validity | Steps with no output | Cost/task |
|---|---|---|---|---|---|---|
2026.8.9-70-g5d247f72 (shipped as 2026.9.0) |
30 | 100.0% | 0.985 | 1.000 | 9.7% | $0.29 |
Efficiency, same arm: 667,801 tokens and 62.4 tool calls per task, 0 escalations, 0 schema retries. Total spend $8.79.
Cost corrected downwards, 2026-09-19. This row first published $0.74 per task and $25.23 total. Those figures were wrong and we are the ones who found it: the observed model was absent from our pricing table, so every call was billed at the table's
defaultrate of $1.00/$3.00 per million tokens instead of the model's real $0.35/$2.75. The corrected figures are recomputed from the arm's own recorded token counts — 22,353,287 prompt and 351,933 completion — and nothing else about the run changed.The true cost may be lower still. This arm reached the model over a path that did not report prompt-cache reads, and a later pass on the same model measured 79.5% of prompt tokens served from cache at a tenth of the input rate. We do not know that share for this arm, so we have not applied it: the number above is an upper bound, stated as one.
Quality figures in this row are unaffected — pricing enters no scoring path.
Read both layers, because the first one alone flatters us. Task success is 100%, and underneath it 14 of 144 terminal steps (9.7%) produced no output at all — 8 hit the iteration cap, 4 entered a degenerate tool loop, 2 failed outright. Recovery absorbs those, which is the system working as designed, but a headline "100%" without the step figure would describe a smoother product than exists.
The model is part of the result, not a footnote¶
Model: Qwen/Qwen3.8-27B-FP8, self-hosted, one box. Every figure above is
that model's behaviour as much as the control logic's, and the difference is
large enough to change how the numbers read.
Measured over ~244,000 tool calls across two deployments: this model enters identical-repeat tool loops 26x more often than the mix of larger hosted models (0.52% of calls against 0.02%), and once nudged out of one it changes approach 36% of the time against 82%.
That is deliberate and it is the point. A self-hosted 27B on a single machine is what someone running this at home actually has. An organisation putting a large hosted model behind the same control logic sees the lower rate; these figures describe the harder case rather than the flattering one. The comparability key pins the model identity, so a future arm on a different model refuses to compare against this row rather than quietly superseding it.
What this row does NOT establish¶
- No pass or fail.
bench agent gaterefused a verdict, correctly: resolving a 5-point effect at the inherited σ=0.0604 needs 12 paired tasks and this arm has 5, which can only resolve 7.6 points. A smaller movement must be reported as inconclusive with that floor, never as "no change". - No trend. It is the first scored arm; there is nothing to compare it against. The releases before it have no agent row at all.
- No noise floor for this task set. The σ above is inherited from an earlier 3-task measurement and describes a different set. The honest σ for these 30 tasks does not exist yet, which is why the gate refuses.
- Not independently reproducible. Unlike the LongMemEval rows, which run against a public dataset anyone can fetch, the agent task set and its answer key are not published. An external reader can see the method and the figures and cannot re-run them. That is a real limitation of this row and is stated rather than left to be discovered.
Provenance¶
Both the daemon binary and the agent image are recorded by content digest in the run's arm key, and the run refuses to merge batches that disagree on any axis. Arms recorded before 2026-08-29 carry no agent-image identity at all — the executor discarded every image ID it observed, so those runs are marked untrustworthy and cannot serve as a baseline. This is the first arm whose image provenance is real.
Releases with no row, and why¶
The policy in RELEASE.md requires a row per release from 2026.8.4. It was not
followed. Rather than leave the gaps silent — which reads as "nothing regressed" —
each is stated:
| Release | Row | Why |
|---|---|---|
2026.8.4 |
none | No benchmark run was taken. Not backfillable: it would need the tagged binary rebuilt and the bench deployment's models restored to what they were, and those were changed during 2026.8.7. |
2026.8.5 |
none | As above. |
2026.8.6 |
none | As above. |
2026.8.7 |
none | No run. The reason was recorded at the time in the 2026.8.7 release notes: every outward-facing provider was disabled and all roles moved to a locally served model during that cycle, so a run could not be pinned to the baseline row's axes. |
2026.8.8 |
yes | Both axes above. The warm row was measured at 0e450f9c; the cold row at 2698eb67. |
2026.8.9 |
none at the tag | No run was taken at 2026.8.9 itself. The tree past it is measured on both memory axes above. |
2026.9.0 |
agent row only | The agent arm above was taken on the release candidate (2026.8.9-70-g5d247f72). No LongMemEval row: the memory axes were measured 50 commits earlier in the same cycle and nothing in the intervening work touches the retrieval path. Stated rather than left as a silent gap. |
2026.9.1 |
none | No run was taken. The cycle is bug fixes and the forge re-review feature; nothing in it touches ingestion, embedding, retrieval or the agent harness, so the axes stand where 2026.8.9-50-g4b343821 (memory, both axes) and 2026.8.9-70-g5d247f72 (agent arm) left them. A row measured on an unchanged path would add a data point without adding evidence, and the cost of a 120-item n=3 run is not free. Stated rather than left as a silent gap. |
2026.9.2 |
none | No run was taken. The cycle moves the agent loop's eleven filesystem/git tools from bash into a Go helper with behaviour pinned by a 64-case golden, and adds persistence and replay instruments; it does not touch ingestion, embedding, retrieval or the harness's scoring path, so the axes stand where the rows above left them. The agent arm should be re-measured on the Go helper before the next loop slice moves more of it — that is filed as work, not claimed here. Stated rather than left as a silent gap. |
2026.9.3 |
none | No run was taken. The cycle is the update path and the pulled agent image; it does not touch ingestion, embedding, retrieval or the harness. Stated late — this row was added with 2026.9.4's — rather than left as a silent gap. |
2026.9.4 |
none | No run could be taken: the benchmark arm's model endpoint has been unreachable since 2026-09-04 (it now accepts and immediately resets, which a port check reads as up). The cycle does not touch ingestion, embedding, retrieval or the scoring path; it does change what a multi-system-step workflow's agent receives, which is not measured here. Stated rather than left as a silent gap. |
2026.9.5 |
none — a gap, not a waiver | No run was taken, and this cycle DID change the retrieval path: recall began ranking by recency and letting a recurring series supersede its earlier members. It shipped unmeasured on the axes above. Stated late — this row was added with 2026.9.7's, whose run measures the tree that includes it — rather than left as a silent gap. |
2026.9.6 |
none | No run was taken. The cycle bounds agent containers and fixes the adoption board; the board reads the retrieval and ingest audit ledgers but changes neither path. Stated late, with 2026.9.7's, rather than left as a silent gap. |
2026.9.7 |
yes | The memory open axis above, on the release build (2026.9.6-117-gee40525c8, n=3). The cycle changes ingestion and retrieval, so a waiver was not open to it. The v9 agent-harness arms run on the slow-hardware track and are reported in the release notes, not in this table. |
2026.9.8 |
none | No run was taken. The patch changes only the companion endpoint's reply to a GET stream request, plus test fixtures. It touches no ingestion, embedding, retrieval or scoring path, so the axes stand where 2026.9.7's measured row left them. Stated rather than left as a silent gap. |
2026.10.1 |
memory row only | The memory open axis above, on the release build (n=3). The cycle lowers the minimum size of a deliberate memory deposit from a companion client, and this benchmark ingests through that path, so a waiver was not open to it. No agent row. The cycle changes the agent harness substantially (the router step, the retry ladder, the prompt-token budget and the tool-result cap), so this is a gap, not a waiver. The agent arm's model is the self-hosted Qwen3.8-27B on the endpoint that has been unreachable since 2026-09-04, and the comparability key pins the model, so a run on any other model would start a new table that cannot show a regression against the 2026.9.0 row. A substitute arm on a locally served 20B model was considered and not run: one arm is about 20M prompt tokens, which on the reference host's integrated GPU is days, not hours. The changes are covered instead by regression tests that replay each incident, and by an end-to-end lane that drives a front-end agent through the broker on a locally served model. |
A missing row stated as missing is honest. An absent row is not — which is the whole reason this section exists.
LongMemEval — judged accuracy (tier 1) and the judge's noise floor¶
A judge with an unmeasured noise floor is not a standard. Before any judged number can be gated or quoted as a win, the variance of the judge itself has to be known — otherwise a "win" is indistinguishable from the judge disagreeing with itself.
Ten identical runs, same items, same models, warm corpus:
| Metric | Mean | sd | Min | Max | Narrowest defensible gate (3σ) |
|---|---|---|---|---|---|
| judged accuracy | 0.8241 | 0.0340 | 0.7500 | 0.8750 | ±0.1020 |
| context recall | 0.9764 | 0.0059 | 0.9722 | 0.9861 | ±0.0176 |
| context precision | 0.9410 | 0.0000 | 0.9410 | 0.9410 | exact |
| MRR | 0.9917 | 0.0108 | 0.9792 | 1.0000 | ±0.0323 |
n=10 runs × 24 items, temporal-reasoning, comparability key a2fd4d245b21.
Judge: gemma4:31b. Answerer: gpt-oss:120b — a different model family, so the
judge is not grading its own output. Both open-weight and version-pinned, served from
a single provider (Ollama Cloud). A multi-provider router was deliberately avoided:
switching backends between calls would inject variance into the very number being
measured.
What this means in practice. Judged accuracy moved between 0.750 and 0.875 with nothing changing but the models' own sampling — 3 of 24 items flipping verdict. A tier-1 gate would therefore need a ±10.2 point threshold to avoid firing on noise, which is wider than most real improvements. That is the entire argument for leading with tier 2: on the same runs, context precision was exactly constant.
Note that context recall is not deterministic here (sd 0.0059) although it is on the tier-2 rows above. These judged runs measure the system as shipped, with the LLM reranker in the retrieval path; the tier-2 rows measure the RRF path, which has no model in it. Same product, two paths, and only one of them can carry an exact-equality gate.
Reproducing the judge-variance figure¶
export VORNIK_BENCH_LLM_URL=https://ollama.com/v1
export VORNIK_BENCH_LLM_KEY=<your key>
for i in $(seq 1 10); do
vornikctl bench memory run --system vornik --dataset longmemeval \
--dataset-path ./longmemeval_oracle.json \
--database <your-bench-db> --i-know-this-wipes <your-bench-db> \
--answer-model gpt-oss:120b --judge-model gemma4:31b \
--category temporal-reasoning --max-items 24 --max-tokens 4096 \
--run-dir ./jv-$i
done
vornikctl bench memory aggregate ./jv-*
Roughly 4.5 minutes per run on the reference deployment; the first run is much slower because it ingests the corpus.
Comparison: hindsight 0.9.0, same items, same budget¶
Measured under identical conditions on the same machine, each system as it ships — both with their own rerankers active. Recorded because a comparison that only flatters us is worth less than one that is obviously honest.
| System | Context recall | Context precision | MRR | Runs |
|---|---|---|---|---|
| Vornik | 1.0000 ±0.0000 | 0.9444 ±0.0000 | 0.9444 ±0.0481 | n=3 |
| hindsight 0.9.0 | 1.0000 ±0.0000 | 0.6806 ±0.0833 | 0.7292 ±0.1250 | n=4 |
Both arms are cold here — database truncated / bank deleted before every run — so the two are measured the same way. That is why Vornik's MRR reads 0.9444 ±0.0481 in this table and 1.0000 ±0.0000 in the release row above: the release row repeats against a warm corpus, as the reproduction script does. Comparing a warm number with a cold one would flatter whichever arm was warm.
Recall is tied at 1.000. Both systems retrieved every gold document on all six
items, so this subset does not separate them on recall — it is the oracle haystack,
which carries about three sessions per item where the full longmemeval_s carries
38–62. Expect recall to separate the systems on the larger haystack, and treat a tie
on an easy subset as a statement about the subset.
Vornik's retrieved set is tighter (precision 0.944 vs 0.681) and better ordered (MRR 0.944 vs 0.729). hindsight reaches the same recall by returning more.
One difference matters more than the scores. Vornik's tier-2 figures are deterministic: recall and precision have standard deviation exactly 0.0000 across repeated runs, because its ingest path is chunk-and-embed with no model deciding what gets stored. hindsight's precision and MRR vary run to run (sd 0.083 and 0.125, with MRR ranging 0.667–0.917 across four runs) because its ingest consolidates memories through an LLM, so the corpus itself differs each time.
That has a practical consequence: a single hindsight run cannot be quoted, and any comparison against it needs repeated runs and a spread. Vornik can be gated on exact equality; hindsight would need a ±0.375 MRR tolerance to avoid firing on its own noise.
Method and honest caveats¶
- n=6, one of six abilities. A smoke-scale result. It exists to prove the pipeline end to end and to establish a baseline, not to rank products.
- Both arms are marked
retrieval_path_unverifiedin their comparability keys: each system was measured with its own reranker in the path, so neither run establishes a deterministic retrieval path. This is the correct setting for a product comparison and the wrong one for a CI gate. - hindsight's internal extraction model was pinned to an open-weight hosted model
(
google/gemma-4-26b-a4b-it). Its embedder and reranker were its own bundled local defaults. Vornik used its configured local embedder. - Hindsight banks are deleted between runs. Without that its corpus accumulates and precision falls run over run — the first measurement taken before this was wired read 0.806 precision against 0.639 on an identical repeat.