The counter-metric
A wrong hit is the failure that grows with context size, because one pooled embedding of a huge prompt is lossy. --verify-sample 0.1 re-asks the model for a tenth of hits and compares the served answer to a fresh one.
recall sits in front of any OpenAI- or Anthropic-compatible endpoint and asks one question of every prompt: have we already answered something close enough to this? If yes, it replays the stored completion and the model is never called. If no, the call goes through once and the answer is kept. One static binary, no vector database, no Python.
Apache-2.0. v0.2.0, an alpha, built from source. Every measured number here comes from a command that ships in the binary.
A full hit is four steps in one process: an exact-hash shortcut, embed,
nearest-neighbour search, and the threshold decision. recall bench
times the two terms that matter separately, and they are not close.
111.5 µs prompt to vector 1.6 µs search and decide 0 network hops
That ratio is the architectural bet. A typical semantic cache pays a round trip to an embedding service and a round trip to a vector database. recall runs a static embedder and the index in the same process, so the only thing left to speed up is the embedder, and the only thing left to choose is the threshold. p99 is 258 µs. There is no garbage collector to pause in the tail.
| embedder | hit-rate | embed p50 | lookup p50 | lookup p99 |
|---|---|---|---|---|
| hash-v1, the default stub | 50.0% | 10.6 µs | 22.8 µs | 38.1 µs |
| model2vec/potion-base-8M | 100.0% | 111.5 µs | 113.1 µs | 257.5 µs |
The default embedder is a blake3 stub. It is fast because it captures no
meaning: on the same workload it reaches 50% hit-rate against the real
embedder's 100%, and that 50% is only the exact-match half. Speed and
quality are set by the same choice, so the number worth quoting is the
slower one. Measured on a dev laptop, 2,000 iterations, nothing on the
network; the shape is what carries, and recall bench reproduces it.
The commonly documented cutoff, cosine 0.8, is wrong across a real embedding space, because different namespaces have different densities. The cost of being wrong is not symmetric either: a missed hit costs one model call, and a false hit serves somebody the wrong answer.
So the adaptive policy does not take a similarity number at all. It takes a
false-hit budget, learns the distribution of similarities that were
confirmed wrong, and places each namespace's cutoff a margin above that
distribution, with a guard so it never climbs into the cluster feedback has
confirmed correct. It is off by default. recall-eval runs the
fixed 0.8, the best fixed cutoff, and the adaptive one over a workload whose
density you control, and reports the hit-rate each reaches at the same
false-hit budget.
A wrong hit is the failure that grows with context size, because one pooled embedding of a huge prompt is lossy. --verify-sample 0.1 re-asks the model for a tenth of hits and compares the served answer to a fresh one.
That check means something only when the model would agree with itself. A non-zero mismatch rate says the cutoff is too loose for this traffic: tighten --tau, or verify on every hit.
--ambiguity-eps serves a miss when the top two neighbours sit within a hair of each other but store different answers. Two near-duplicates with the same answer still replay. Refusals are counted, so lost hits are never invisible.
How fast a hit is has nothing to do with what you save. Savings are hit-rate times tokens times price, so on unique, large-context traffic they are close to zero however quick the cache is. Measure the hit-rate on your own log first. Everything else on this page matters only after that.
recall replay gives you yours. Then it is your numbers and the
page's arithmetic. A wrong hit is not a saving, which is why the net line exists. Input and output are priced apart because output costs three to five
times input and a long-context workload is almost all input. The hit-rate is the only
field you cannot guess: recall replay measures it, below.
The proxy prices a replay itself and never guesses a rate: prices come from
a [pricing.<model>] table you write, in dollars per million
tokens, and a model with no rule is counted as unpriced rather than
as nothing saved. With --verify-sample the same report prints the
figure discounted by the wrong hits it found. It also remembers what it paid
to fill the cache, so the two numbers side by side are the instance's
payback.
This is whole-prompt to whole-response reuse: the saving comes from not making the call when a request recurs. It does not discount a unique million-token prompt that still needs a fresh answer; that is the provider's own prefix caching, a different mechanism. The two stack. They do not substitute.
Brute force is exact but linear. The optional pure-Rust HNSW index is sublinear and 14 to 25 times faster at 50,000 entries. In v0.1.1 its accuracy gate, recall@1 of at least 0.98, held at 32 dimensions and collapsed at 256, which is exactly the width of the bundled embedder. This is the table the binary printed.
| v0.1.1, dims | corpus | brute p50 | hnsw p50 | speedup | recall@1 |
|---|---|---|---|---|---|
| 32 | 10,000 | 566 µs | 85 µs | 6.66× | 0.9960 |
| 32 | 50,000 | 4,470 µs | 181 µs | 24.64× | 0.9860 |
| 256 | 10,000 | 1,563 µs | 400 µs | 3.91× | 0.5900 |
| 256 | 50,000 | 10,607 µs | 733 µs | 14.48× | 0.2560 |
The cause was neighbour selection: keeping the closest candidates wires
near-duplicate edges and starves the graph of the long-range links routing
needs in high dimensions. In v0.2.0 the insert and prune use the paper's
diversity heuristic and the defaults move to m=32,
ef_construction=400, ef_search=256, all three now
flags. Re-measured on the same adversarial workload:
| v0.2.0, dims 256 | corpus | ef_search | hnsw p50 | speedup | recall@1 |
|---|---|---|---|---|---|
| diversity heuristic | 10,000 | 128 | — | — | 0.9580 |
| diversity heuristic, the new default | 10,000 | 256 | — | — | 0.9980 |
| diversity heuristic | 50,000 | 128 | 6,890 µs | 11.44× | 0.6660 |
| diversity heuristic, the new default | 50,000 | 256 | 12,270 µs | 4.77× | 0.8760 |
So the gate holds at 10,000 entries with margin and still falls short at
50,000, where recall decays with corpus size. Above that, raise
ef_search toward 512 or 1,024, which is still several times faster
than the brute scan at that size, raise m, or stay on brute force until you
have measured your own embeddings. These are random unit vectors, the worst
input any ANN index can be given; real embeddings cluster and score
materially higher. Latencies in the second table were taken on a loaded CI
box and are noisy; the recall figures are exact.
Correctness is unaffected unless you opt in with --index hnsw. A
missed neighbour is never a wrong answer, because the threshold still gates
what is served; it is a lost hit, and so a lost saving. Before enabling it at
scale, run recall ann-bench for your own dimensions and corpus and
raise ef_search until recall@1 holds. The measurement that
produced both tables ships in the binary.
Embed, search, decide, hit or miss: the whole loop runs, and the proxy caches
both OpenAI /v1/chat/completions and Anthropic /v1/messages,
each in its own namespace, streaming and not.
redb store that rebuilds the index on restart, so hits survive a restart, not just the bytesrecall calibrate from your own labelled pairs[pricing] table, gross and net of sampled wrong hitsrecall status, one screen per proxy; recall export, one versioned JSON documentm and ef as flagsBuilt from source. There is no published crate, because the name recall on crates.io belongs to an unrelated flashcard tool, and no prebuilt binary or image yet.
potion-base-8M scores 51.08 on MTEB against all-MiniLM-L6-v2's 55.93, about 91% of the quality at roughly 8 MB and microsecond encodes. The whole latency story rests on that trade, and it is a real trade.
Competitor figures in the project's benchmark document are published numbers, not runs on this hardware. They frame the architecture, removing the embedding round trip and the vector-database round trip, and claim no head-to-head win.
Every measured number on this page comes from a command that ships in the binary, and the two figures you can drag are labelled illustrative. Run the commands on your own traffic. The answer may be that a semantic cache is not what your workload needs, and that is the right answer to have before installing one.