The fastest model call is the one you don’t make.

recall sits in front of any OpenAI- or Anthropic-compatible endpoint and asks one question of every prompt: have we already answered something close enough to this? If yes, it replays the stored completion and the model is never called. If no, the call goes through once and the answer is kept. One static binary, no vector database, no Python.

Apache-2.0. v0.2.0, an alpha, built from source. Every measured number here comes from a command that ships in the binary.

Three prompts rise to their similarity with the nearest stored answer. Two cross the threshold and are served from memory; one falls short, goes upstream, and its answer is stored. 1.0 0.5 nearest-answer similarity prompts, as they arrive τ 0.85 to the model, then stored 0.91 0.62 0.87
  1. prompt 1, nearest 0.91served from memory, 113 µs
  2. prompt 2, nearest 0.62upstream, once; the answer is stored
  3. prompt 3, nearest 0.87served from memory
One line decides everything. Above it costs nothing; below it costs one model call. Illustrative similarities; the latency is measured.

A hit costs 113 µs, and almost all of it is the embedder

A full hit is four steps in one process: an exact-hash shortcut, embed, nearest-neighbour search, and the threshold decision. recall bench times the two terms that matter separately, and they are not close.

Full hit, 113.1 µs p50, with model2vec/potion-base-8M

embed, 111.5 µs

111.5 µs prompt to vector 1.6 µs search and decide 0 network hops

That ratio is the architectural bet. A typical semantic cache pays a round trip to an embedding service and a round trip to a vector database. recall runs a static embedder and the index in the same process, so the only thing left to speed up is the embedder, and the only thing left to choose is the threshold. p99 is 258 µs. There is no garbage collector to pause in the tail.

embedderhit-rateembed p50lookup p50lookup p99
hash-v1, the default stub50.0%10.6 µs22.8 µs38.1 µs
model2vec/potion-base-8M100.0%111.5 µs113.1 µs257.5 µs

The 22 µs number is not the product

The default embedder is a blake3 stub. It is fast because it captures no meaning: on the same workload it reaches 50% hit-rate against the real embedder's 100%, and that 50% is only the exact-match half. Speed and quality are set by the same choice, so the number worth quoting is the slower one. Measured on a dev laptop, 2,000 iterations, nothing on the network; the shape is what carries, and recall bench reproduces it.

0.8 is a magic number, and it is wrong somewhere

The commonly documented cutoff, cosine 0.8, is wrong across a real embedding space, because different namespaces have different densities. The cost of being wrong is not symmetric either: a missed hit costs one model call, and a false hit serves somebody the wrong answer.

namespace faq 100% of repeats served, 1.6% of them wrong, τ 0.800
namespace support 100% of repeats served, 11.8% of them wrong, τ 0.800
0.60similarity to the nearest stored answer1.00
Illustrative pairs, two namespaces, one cutoff. Filled plum points are repeats served correctly; rust points are lookalikes served with the wrong answer; hollow points went upstream. At 0.8 the FAQ namespace is nearly clean and the support namespace serves a wrong answer about one time in eight. The adaptive policy sets a separate τ per namespace from the similarities that proved wrong, targeting a false-hit rate you choose.

So the adaptive policy does not take a similarity number at all. It takes a false-hit budget, learns the distribution of similarities that were confirmed wrong, and places each namespace's cutoff a margin above that distribution, with a guard so it never climbs into the cluster feedback has confirmed correct. It is off by default. recall-eval runs the fixed 0.8, the best fixed cutoff, and the adaptive one over a workload whose density you control, and reports the hit-rate each reaches at the same false-hit budget.

The counter-metric

A wrong hit is the failure that grows with context size, because one pooled embedding of a huge prompt is lossy. --verify-sample 0.1 re-asks the model for a tenth of hits and compares the served answer to a fresh one.

Only at temperature 0

That check means something only when the model would agree with itself. A non-zero mismatch rate says the cutoff is too loose for this traffic: tighten --tau, or verify on every hit.

Near-ties are refused

--ambiguity-eps serves a miss when the top two neighbours sit within a hair of each other but store different answers. Two near-duplicates with the same answer still replay. Refusals are counted, so lost hits are never invisible.

$ recall replay --file traffic.jsonl --verify-sample 0.1 --upstream https://api.openai.com verify : 12 sampled, 0 mismatch (0.0% candidate false-hit), 0 unchecked

Hit-rate is the number. Latency is not.

How fast a hit is has nothing to do with what you save. Savings are hit-rate times tokens times price, so on unique, large-context traffic they are close to zero however quick the cache is. Measure the hit-rate on your own log first. Everything else on this page matters only after that.

calls not made, per day
3,000
saved per day, gross
$23.25
net of wrong hits
$22.79
per month, net
$683.55
The figures shown are a worked example, not a measurement; the 30% is a placeholder until recall replay gives you yours. Then it is your numbers and the page's arithmetic. A wrong hit is not a saving, which is why the net line exists. Input and output are priced apart because output costs three to five times input and a long-context workload is almost all input. The hit-rate is the only field you cannot guess: recall replay measures it, below.
# 1. run the proxy over your upstream, or over a mock for an isolated baseline $ recall serve --config recall.toml & # 2. replay a real request log; the proxy prices it from your own [pricing] table $ recall replay --file traffic.jsonl --target http://127.0.0.1:8080 recall replay (target: http://127.0.0.1:8080, 8 lines) est. saved : $0.0015 (in $2.5/Mtok, out $10/Mtok, un-discounted: no verify sample) requests : 8 hits : 4 hit-rate : 50.0% (over cacheable) tokens saved : 234 total, 106 input, 128 output

The proxy prices a replay itself and never guesses a rate: prices come from a [pricing.<model>] table you write, in dollars per million tokens, and a model with no rule is counted as unpriced rather than as nothing saved. With --verify-sample the same report prints the figure discounted by the wrong hits it found. It also remembers what it paid to fill the cache, so the two numbers side by side are the instance's payback.

recall is not prompt caching

This is whole-prompt to whole-response reuse: the saving comes from not making the call when a request recurs. It does not discount a unique million-token prompt that still needs a fresh answer; that is the provider's own prefix caching, a different mechanism. The two stack. They do not substitute.

The index was fast and wrong at 256 dimensions. In v0.2.0 it is mostly right.

Brute force is exact but linear. The optional pure-Rust HNSW index is sublinear and 14 to 25 times faster at 50,000 entries. In v0.1.1 its accuracy gate, recall@1 of at least 0.98, held at 32 dimensions and collapsed at 256, which is exactly the width of the bundled embedder. This is the table the binary printed.

v0.1.1, dimscorpusbrute p50hnsw p50speeduprecall@1
3210,000566 µs85 µs6.66×0.9960
3250,0004,470 µs181 µs24.64×0.9860
25610,0001,563 µs400 µs3.91×0.5900
25650,00010,607 µs733 µs14.48×0.2560

The cause was neighbour selection: keeping the closest candidates wires near-duplicate edges and starves the graph of the long-range links routing needs in high dimensions. In v0.2.0 the insert and prune use the paper's diversity heuristic and the defaults move to m=32, ef_construction=400, ef_search=256, all three now flags. Re-measured on the same adversarial workload:

v0.2.0, dims 256corpusef_searchhnsw p50speeduprecall@1
diversity heuristic10,000128——0.9580
diversity heuristic, the new default10,000256——0.9980
diversity heuristic50,0001286,890 µs11.44×0.6660
diversity heuristic, the new default50,00025612,270 µs4.77×0.8760

So the gate holds at 10,000 entries with margin and still falls short at 50,000, where recall decays with corpus size. Above that, raise ef_search toward 512 or 1,024, which is still several times faster than the brute scan at that size, raise m, or stay on brute force until you have measured your own embeddings. These are random unit vectors, the worst input any ANN index can be given; real embeddings cluster and score materially higher. Latencies in the second table were taken on a loaded CI box and are noisy; the recall figures are exact.

Which is why the default index is still brute force

Correctness is unaffected unless you opt in with --index hnsw. A missed neighbour is never a wrong answer, because the threshold still gates what is served; it is a lost hit, and so a lost saving. Before enabling it at scale, run recall ann-bench for your own dimensions and corpus and raise ef_search until recall@1 holds. The measurement that produced both tables ships in the binary.

An alpha. The loop runs end to end; the edges are being measured.

Embed, search, decide, hit or miss: the whole loop runs, and the proxy caches both OpenAI /v1/chat/completions and Anthropic /v1/messages, each in its own namespace, streaming and not.

The loop, since v0.1.1

  • Optional static model2vec/potion embedder, loaded locally, no network
  • Durable redb store that rebuilds the index on restart, so hits survive a restart, not just the bytes
  • Adaptive threshold engine, off by default
  • Streamed completions stored as raw SSE and replayed as a stream
  • recall calibrate from your own labelled pairs

New in v0.2.0

  • Savings in dollars from your own [pricing] table, gross and net of sampled wrong hits
  • recall status, one screen per proxy; recall export, one versioned JSON document
  • Savings split by namespace, and what the cache cost to fill
  • The HNSW retune above, with m and ef as flags
  • The near-tie ambiguity gate

Getting it

Built from source. There is no published crate, because the name recall on crates.io belongs to an unrelated flashcard tool, and no prebuilt binary or image yet.

$ cargo run -p recall -- bench $ cargo run -p recall --features static -- \ bench --model <potion-base-8M>

The quality trade, stated

potion-base-8M scores 51.08 on MTEB against all-MiniLM-L6-v2's 55.93, about 91% of the quality at roughly 8 MB and microsecond encodes. The whole latency story rests on that trade, and it is a real trade.

Competitor figures in the project's benchmark document are published numbers, not runs on this hardware. They frame the architecture, removing the embedding round trip and the vector-database round trip, and claim no head-to-head win.

Measure your hit-rate before you believe any of this

Every measured number on this page comes from a command that ships in the binary, and the two figures you can drag are labelled illustrative. Run the commands on your own traffic. The answer may be that a semantic cache is not what your workload needs, and that is the right answer to have before installing one.