# Retrieval experiment results Completed real calls on 2026-09-21. Synthetic browser incidents only. | Docs | Method | Hit@1 | MRR | Recall@3 | All evidence@3 | Multi all@3 | Negative specificity | Decision accuracy | Online p50/p95 ms | |---:|---|---:|---:|---:|---:|---:|---:|---:|---:| | 24 | bm25 | 75% | 0.827 | 85.0% | 85% | 100% | 0% | 62.5% | 0.67 / 1.07 | | 24 | embedding | 85% | 0.925 | 100.0% | 100% | 100% | 50% | 62.5% | 64.42 / 88.17 | | 24 | hybrid | 100% | 1.000 | 100.0% | 100% | 100% | 100% | 100.0% | 174.71 / 296.39 | | 24 | direct | 100% | 1.000 | 100.0% | 100% | 100% | 100% | 100.0% | 172.41 / 200.96 | | 96 | bm25 | 80% | 0.853 | 82.5% | 80% | 75% | 0% | 66.7% | 1.11 / 1.61 | | 96 | embedding | 85% | 0.925 | 100.0% | 100% | 100% | 50% | 62.5% | 63.44 / 85.05 | | 96 | hybrid | 95% | 0.950 | 92.5% | 90% | 75% | 100% | 95.8% | 163.83 / 194.19 | | 96 | direct | 100% | 1.000 | 100.0% | 100% | 100% | 100% | 100.0% | 202.07 / 272.62 | | 240 | bm25 | 80% | 0.849 | 82.5% | 80% | 75% | 0% | 66.7% | 1.86 / 3.32 | | 240 | embedding | 85% | 0.925 | 100.0% | 100% | 100% | 50% | 62.5% | 61.26 / 74.91 | | 240 | hybrid | 95% | 0.950 | 92.5% | 90% | 75% | 100% | 95.8% | 161.59 / 182.08 | | 240 | direct | 100% | 1.000 | 100.0% | 100% | 100% | 100% | 100.0% | 297.41 / 328.28 | 144 live requests; 0 API failures; 850,704 reported input tokens; estimated input charge $0.03572957 at $0.042/M; free output. Not an invoice. Pinned and returned model: jev-1.13.0. 288 measured method/query/scale rows. ## Interpretation and limits BM25 and cosine rank all docs. Hybrid independently scores each of BM25 top 8 with a Noul in one request; rank descending, abstain if maximum < .5. Direct Choice includes all docs and NONE; rank doc probability descending, abstain if Choice=NONE or answer-existence Noul < .5. Ties use doc ID. No threshold tuning; default cosine .5 is arbitrary, not calibrated. Hit@1, MRR, evidence recall@3 and all-evidence@3 on 20 positive queries ignoring abstention; decision accuracy uses accepted top1 relevant or abstained negative. Direct Choice probability is mutually exclusive, not independent relevance, so multi-evidence ranking is a limited diagnostic, not proof of evidence sufficiency. Four negative queries report specificity separately. One live call per query/method/scale; no repeated trials or generated answers. Corpus scales are nested distractor growth, not independent datasets. Frozen synthetic labels were authored, not externally adjudicated. Online local times include query embedding; offline document embedding measured separately. Jev latency is HTTP roundtrip, not model-only. Failures separate and included as misses in intended-query aggregates. No retries; auth/billing errors stop calls. Token preflight conservatively reserves UTF-8 payload byte count plus 4096 overhead; reconcile actual usage after each call. Each scale uses 20 answerable queries (16 single-evidence, 4 two-evidence) and 4 negatives. Hit@1/MRR/Recall@3 ignore abstention; decision accuracy incorporates it. A correct top document for multi-evidence does not mean the full question is answerable from that one document. Metrics are macro averages; no confidence calibration is evaluated. Hybrid MRR is truncated to its eight-document shortlist; direct ranks a mutually exclusive Choice distribution. Corpus titles are included for all methods. Scales grow mostly templated cosmetic distractors, not diverse production incidents. The model may have seen similar bug patterns in training, but these particular synthetic records and labels were authored here. No answer generation, citation faithfulness, end-to-end RAG, production indexing updates, concurrency, or repeated timing trials. One local ONNX embedding baseline, not a best-possible vector stack. Local methods do not incur remote network latency. Thresholds are untuned and score spaces differ. Semantic support labels may be debatable and have no independent annotation. Results cannot establish that Jev replaces RAG. ## Reproduce Create an isolated .venv and install requirements-lock.txt. Frozen corpus.json, queries.json, orders.json, preregistered.json and freeze-manifest.json are inputs; do not rebuild them to reproduce this exact run. run_benchmark.py pins downloadable ONNX assets using embedding-repo.json and checks freeze hashes. It refuses to overwrite results.jsonl. Copy benchmark code plus frozen inputs into a fresh isolated artifact directory for a new run; provide the authorized local secret through the code's private loader (never publish the credential file). Running analyze.py again is offline and uses saved real outputs only. Document embedding cache: document-embeddings.npy; exact model asset digests: embedding-manifest.json. Dependencies: requirements-lock.txt. Sanitized requests and complete response bodies: raw/. No auth headers are saved.