|国家预印本平台
| 注册
首页|Proof-Valid Benchmarking for Approximate LLM Caching under Premise Erasures

Proof-Valid Benchmarking for Approximate LLM Caching under Premise Erasures

Jianfeng Xu

Proof-Valid Benchmarking for Approximate LLM Caching under Premise Erasures

Proof-Valid Benchmarking for Approximate LLM Caching under Premise Erasures

Jianfeng Xu1

作者信息

  • 1. Koguan School of Law, China Institute for Smart Justice, School of Computer Science, Shanghai Jiao Tong University
  • 折叠

摘要

Approximate caches for large language models reuse a stored answer when a new prompt is semantically close to a cached one,yet similarity alone does not establish that the evidence supporting the stored answer is still available. This paper introduces proof-validbenchmarking (PVB): a cached answer may be served only with a certificate whose premises survive a controlled erasure stress test.We map class–response pairs to response-conditioned evidence-support graphs, derive the exact residual-premise law andthe exact price of transparent, shared-module, and coded protection. General semantic-module selection is NP-complete, but a greedyrule retains a logarithmic guarantee; a certificate-based converse shows that query-only coded caching must protect a proof certificaterather than an embedding key; and a hybrid theorem under a common erasure realization gives the exact transparent–coded trade-off.Experiments close a three-layer loop: a trace-conditioned stress test on 40,000 prompts is consistent with the exact law; a grounded auditon 7,405 HotpotQA questions confirms it with measured exponents and quantifies proxy bias; and an instrumented pipeline constructscertificates at 38 μs and verifies them at 1.6 μs per question, predicting measured availability to within 0.0016. Correlated and burstyerasure delimit the i.i.d. assumption quantitatively, and PVB serving dominates the error–throughput region reachable by thresholdtuning alone. A real-dynamics audit on five monthly Wikipedia and four Wikidata snapshots measures the erasure process in vivo andaudits the serving gate end to end against statement-level ground truth.

Abstract

Approximate caches for large language models reuse a stored answer when a new prompt is semantically close to a cached one, yet similarity alone does not establish that the evidence supporting the stored answer is still available. This paper introduces proof-valid benchmarking (PVB): a cached answer may be served only with a certificate whose premises survive a controlled erasure stress test. We map classresponse pairs to response-conditioned evidence-support graphs, derive the exact residual-premise law and the exact price of transparent, shared-module, and coded protection. General semantic-module selection is NP-complete, but a greedy rule retains a logarithmic guarantee; a certificate-based converse shows that query-only coded caching must protect a proof certificate rather than an embedding key; and a hybrid theorem under a common erasure realization gives the exact transparentcoded trade-off. Experiments close a three-layer loop: a trace-conditioned stress test on 40,000 prompts is consistent with the exact law; a grounded audit on 7,405 HotpotQA questions confirms it with measured exponents and quantifies proxy bias; and an instrumented pipeline constructs certificates at 38 s and verifies them at 1.6 s per question, predicting measured availability to within 0.0016. Correlated and bursty erasure delimit the i.i.d. assumption quantitatively, and PVB serving dominates the errorthroughput region reachable by threshold tuning alone. A real-dynamics audit on five monthly Wikipedia and four Wikidata snapshots measures the erasure process in vivo and audits the serving gate end to end against statement-level ground truth.

关键词

proof-valid benchmarking/ approximate cache/ semantic cache/ large language model/ premise erasure/ erasure coding/ retrieval-augmented generation

Key words

proof-valid benchmarking/ approximate cache/ semantic cache/ large language model/ premise erasure/ erasure coding/ retrieval-augmented generation

引用本文复制引用

Jianfeng Xu.Proof-Valid Benchmarking for Approximate LLM Caching under Premise Erasures[EB/OL].(2026-08-30)[2026-09-01].https://chinaxiv.org/abs/202608.00158.

学科分类

计算技术、计算机技术
首发时间 2026-08-30
下载量:0
|
点击量:7
段落导航相关论文