|国家预印本平台
| 注册
首页|BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Harmon Bhasin Kevin Flyangolts Dianzhuo Wang Evan Seeyave Arjun Banerjee Amanda Darling Joshua Stallings David Stern Shawn Higdon Claire Duvallet Bryan Tegomoh Kenny Workman

Arxiv_logoArxiv

BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance

Harmon Bhasin Kevin Flyangolts Dianzhuo Wang Evan Seeyave Arjun Banerjee Amanda Darling Joshua Stallings David Stern Shawn Higdon Claire Duvallet Bryan Tegomoh Kenny Workman

作者信息

Abstract

As pathogen genomic surveillance scales, the bottleneck is shifting from data generation to analysis. We present BioSecBench-Surveillance, a verifiable benchmark of 100 evaluations testing whether AI agents can infer the right analysis pipeline from raw sequencing data and surveillance context. Each evaluation gives an agent only the data and context a human analyst would have, then grades its structured answer deterministically. The tasks span seven categories, from taxonomic classification to genetic-engineering detection, across diverse sample types and sequencing technologies. Across 3,962 gradable attempts from sixteen model-harness pairs, the strongest configuration cleared only about half. Opus 4.8 with PI led at 50.2 percent, with a 95 percent confidence interval of 40.1 to 60.3 percent across 83 evaluations, tied with GPT-5.5 with Codex at 50.2 percent, with a 95 percent confidence interval of 40.8 to 59.6 percent, followed by Opus 4.7 with PI at 49.6 percent, with a 95 percent confidence interval of 40.0 to 59.2 percent, and Sonnet 4.6 with PI at 48.6 percent, with a 95 percent confidence interval of 38.9 to 58.3 percent. Even when agents invoked the correct workflows, their mistakes came from the choices around them, such as which references, thresholds, filters, and normalization to apply. BioSecBench-Surveillance provides a standard for measuring whether agents can be trusted to perform genomic surveillance when the next outbreak arrives.

引用本文复制引用

Harmon Bhasin,Kevin Flyangolts,Dianzhuo Wang,Evan Seeyave,Arjun Banerjee,Amanda Darling,Joshua Stallings,David Stern,Shawn Higdon,Claire Duvallet,Bryan Tegomoh,Kenny Workman.BioSecBench-Surveillance: A Verifiable Benchmark for AI Agents in Pathogen Genomic Surveillance[EB/OL].(2026-07-21)[2026-08-11].https://arxiv.org/abs/2607.19262.

学科分类

生物科学研究方法、生物科学研究技术
首发时间 2026-07-21
下载量:0
|
点击量:17
段落导航相关论文