Story
arxiv_cs_lg ยท Oct 2, 2026 ยท paper
arxiv.orgOct 2, 2026
original source linked
In brief
Security benchmarks for LLM-based agents often report the attack success rate (ASR) as a measure of model robustness and use these scores to compare different models and defense mechanisms, assuming that they describe...
Feed lens
agentevaluation