Story
arxiv_agent_systems_research ยท Oct 1, 2026 ยท paper
arxiv.orgOct 1, 2026
original source linked
In brief
LLMs are increasingly applied to cybersecurity workflows, where they are expected to translate analysts' intent into tool invocations. However, existing evaluations focus on knowledge-based assessments or end-to-end a...
Feed lens
agenticevaluation