R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Yuan, Tongxin; He, Zhiwei; Dong, Lingzhong; Wang, Yiming; Zhao, Ruijie; Xia, Tian; Xu, Lizhen; Zhou, Binglin; Li, Fangqi; Zhang, Zhuosheng; Wang, Rui; Liu, Gongshen

Computer Science > Computation and Language

arXiv:2401.10019 (cs)

[Submitted on 18 Jan 2024 (v1), last revised 18 Feb 2024 (this version, v2)]

Title:R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Authors:Tongxin Yuan, Zhiwei He, Lingzhong Dong, Yiming Wang, Ruijie Zhao, Tian Xia, Lizhen Xu, Binglin Zhou, Fangqi Li, Zhuosheng Zhang, Rui Wang, Gongshen Liu

View PDF HTML (experimental)

Abstract:Large language models (LLMs) have exhibited great potential in autonomously completing tasks across real-world applications. Despite this, these LLM agents introduce unexpected safety risks when operating in interactive environments. Instead of centering on LLM-generated content safety in most prior studies, this work addresses the imperative need for benchmarking the behavioral safety of LLM agents within diverse environments. We introduce R-Judge, a benchmark crafted to evaluate the proficiency of LLMs in judging and identifying safety risks given agent interaction records. R-Judge comprises 162 records of multi-turn agent interaction, encompassing 27 key risk scenarios among 7 application categories and 10 risk types. It incorporates human consensus on safety with annotated safety labels and high-quality risk descriptions. Evaluation of 9 LLMs on R-Judge shows considerable room for enhancing the risk awareness of LLMs: The best-performing model, GPT-4, achieves 72.52% in contrast to the human score of 89.07%, while all other models score less than the random. Moreover, further experiments demonstrate that leveraging risk descriptions as environment feedback achieves substantial performance gains. With case studies, we reveal that correlated to parameter amount, risk awareness in open agent scenarios is a multi-dimensional capability involving knowledge and reasoning, thus challenging for current LLMs. R-Judge is publicly available at this https URL.

Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2401.10019 [cs.CL]
	(or arXiv:2401.10019v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2401.10019

Submission history

From: Tongxin Yuan [view email]
[v1] Thu, 18 Jan 2024 14:40:46 UTC (998 KB)
[v2] Sun, 18 Feb 2024 03:01:11 UTC (698 KB)

Computer Science > Computation and Language

Title:R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators