Benchmark Research on Safety Value Alignment Evaluation of Open-Domain Dialogue Systems based on NLP

Authors

  • Peirong Li School of Computer and Artificial Intelligence, Hubei University of Education, Wuhan, Hubei, 430000, China

DOI:

https://doi.org/10.54691/p4zgas48

Keywords:

Value Alignment; Open-domain Dialogue System; Full-link Framework; Evaluation Benchmark; LLM Safety.

Abstract

Large language models sometimes behave in puzzling ways. They pass various safety tests, yet in multi-turn dialogues, a few carefully crafted sentences can lead them astray into making dangerous judgments. We call this "extreme value alignment failure." A review of recent research reveals an awkward situation: attack, defense, evaluation, and theoretical studies operate in isolation, with little connection among them. This fragmentation results in repeated extreme risks that remain unresolved. This paper maps the four research directions onto a unified framework---"failure mode, attack vector, defense level, evaluation benchmark"---providing a theoretical coordinate for the field and directions for future evaluation research.

Downloads

Download data is not yet available.

References

[1] Sun, H. (2023). Research on safety methods for open domain dialogue systems [Doctoral dissertation]. Tsinghua University.

[2] Lu, C., et al. (2026). The assistant axis: Situating and stabilizing the default persona of language models. arXiv preprint arXiv:2601.10387.

[3] Li, L., et al. (2024). JailBench: Assessing and eliciting harmful content in Chinese LLMs. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (pp. 567–582). Bangkok, Thailand.

[4] Jiang, L., et al. (2024). MoralBench: Moral evaluation of large language models. arXiv preprint arXiv:2406.04428.

[5] Cui, J., et al. (2024). CValues: Measuring the values of Chinese large language models. arXiv preprint arXiv:2307.09705.

[6] Zhang, Z., et al. (2024). EasyJailbreak: A unified framework for jailbreaking large language models. arXiv preprint arXiv:2403.12171.

[7] Zhang, X., et al. (2025). DAMON: A dialogue aware MCTS framework for jailbreaking large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing (pp. 112–128). Miami, FL.

[8] Liu, Y., et al. (2024). Jailbreaker: Automated jailbreaking of large language models via prompt decomposition. arXiv preprint.

[9] Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073.

[10] Hu, H., et al. (2025). Steering dialogue dynamics for robustness against multi turn jailbreaking attacks. arXiv preprint arXiv:2503.00187.

[11] Anthropic. (2024). The Claude 3 model family: Opus, Sonnet, Haiku. https://www.anthropic.com/news/claude 3 family

[12] Motnikar, L., et al. (2025). The value of Gen AI conversations: A bottom up framework for AI value alignment. In Proceedings of the 33rd European Conference on Information Systems (ECIS). Oslo, Norway.

[13] Tang, et al. (2026). Priority hacking: The vulnerability of LLM value alignment. In Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing. Miami, FL.

[14] Mazeika, M., et al. (2024). HarmBench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249.

[15] Huang, Z., Ramnath, K., Chen, Y., et al. (2026). Diffusion language model inference with Monte Carlo tree search. In Findings of the Association for Computational Linguistics: EACL 2026 (pp. 3493–3512). Rabat, Morocco.

[16] Carlini, N., et al. (2024). How many unicorns are in this image? A safety evaluation benchmark for vision language models. arXiv preprint.

[17] Zou, A., et al. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.

[18] Journal of Sichuan University. (2026). A review of black box jailbreak attacks on large language models. Journal of Sichuan University (Natural Science Edition), 63(2), 1–15.

[19] Bengio, Y., et al. (2025). International AI Safety Report 2025: Second key update: Technical safeguards and risk management. arXiv preprint arXiv:2511.19863.

[20] Hong, Y., et al. (2026). Symbolic guardrails for domain specific agents: Stronger safety and security guarantees without sacrificing utility. arXiv preprint arXiv:2604.15579.

[21] Qi, X., et al. (2024). Fine tuning aligned language models compromises safety, even when users do not intend to! arXiv preprint.

[22] Denison, C., et al. (2024). Sycophancy to subterfuge: Investigating reward tampering in large language models. arXiv preprint arXiv:2406.10162.

Downloads

Published

2026-08-26

Issue

Section

Articles

How to Cite

Li, Peirong. 2026. “Benchmark Research on Safety Value Alignment Evaluation of Open-Domain Dialogue Systems Based on NLP”. Scientific Journal of Intelligent Systems Research 8 (8): 1-8. https://doi.org/10.54691/p4zgas48.