A Survey of Zero-Shot Sensitive Information Detection Techniques based on Large Language Models
DOI:
https://doi.org/10.54691/3kzsns60Keywords:
Large Language Models; Zero-Shot Learning; Sensitive Information Detection; Named Entity Recognition; Privacy Protection; Prompt Learning.Abstract
With the rapid growth of digital information, the risk of sensitive information leakage in textual data, including personally identifiable information, medical privacy, financial data, and corporate confidential information, has become increasingly prominent. Traditional sensitive information detection methods, which mainly rely on rule matching, supervised learning, and manual annotation, struggle to meet the requirements of identifying diverse, open-domain, and dynamically evolving sensitive information. In recent years, Large Language Models (LLMs) have provided a new technical paradigm for zero-shot sensitive information detection without annotated data, owing to their powerful semantic understanding, contextual reasoning, and knowledge transfer capabilities. Through prompt learning, in-context learning, and instruction-driven information extraction approaches, LLMs can achieve flexible sensitive information identification in scenarios involving unknown sensitive categories and cross-domain applications. However, LLMs themselves introduce new security risks, including training data leakage, privacy memorization, and prompt injection attacks, posing significant challenges to sensitive information detection technologies. This paper presents a systematic survey of the development of LLM-based zero-shot sensitive information detection techniques. First, it introduces the development path of sensitive data detection technology, including rule-based methods, older machine-learning techniques, and currently popular pre-trained language models. Then it introduces the research content of LLM-driven zero-shot named entity recognition, open information extraction and privacy detection. List the privacy-leakage risks and corresponding defence measures for current applications of LLMs. Finally, this paper presents some future research directions for the above work and provides a path for the development of efficient, secure and trustworthy intelligent sensitive information detection systems.
Downloads
References
[1] Vaswani, A., Shazeer, N., Parmar, N., et al. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008).
[2] Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre training of deep bidirectional transformers for language understanding. In Proceedings of NAACL HLT (pp. 4171–4186).
[3] Liu, Y., Ott, M., Goyal, N., et al. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv preprint arXiv:1907.11692.
[4] Raffel, C., Shazeer, N., Roberts, A., et al. (2020). Exploring the limits of transfer learning with a unified text to text transformer. Journal of Machine Learning Research, 21, 1–67.
[5] Brown, T. B., Mann, B., Ryder, N., et al. (2020). Language models are few shot learners. In Advances in Neural Information Processing Systems (Vol. 33, pp. 1877–1901).
[6] Carlini, N., Liu, C., Erlingsson, U., Kos, J., & Song, D. (2019). The secret sharer: Evaluating and testing unintended memorization in neural networks. In Proceedings of the 28th USENIX Security Symposium (pp. 267–284).
[7] Nasr, M., Carlini, N., Hayase, J., et al. (2023). Membership inference attacks against language models. arXiv preprint arXiv:2302.03055.
[8] Lample, G., Ballesteros, M., Subramanian, S., Kawakami, K., & Dyer, C. (2016). Neural architectures for named entity recognition. In Proceedings of NAACL HLT (pp. 260–270).
[9] Uzuner, O., Luo, Y., & Szolovits, P. (2007). Evaluating the state of the art in automatic de identification of medical records. Journal of the American Medical Informatics Association, 14(5), 550–563. https://doi.org/10.1197/jamia.M2427
[10] Garbade, F., Jurgens, D., & Cohan, A. (2021). A survey on text de identification: Techniques, datasets and challenges. arXiv preprint arXiv:2101.08327.
[11] Lester, B., Al Rfou, R., & Constant, N. (2021). The power of scale for parameter efficient prompt tuning. In Proceedings of EMNLP (pp. 3045–3059).
[12] Zhou, K., Yang, J., Loy, C. C., & Liu, Z. (2022). Learning to prompt for vision language models. International Journal of Computer Vision, 130, 2337–2348. https://doi.org/10.1007/s11263 022 01647 2
[13] Shen, Y., Liu, J., Zhang, Y., et al. (2023). PromptNER: Prompting for named entity recognition. arXiv preprint arXiv:2305.15444.
[14] Zhang, W., Li, Y., Xiao, Y., et al. (2023). UniversalNER: Targeted distillation from large language models for open named entity recognition. arXiv preprint arXiv:2308.03279.
[15] Zaratiana, U. H., Tomeh, N., Holat, P., & Charnois, T. (2023). GLiNER: Generalist model for named entity recognition using bidirectional transformer. arXiv preprint arXiv:2311.08526.
[16] Xie, T., Li, Q., Zhang, J., Liu, Z., & Wang, H. (2023). Empirical study of zero shot NER with ChatGPT. arXiv preprint arXiv:2310.10035.
[17] Xie, T., Li, Q., Zhang, Y., Liu, Z., & Wang, H. (2023). Self improving for zero shot named entity recognition with large language models. arXiv preprint arXiv:2311.08921.
[18] Villena, F., Miranda, L., & Aracena, C. (2024). llmNER: Zero/few shot named entity recognition exploiting the power of large language models. arXiv preprint arXiv:2406.04528.
[19] Zhang, N., Bi, Z., Yu, X., et al. (2024). Large language models for information extraction: A survey. arXiv preprint arXiv:2312.10458.
[20] Chen, Y., Li, Y., Zhang, N., et al. (2024). A survey on large language models for named entity recognition. arXiv preprint arXiv:2405.01212.
[21] Carlini, N., Tramer, F., Wallace, E., et al. (2021). Extracting training data from large language models. In Proceedings of the 30th USENIX Security Symposium (pp. 2633–2650).
[22] Carlini, N., Jagielski, M., Zhang, C., et al. (2023). Scalable extraction of training data from production language models. arXiv preprint arXiv:2311.17035.
[23] Nasr, M., Carlini, N., Hayase, J., et al. (2023). Membership inference attacks against language models. In Proceedings of IEEE Symposium on Security and Privacy (pp. 1–18).
[24] Zhang, Z., Zhang, A., Li, M., & Smola, A. J. (2024). Privacy leakage of large language models: A survey. arXiv preprint arXiv:2402.07090.
[25] Das, B. C., Amini, M. H., & Wu, Y. (2025). Security and privacy challenges of large language models: A survey. ACM Computing Surveys.
[26] Greshake, K., Abdelnabi, S., Mishra, S., et al. (2023). More than you've asked for: A comprehensive analysis of novel prompt injection threats to application integrated large language models. arXiv preprint arXiv:2302.12173.
[27] Zou, A., Wang, Z., Kolter, J. Z., & Fredrikson, M. (2023). Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Scientific Journal of Intelligent Systems Research

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.




