《電子技術應用》
您所在的位置:首頁 > 其他 > 设计应用 > 基于梯度优化的大语言模型后门识别探究
基于梯度优化的大语言模型后门识别探究
网络安全与数据治理
陈佳华1,陈宇2,曹婍3
1 电子科技大学信息与软件工程学院,四川成都610066;2 北京邮电大学计算机学院,北京100876; 3 中国科学院计算技术研究所智能算法安全重点实验室,北京100190
摘要: 随着大语言模型的流行并且应用在越来越多的领域,大语言模型的安全问题也随之而来。 通常训练大语言模型对数据集以及计算资源有着极为苛刻的要求,所以有使用需求的用户大部分都直接利用网络上开源的数据集以及模型,这给后门攻击提供了绝佳的温室。后门攻击是指用户在模型中输入正常数据时模型表现像没有注入后门时一样正常,但当输入带有后门触发器的数据时模型输出异常。防止后门攻击的有效方法就是进行后门识别。目前基于梯度的优化方法是比较常用的,但使用这些方法时内部影响因子的设定对识别效果具有一定影响。文章就词令牌数量、最邻近数量、噪声大小进行了实验测量和作用机制的分析,以便为后续使用这些方法的研究者提供参考。
中圖分類號:TP309文獻標識碼:ADOI:10.19358/j.issn.2097-1788.2023.12.003
引用格式:陳佳華,陳宇,曹婍.基于梯度優化的大語言模型后門識別探究[J].網絡安全與數據治理,2023,42(12):14-19.
Research on gradient optimization based backdoor identification of large language model
Chen Jiahua1,Chen Yu 2,Cao Qi3
1 School of Information and Software Engineering,University of Electronic Science and Technology of China,Chengdu 610066, China; 2 School of Computer Science,Beijing University of Posts and Telecommunications, Beijing 100876, China; 3 CAS Key Laboratory of AI Security, Institute of Computing Technology, Chinese Academy of Sciences, Beijing 100190, China
Abstract: With the popularity of large language models (LLM) and their application in more fields, the security concerns of large language models also arise. In general, training LLM has extremely demanding requirements for datasets and computing resources, so most users who need to use them directly use opensource datasets and models on the Internet, which provides an excellent greenhouse for backdoor attacks. A backdoor attack is when a user enters normal data into the model as if it were not injected with a backdoor, but the model output is abnormal when data with a backdoor trigger is input. An effective way to prevent backdoor attacks is to perform backdoor identification. At present, gradientbased optimization methods are commonly used, but the setting of internal impact factors has a great impact on the recognition effect when using these methods. In this paper, the word token length, the number of nearest neighbors, and the noise scale are measured experimentally and the mechanism of action is analyzed, so as to provide reference for researchers who use these methods in the future.
Key words : large language models; backdoor attack; gradient based backdoor identification; impact factor

引言

近年來,大語言模型越來越多地運用在了人們的日常生活中,也誕生了很多著名的模型比如ChatGPT、GPT4[1]、LLaMA[2]等。這些模型能夠進行廣泛的任務如文本總結、情感分析等,有研究表明大模型具有小模型沒有的能力[3],如推理能力等。大語言模型也成為現在研究的熱點之一。但任何事物都有它的兩面性。大語言模型的訓練需要有足夠且良好的訓練數據集,且由于其龐大的參數量,對計算資源的需求也極高。例如GPT35具有1 750億的參數量,使用數據集達到了45 TB的大小[4]。在大部分情況下,使用者可能會選擇直接使用網絡上開源的大模型來進行下游任務的完成,或者使用領域特定數據集在開源大模型的基礎上進行微調從而定制化領域特定模型。在這種大環境下,開源大模型如果存在安全問題將造成嚴重的危害。


作者信息

陳佳華1,陳宇2,曹婍3

(1 電子科技大學信息與軟件工程學院,四川成都610066;2 北京郵電大學計算機學院,北京100876;

3 中國科學院計算技術研究所智能算法安全重點實驗室,北京100190)


文章下載地址:http://m.tom3567.com/resource/share/2000005871



weidian.jpg

此內容為AET網站原創,未經授權禁止轉載。
主站蜘蛛池模板: 99久久自偷自偷国产精品不卡| 国产精品一区二区你懂得| 精品少妇人欧美激情在线观看| 色综合色综合网色综合| 久久精品国产理论片免费| 色综合久久久久久久久五月| 日韩中文在线中文网三级| 日韩在线一区二区三区免费视频| 日韩一级特黄毛片| 久久亚洲中文字幕无码| 日韩中文字幕网址| 欧美一区二视频在线免费观看| 久久综合狠狠综合久久综青草| 一区二区三区在线观看www| 国产精品高潮视频| 久久99久久亚洲国产| 日本久久久精品视频| 日韩在线免费视频V| 豆国产97在线| 国产精品美女在线| 国产精品视频久久| 欧美亚洲国产日韩2020| 久久国产精品99久久久久久丝袜 | 国产亚洲精品网站| 午夜精品99久久免费| 色综合久久久久久中文网| 久久99中文字幕| 久久99精品久久久水蜜桃| 日本一区二区三区精品视频| 亚洲欧洲精品在线观看| 91国内揄拍国内精品对白| 俺也去精品视频在线观看| 国产精品久久久久久久久久99 | 国产精品美女久久久免费| 久久国产精彩视频| 欧美日韩无遮挡| 日本不卡一区二区三区在线观看| 日韩福利视频| 国产精品一区二区在线观看| 欧美精品在线一区| 日韩一级片一区二区|