《電子技術(shù)應(yīng)用》
您所在的位置:首頁(yè) > 人工智能 > 设计应用 > 一种基于状态预测的多线程数据过滤算法
一种基于状态预测的多线程数据过滤算法
电子技术应用
杨嘉佳,李正,郑儿,姚旺君,赵静,关健
中国电子信息产业集团有限公司第六研究所
摘要: 数据过滤算法在大数据处理领域有着重要的作用。基于正则表达式匹配技术的数据过滤算法凭借强大的特征表达能力适合于处理大规模复杂数据。然而,传统的正则表达式匹配过程为串行匹配,造成性能低,无法满足现代数据处理的需求。针对传统正则表达式匹配性能低的问题,提出一种基于多线程和状态预测的正则表达式加速匹配算法,称之为μFA:基于向量指令执行字符值比较,获取可直接跳过的信任字符数。同时,基于多线程加速和状态猜测技术,实现字符串的分段匹配处理,通过圈定字符危险区域,研判各分段最终匹配结果的正确性。实验结果表明,μFA算法的吞吐率是原始DFA算法的10.12~91.36倍、ßFA算法的1.08~2.97倍。
中圖分類號(hào):TP391.1 文獻(xiàn)標(biāo)志碼:A DOI: 10.16157/j.issn.0258-7998.245321
中文引用格式: 楊嘉佳,李正,鄭兒,等. 一種基于狀態(tài)預(yù)測(cè)的多線程數(shù)據(jù)過(guò)濾算法[J]. 電子技術(shù)應(yīng)用,2024,50(12):87-91.
英文引用格式: Yang Jiajia,Li Zheng,Zheng Er,et al. An accelerated regular expression matching algorithm based on multi-threading and state prediction[J]. Application of Electronic Technique,2024,50(12):87-91.
An accelerated regular expression matching algorithm based on multi-threading and state prediction
Yang Jiajia,Li Zheng,Zheng Er,Yao Wangjun,Zhao Jing,Guan Jian
The Sixth Research Institute of China Electronics Corporation
Abstract: Data filtering algorithms play a crucial role in the field of big data processing. Data filtering algorithms based on regular expression matching technology are suitable for processing large-scale complex data due to their powerful feature expression capabilities. However, the traditional regular expression matching process is serial matching, resulting in low performance that cannot meet the needs of modern data processing. To address the issue of low performance in traditional regular expression matching, an accelerated regular expression matching algorithm based on multithreading and state prediction is proposed, named μFA. This algorithm performs character value comparison based on vector instructions to obtain the number of trusted characters that can be skipped directly. Simultaneously, it utilizes multithreading acceleration and state prediction techniques to achieve segmented matching processing of strings. By delimiting dangerous character regions, it determines the correctness of the final matching results for each segment. Experimental results show that the throughput is 10.12 to 91.36 times higher than the original DFA algorithm and 1.08 to 2.97 times higher than the ßFA algorithm.
Key words : regular expression matching;state prediction;data filtering

引言

在人工智能時(shí)代[1],正則表達(dá)式匹配技術(shù)有助于數(shù)據(jù)的預(yù)處理過(guò)濾,可為業(yè)務(wù)應(yīng)用提供更高質(zhì)量的數(shù)據(jù)。例如,正則表達(dá)式規(guī)則由于其展現(xiàn)出強(qiáng)大的表征能力,可從大規(guī)模數(shù)據(jù)中過(guò)濾出復(fù)雜且符合深度學(xué)習(xí)模型要求的數(shù)據(jù),提升模型的推理精度。

數(shù)據(jù)預(yù)處理吞吐率是衡量過(guò)濾算法的重要性能因素之一,反映出在特定環(huán)境下算法可以運(yùn)行的性能極限,決定其是否適用于高性能大數(shù)據(jù)預(yù)處理領(lǐng)域。因此,本文重點(diǎn)研究如何提高基于正則表達(dá)式匹配的數(shù)據(jù)過(guò)濾性能。

當(dāng)前,已涌現(xiàn)出許多優(yōu)秀的基于正則表達(dá)式技術(shù)的數(shù)據(jù)過(guò)濾算法[2],包括基于非確定型有限自動(dòng)機(jī)(Nondeterministic Finite Automata, NFA)、基于確定型有限自動(dòng)機(jī)(Deterministic Finite Automata, DFA)和基于混合自動(dòng)機(jī)(Hybrid Finite Automata, Hybrid-FA)等實(shí)現(xiàn)方式。其中,因DFA的數(shù)據(jù)過(guò)濾性能較為穩(wěn)定,備受研究人員和開發(fā)人員的青睞。

然而,現(xiàn)有的正則表達(dá)式過(guò)濾算法性能較低,無(wú)法滿足大數(shù)據(jù)背景下的高性能過(guò)濾需求。因此,本文提出一種基于狀態(tài)預(yù)測(cè)的多線程數(shù)據(jù)過(guò)濾算法:通過(guò)向量指令字符值比較、多線程加速、狀態(tài)猜測(cè)等技術(shù),實(shí)現(xiàn)字符串的分段匹配處理,從而提高算法的吞吐率。


本文詳細(xì)內(nèi)容請(qǐng)下載:

http://m.tom3567.com/resource/share/2000006254


作者信息:

楊嘉佳,李正,鄭兒,姚旺君,趙靜,關(guān)健

(中國(guó)電子信息產(chǎn)業(yè)集團(tuán)有限公司第六研究所,北京 100083)


Magazine.Subscription.jpg

此內(nèi)容為AET網(wǎng)站原創(chuàng),未經(jīng)授權(quán)禁止轉(zhuǎn)載。
主站蜘蛛池模板: 91精品视频观看| 一区二区三区四区视频在线观看| 激情小说综合网| 日本一区视频在线| 欧美亚洲日本黄色| 欧美激情网站在线观看| 激情五月婷婷六月| 99在线精品免费视频| 亚洲狠狠婷婷综合久久久| 欧美综合国产精品久久丁香| 久久精品视频91| 国产精品嫩草视频| 日韩欧美亚洲精品| 国产日韩欧美在线播放| 久久精精品视频| 国产精品区免费视频| y97精品国产97久久久久久| 日韩精品资源| 国产综合欧美在线看| 亚洲欧美综合一区| 精品日本一区二区三区在线观看| 国产精品人成电影在线观看| 欧美在线亚洲在线| 国产不卡av在线| 久久精品久久久久| 日韩色av导航| 国产福利久久精品| 狠狠干视频网站| 免费在线观看一区二区| av在线亚洲男人的天堂| 国模无码视频一区二区三区| 欧美日韩一区二区三区在线观看免| 久久国产精品亚洲va麻豆| 无码免费一区二区三区免费播放| 麻豆一区二区三区在线观看| 欧美日韩一区二区三| 欧美亚洲日本黄色| 欧美激情国产精品日韩| 欧美亚洲日本网站| 日韩在线三区| 欧美一级片久久久久久久|