《電子技術(shù)應用》
您所在的位置:首頁 > 人工智能 > 设计应用 > 基于单页语义特征的垃圾网页检测
基于单页语义特征的垃圾网页检测
电子技术应用
陈木生1,2,高斐1,吴俊华1
(1.江西理工大学 软件工程学院,江西 南昌 330013;2.南昌市虚拟数字工程与文化传播重点实验室,江西 南昌 330013)
摘要: 为解决垃圾网页检测中特征提取难度高、计算量大的问题,提出一种仅基于当前网页的HTML脚本提取语义特征的方法。首先使用深度优先搜索和动态规划相结合的记忆化搜索算法对域名进行单词切割,采用隐含狄利克雷分布提取主题词,基于Word2Vec词向量和词移距离计算3个单页语义相似度特征;然后将单页语义相似度特征融合单页统计特征,使用随机森林等分类算法构建分类模型进行垃圾网页检测。实验结果表明,基于单页内容提取语义特征融合单页统计特征进行分类的AUC值达到88.0%,比对照方法提高4%左右。
中圖分類號:TP391.6
文獻標志碼:A
DOI: 10.16157/j.issn.0258-7998.223376
中文引用格式: 陳木生,高斐,吳俊華. 基于單頁語義特征的垃圾網(wǎng)頁檢測[J]. 電子技術(shù)應用,2023,49(6):24-29.
英文引用格式: Chen Musheng,Gao Fei,Wu Junhua. Web spam detection based on semantic features from current page[J]. Application of Electronic Technique,2023,49(6):24-29.
Web spam detection based on semantic features from current page
Chen Musheng1,2,Gao Fei1,Wu Junhua1
(1.School of Software Engineering, Jiangxi University of Science and Technology, Nanchang 330013, China; 2.Nanchang Key Laboratory of Virtual Digital Engineering and Cultural Communication, Nanchang 330013, China)
Abstract: In order to solve the problem of high difficulty and large amount of computation in feature extraction for web spam detection, a method for extracting semantic features only based on the HTML script of the current page is proposed. Firstly, the domain name is segmented by a memorization search algorithm combining depth-first search and dynamic programming. Secondly, The latent Dirichlet distribution is used to extract subject words of the web page. Lastly, three single-page semantic similarity features are calculated based on Word2Vec and word mover distance. Combining the single-page semantic similarity features with single-page statistical features, classification algorithms such as random forest are used to build classification models for web spam detection. The experimental results show that the AUC value of single-page content extraction based on semantic and statistical features for classification reaches 88.0%, which is about 4% higher than that of the control method.
Key words : web spam detection;feature extraction;memory search;latent Dirichlet distribution;Word2Vec;word mover distance;random forest

0 引言

如今,隨著互聯(lián)網(wǎng)信息的快速增長,搜索引擎被認為是訪問網(wǎng)站的關(guān)鍵工具,其用戶占到網(wǎng)絡用戶的80%以上[1]。但是有研究表明,大約60%的用戶只查看第一頁中最初的5個結(jié)果[2]。可以看出,在搜索結(jié)果中排名靠前的網(wǎng)頁會擁有更多的訪問者,由此帶來更多的收入。由于通過正常手段提高網(wǎng)頁排名非常困難,于是某些網(wǎng)站便通過非正常手段和技術(shù)欺騙搜索引擎提高網(wǎng)頁排名,這些網(wǎng)頁被稱為垃圾網(wǎng)頁[3]。垃圾網(wǎng)頁會降低搜索結(jié)果的質(zhì)量,浪費用戶的時間,侵占搜索引擎公司和其他內(nèi)容網(wǎng)站的合法利益[4]。盡管搜索引擎公司已經(jīng)使用了各種方法來應對垃圾網(wǎng)頁,但至今為止,垃圾網(wǎng)頁檢測依然是搜索引擎需要重點突破的難題,也是學術(shù)領域的一個前沿課題。因此,高效、準確地檢測垃圾網(wǎng)頁具有重要意義。



本文詳細內(nèi)容請下載:http://m.tom3567.com/resource/share/2000005343




作者信息:

陳木生1,2,高斐1,吳俊華1

(1.江西理工大學 軟件工程學院,江西 南昌 330013;2.南昌市虛擬數(shù)字工程與文化傳播重點實驗室,江西 南昌 330013)


微信圖片_20210517164139.jpg

此內(nèi)容為AET網(wǎng)站原創(chuàng),未經(jīng)授權(quán)禁止轉(zhuǎn)載。
主站蜘蛛池模板: 久久精品无码中文字幕| 热久久视久久精品18亚洲精品| 日本一区二区三区在线视频| 久久婷婷国产精品| 亚洲综合国产精品| 国产精品久久久久影院日本 | 国产亚洲欧美在线视频| 欧美亚洲日本黄色| 亚洲午夜高清视频| 国产成人精品a视频一区www| 精品国偷自产在线| 久久久久国产精品熟女影院| 欧美亚洲国产日韩2020| 国产精品美女免费看| 日本一区二区三区在线视频| 日韩精品一区二区三区丰满 | 91精品综合久久| 国产精品高清在线观看| 亚洲a∨一区二区三区| www国产精品com| 国产99在线免费| 99视频免费观看蜜桃视频| 91国产在线免费观看| 伊人色综合久久天天五月婷| 亚洲国产精品久久久久婷婷老年| 91国内在线视频| www.久久草| 亚洲伊人久久大香线蕉av| 亚洲a在线观看| 日韩一区免费观看| 热久久免费国产视频| 欧美亚洲第一页| 久久riav| 国产成人精品999| 午夜久久久久久久久久久| 欧美中日韩在线| 国产日韩第一页v| 国产福利精品在线| 性高潮久久久久久久久| 欧美国产综合视频| 国产日韩在线观看av|