文献详情 - Gdtheory理论粤军网|广东智库信息化平台

文献详细_{Journal detailed}

Web搜索结果多层聚类方法研究
Research on Multi-level Clustering for Web Search Results

下载全文在线阅读

收藏

作　　者： ; ; ; ; ;

出　　处： 《情报学报》 2011年第5期464-470,共7页

摘　　要： 为了便于用户浏览搜索引擎返回结果,本文提出了一种基于TFIDF新的文本相似度计算方法,并提出使用具有近似线性时间复杂度的增量聚类算法对文本进行多层聚类的策略。同时,提出了一种从多文本中提取关键词的策略：提取簇中的名词或名词短语作为候选关键词,综合考虑每个候选关键词的词频、出现位置、长度和文本长度设置加权函数来计算其权重,不需要人工干预以及语料库的协助,自动提取权重最大的候选关键词作为类别关键词。在收集的百度、ODP语料以及公开测试的实验结果表明本文提出方法的有效性。 In order to facilitate the browse of the search results produced by search engines,this paper proposed a TFIDF-based new method to calculate the similarity of the documents and Web search results multi-level clustering by using one-pass clustering algorithm with linear time complexity.At the same time,we proposed a strategy to extract cluster keyword from multi-texts：selected noun or noun phrase as candidate cluster keywords,and took term frequency,the position of term occurring,the length of term and text into consideration to set a weighting function to compute every words weights of the search results,then automatically extracted the weightiest candidate keyword for each cluster generated by multi-level clustering without the intervene of human and the assistance of corpus.Experimental results on Baidu,ODP corpus and user investigation show the efficient and acceptance of our algorithm.

关键词： 文本聚类多层聚类类别关键词提取加权函数

领　　域： [自动化与计算机技术] [自动化与计算机技术]

Web搜索结果多层聚类方法研究
Research on Multi-level Clustering for Web Search Results

参考文献更多+

二级参考文献更多+

引证文献更多+

二级引证文献更多+

同被引文献更多+

耦合作品文献更多+

相关文献更多+

相关作者

相关机构对象

相关领域作者

Web搜索结果多层聚类方法研究 Research on Multi-level Clustering for Web Search Results

参考文献 更多+

二级参考文献 更多+

引证文献 更多+

二级引证文献 更多+

同被引文献 更多+

耦合作品文献 更多+

相关文献 更多+

相关作者

相关机构对象

相关领域作者

Web搜索结果多层聚类方法研究
Research on Multi-level Clustering for Web Search Results

参考文献更多+

二级参考文献更多+

引证文献更多+

二级引证文献更多+

同被引文献更多+

耦合作品文献更多+

相关文献更多+