토큰 크기 및 출현 빈도에 기반한 웹 페이지 유사도

토큰 크기 및 출현 빈도에 기반한 웹 페이지 유사도

ㆍ 저자명: 이은주,정우성,Lee. Eun-Joo,Jung. Woo-Sung
ㆍ 간행물명: 한국IT서비스학회지= Journal of the Korea society of IT services
ㆍ 권/호정보: 2012년|11권 4호|pp.263-275 (13 pages)
ㆍ 발행정보: 한국IT서비스학회
ㆍ 파일정보: 정기간행물|
PDF텍스트
ㆍ 주제분야: 기타

이 논문은 한국과학기술정보연구원과 논문 연계를 통해 무료로 제공되는 원문입니다.

서지반출

기타언어초록

It is becoming hard to maintain web applications because of high complexity and duplication of web pages. However, most of research about code clone is focusing on code hunks, and their target is limited to a specific language. Thus, we propose GSIM, a language-independent statistical approach to detect similar pages based on scarcity and frequency of customized tokens. The tokens, which can be obtained from pages splitted by a set of given separators, are defined as atomic elements for calculating similarity between two pages. In this paper, the domain definition for web applications and algorithms for collecting tokens, making matrics, calculating similarity are given. We also conducted experiments on open source codes for evaluation, with our GSIM tool. The results show the applicability of the proposed method and the effects of parameters such as threshold, toughness, length of tokens, on their quality and performance.

키워드

Web Application Page Clone Page Similarity Token Frequency

다운URL