- 대용량 자료에서 핵심적인 소수의 변수들의 선별과 로지스틱 회귀 모형의 전개
- ㆍ 저자명
- 임용빈,조재연,엄경아,이선아,Lim. Yong-B.,Cho. J.,Um. Kyung-A,Lee. Sun-Ah
- ㆍ 간행물명
- 品質經營學會誌
- ㆍ 권/호정보
- 2006년|34권 2호|pp.129-135 (7 pages)
- ㆍ 발행정보
- 한국품질경영학회
- ㆍ 파일정보
- 정기간행물| PDF텍스트
- ㆍ 주제분야
- 기타
In the advance of computer technology, it is possible to keep all the related informations for monitoring equipments in control and huge amount of real time manufacturing data in a data base. Thus, the statistical analysis of large data sets with hundreds of thousands observations and hundred of independent variables whose some of values are missing at many observations is needed even though it is a formidable computational task. A tree structured approach to classification is capable of screening important independent variables and their interactions. In a Six Sigma project handling large amount of manufacturing data, one of the goals is to screen vital few variables among trivial many variables. In this paper we have reviewed and summarized CART, C4.5 and CHAID algorithms and proposed a simple method of screening vital few variables by selecting common variables screened by all the three algorithms. Also how to develop a logistics regression model on a large data set is discussed and illustrated through a large finance data set collected by a credit bureau for th purpose of predicting the bankruptcy of the company.