學刊論文
中文文本可讀性探討:指標選取、模型建立與效度驗證

DOI: 10.6129/CJP.20120621
中華心理學刊 民102,55卷,1期,75-106
Chinese Journal of Psychology 2013, Vol.55, No.1, 75-106


宋曜廷(國立台灣師範大學教育心理與輔導學系);陳茹玲(國立台灣師範大學教育心理與輔導學系);李宜憲(國立台灣師範大學心理與教育測驗研究發展中心);查日龢(國立台灣師範大學心理與教育測驗研究發展中心);曾厚強(國立台灣師範大學心理與教育測驗研究發展中心);林維駿(國立台灣師範大學教育心理與輔導學系);張道行(國立高雄應用大學資訊工程學系);張國恩(國立台灣師範大學資訊教育研究所)

 

摘要

本研究根據中文特性發展可讀性指標,接著建立中文文本可讀性數學模型,並進行模型效度驗證。本研究以所發展24個可讀性指標為預測變項,386篇教科書文章之年級值為效標變項,建立逐步迴歸(stepwise regression)與SVM可讀性數學模型,再以96篇新文章為測試資料進行模型驗證。研究結果顯示:在逐步迴歸模型中,難詞數、單句數比率、實詞頻對數平均與人稱代名詞數為重要的預測變項;以SVM模型F-score方法所得的重要預測變項則為難詞數、二字詞數、字數與中筆畫字元數等。逐步迴歸模型與SVM模型對新文章的預測正確性分別為55.21%及72.92%,兩種模型預測低年級文章之正確性均高於高年級文章。


關鍵詞:可讀性、正確性、逐步迴歸、SVM數學模型


Investigating Chinese Text Readability: Linguistic Features, Modeling, and Validation

Yao-Ting Sung(Department of Educational Psychology and Counseling, National Taiwan Normal University);Ju-Ling Chen(Department of Educational Psychology and Counseling, National Taiwan Normal University); Yi-Shian Lee(Research Center for Psychological and Educational Testing, National Taiwan Normal University);Jih-Ho Cha(Research Center for Psychological and Educational Testing, National Taiwan Normal University);Hou-Chiang Tseng(Research Center for Psychological and Educational Testing, National Taiwan Normal University);Wei-Chun Lin(Department of Educational Psychology and Counseling, National Taiwan Normal University);Tao-Hsing Chang(Department of Computer Science and Information Engineering, National Kaohsiung University of Applied Sciences);Kuo-En Chang(Graduate Institute of Information and Computer Education, National Taiwan Normal University)

 

Abstract

This study aims to (a) develop readability indicators based on the textual factors that influence reading
comprehension; (b) construct the readability model for Chinese text; and (c) validate the proposed readability models. This study constructs readability models employing step regression and SVM, using 24 readability indicators as its predictive variable and the grade level of 386 textbook articles as the criteria. The proposed models are then validated according to an additional 96 texts. The results show that in step regression, the critical predictors are the number of complex words, proportion of simple sentences, average logarithm of content word frequency, and number of personal pronouns. In the SVM model, the critical predictors selected by using the F-score include the number of complex words, number of two-character words, number of characters, and number of intermediate-stroke characters. The accuracy rates of step regression and SVM are 55.21% and 72.92%, respectively. Both models predict the texts more accurately at the lower grade levels than at the higher grade levels. 


Keywords: accuracy, readability, stepwise regression, support vector machine

 

登入
會員登入
更新驗證碼