This study aims to explore how experienced English speaking raters score speaking responses
in terms of the overall scoring style, strategies for improving scoring consistency, and
interpretation and application of scoring rubric for each scoring area. For this purpose, this
study conducted retrospective interviews with six experienced speaking raters and three
inexperienced speaking raters of NEAT, in order to compare and contrast the two rater groups’
scoring behaviors. The participants were asked to verbally report how they decided the score
of each scoring area and what strategies they employed to improve their scoring consistency.
The main findings are as follows: First, the experienced raters employed note-taking strategy
while they were listening to and scoring each response, in order to make a more accurate
decision of the score for each scoring area. Second, the experience raters applied the so-called
‘Golden-Rule’ of NEAT effectively and consistently, while the inexperienced raters were hardly
consistent in applying this rule. Third, the experienced raters scored each scoring area
independently of the other areas by applying the absolute scoring criteria given for that area
without being seriously affected by the overall impression of the response, while the
inexperienced raters decided the scores of the five scoring areas interdependently. Lastly, it
seems that inexperienced raters apply a narrow range of scale points, that is, give similar scores
for most responses if they do not have a clear understanding of the scoring criteria. Based on
the results, some suggestions are made for English speaking rater training.