Methods of Assessment
- 24. What is test reliability? What are the aspects of reliability that are most important for the teacher/tester?
- 25. How can the reliability of a test be established? What can be done about the findings?
- 26. What is test validity? What are the most important aspects of validity for the teacher/tester?
- 27. How can the validity of a test be established? What can be done about the findings?
See detailed examples of my own attempts to calculate validity in relation to the following test types:
- 24. What is test reliability? What are the aspects of reliability that are most important for the teacher/tester?
The reliability of a measuring device is high when any variations in readings taken represent true differences between the individuals being tested. Any other variation represents error.
The reliability of a test is its consistency. A tape measure that stays the same length all the time is more reliable than a piece of elastic. The same results should be obtained wherever the tape measure is used.
Reliability can also be considered in relation to different versions of a test. Whether a student takes one version of a test, such as the CFE, or another version, the result should be broadly comparable.
Three aspects of reliability:
- The circumstances in which the test is taken.
- The way in which it is marked.
- The uniformity of the assessment it makes.
Extrinsic sources of error:
- Examiner variability.
- Variability in testing conditions.
Intrinsic sources of error:
- Lack of stability.
- Lack of equivalence.
- 25. How can the reliability of a test be established? What can be done about the findings?
Examiner variability can be virtually eliminated by using objective formats.
Variability in testing conditions can be reduced by taking meticulous care over the instructions given to the test administrator and by formulating clear explanations for candidates. If necessary, candidates can be given some preliminary practice with the test rubric so that those unfamiliar with the format are not disadvantaged.
Stability reliability: a measuring device is stable if it gives the same result when used twice on the same object. If the relative positions of individuals within a group remain the same, or nearly the same, on both occasions, the test has high stability reliability.
Equivalence reliability: a measuring device is equivalent to another measuring device if both give the same results when applied to the same object. Two tests can be constructed in order to estimate equivalence reliability.
Two approaches are:
- Parallel versions: administer both versions to the same group of individuals and correlate the two sets of scores.
- Split-half reliability: if a test contains 100 items, calculate scores separately for two sets of 50 items and correlate the results. For example, the odd-numbered and even-numbered items can be treated as two halves.
Variance estimates: M = the mean; n = the number of items in the test; s = standard deviation; r = reliability estimate.
The Kuder-Richardson formula 21 can be used to estimate reliability:
r = 1 - [M(n - M)] / (n × s2)
Equivalence reliability for a certified achievement test should reach approximately 0.7. A lower figure might be acceptable for a diagnostic test which is to be used as the basis for class discussion.
If a test is designed so that the spread of scores is not similar to that of a normal distribution, for example because nearly all the items are answered correctly, the results need to be examined carefully. Useful procedures include producing a distribution histogram, examining the mean, and carrying out item analysis to identify which items were answered incorrectly.
Further information can be obtained through inspection of scripts and class discussion of the test results.
- 26. What is test validity? What are the most important aspects of validity for the teacher/tester?
When a test measures what it is intended to measure and nothing else, it is considered valid. Validity is therefore the extent to which a test measures what it is intended to measure.
The most important kinds of validity for the teacher/tester are content validity and face validity.
Content validity means that the test accurately reflects the syllabus on which it is based. A content specification or assessment plan can help to ensure that the test reflects all the areas to be assessed in appropriate proportions.
The test should provide a balanced sample without being biased towards items that are easiest to write or towards material that happens to be readily available.
Face validity concerns whether the test appears to be a good and appropriate test to teachers and students. Is it a reasonable way of assessing the students? Is it trivial? Is it unnecessarily difficult?
A formal questionnaire can be used to obtain views about face validity. Informal discussion with teachers and students can also provide useful evidence.
Predictive validity concerns the extent to which a test accurately predicts performance in some subsequent situation.
Concurrent validity concerns the extent to which a test gives similar results to existing tests that have already been validated.
Construct validity concerns the extent to which a test accurately reflects the principles of a valid theory of foreign language learning.
- 27. How can the validity of a test be established? What can be done about the findings?
See the case studies in the sections on developing a placement test and a multiple-choice test.
Ultimately, the validity of a test design rests on its relationship with the tester's own goals and objectives: that is, its success in measuring the behaviours the learners are expected to develop or the skills they need to further their own objectives.
There is some consensus within societies about useful skills and socially responsible behaviour. However, within most educational systems, it is for test designers to define and/or take account of the aims and content of the programme of study that is the subject of the assessment.