Are You Testing What You Taught, or Just How You Were Tested?

It is widely acknowledged across the ELT community that pre-service and in-service teacher education programmes frequently ignore language assessment. Teacher training syllabuses regularly dedicate dozens of hours to methods, approaches, classroom management, and ed-tech tools, while assessment is all too often relegated to isolated sessions. As Popham (2001, apud GREEN, 2014, p. 20-21) points out, many teachers suffer from a chronic lack of “assessment literacy”, meaning they have rarely been trained in basic assessment principles or shown how assessment and teaching can systematically complement each other. 

Left without explicit tools, educators naturally rely on intuition. We assess the way we were assessed. We design tests using the same discrete, decontextualized multiple-choice quizzes, tricky gap-fills, and arbitrary grammar translation exercises to which we were subjected as students. However, these inherited practices were rarely grounded in foundational principles to begin with. More often than not, they were designed to catch students out rather than give them a fair chance to show what they can really do.

The purpose of this article is to shed some light on the fundamental principles of language test design, separate teaching from testing, and share a few simple principles to help you write better, fairer tests.

Teaching vs. Assessing: clarifying the mission

To break out of the cycle of this inherited knowledge, we must first make a clear distinction between the nature of teaching and the nature of assessment. As Baxter (1997, p. 9) neatly summarizes:

Every time we ask students to answer a question to which we already know the answer, we are giving them a kind of test. […] In other words, testing is generally concerned with ENUMERATION, that is, turning performance into numbers.

Teaching is fundamentally about leading students to success. When we teach, our mission is to guide, scaffold, raise linguistic awareness, and provide safety nets so that learners can make sense of the language system and its communicative functions. On the other hand, if we step back and ask the student to show us what he or she can actually do independently, under explicit constraints without intervention, this is a test. 

Confusing the two creates serious problems. When teachers ask a long list of display questions in class without providing constructive strategies, they are not teaching reading or listening; they are essentially administering micro-tests. 

Teaching nurtures the plant; testing measures whether the fertilizer worked. 

Core concepts every teacher must have in mind

When we set out to build an assessment tool, four key concepts must guide every choice we make:

  • Validity: A test is valid if it genuinely measures what it claims to measure and nothing else. As Hughes (1989, p. 22) states, a test demonstrates content validity when its content “constitutes a representative sample of the language skills, structures, etc. with which it is meant to be concerned”. If a test claims to assess oral fluency but only asks learners to transform written sentences into reported speech, it lacks construct and content validity. 
  • Reliability: Reliability refers to the consistency of measurement. In Hughes’s (1989, p. 3) terms, a test is reliable if “someone will get more or less the same score, whether they happen to take it on one particular day or on the next”. 
  • Washback (Backwash): Washback refers to the impact that a test has on the teaching and learning leading up to it. As Harris and McCann (1994, p. 2) say, examinations “can have a profound washback effect (the influence of assessment on both teaching and learning)”. If an end-of-term test only measures discrete grammar points, students and teachers will quickly ignore communicative interaction and dedicate class time strictly to grammar drills. Beneficial washback occurs when test tasks reflect authentic, real-life communication, encouraging teachers to teach communicative competence in the classroom. 
  • Practicality: Practicality represents the relationship between the resources required to design, administer, and score a test. An oral interview lasting 40 minutes for every single student might have an amazing validity, but if you have a class of 45 students and two contact hours a week, it is impractical. 

The art of designing effective tests

A common misconception among educators is that there is such a thing as an universally “perfect” test. In practice, assessment design is an exercise in compromise. 

There is an unavoidable, permanent tension between reliability and validity. Highly objective items (such as multiple-choice questions) offer virtually perfect scorer reliability and high practicality, but they measure recognition rather than authentic production, running a high risk of construct under-representation and negative washback. 

The teacher’s task is not to eliminate this tension, but to find a principled balance that honors their syllabus objectives and respects their classroom constraints. 

How not to design test items: lessons from the book Language Test Construction and Evaluation 

To illustrate how easily test item design can go wrong, Alderson, Clapham, and Wall (1995) catalogued common item-writing traps in Language Test Construction and Evaluation.

  1. Testing general knowledge or IQ instead of language: Items should never be answerable without reading the text or listening to the audio track. Alderson, Clapham, and Wall (1995, p. 50) cite the flagrant example of an item asking “Who gets food from trees?” with options like “Only man” and “Man and animals”. Any candidate can answer this based on common sense without reading a single word of the passage. 
  2. Ambiguous stems with multiple correct options: Multiple-choice items often contain distractors that are inadvertently acceptable in specific standard varieties or natural contexts. Alderson, Clapham, and Wall (1995, p. 48) show that in an item meant to test reported speech tense-shifts (“She said she [can’t/won’t/couldn’t] leave the baby”), native speakers routinely accepted “can’t” and “won’t” in natural speech, making the official answer key indefensible. 
  3. The “giveaway” distractor: When test writers try to make the correct option unmistakably accurate, they often load it with qualifiers, making it noticeably longer or more detailed than the distractors. Candidates quickly learn that the longest, most academic-sounding option is almost always the key. 
  4. Grammatical clues in the stem: Stems that end with determiners like “a” or prepositions often make certain distractors grammatically impossible, instantly giving away the answer. For example, in “Someone who designs houses is a [designer / builder / architect / plumber]”, the option “architect” is ruled out by basic phonology, not by reading comprehension. 
  5. The Domino Effect: Items must be entirely independent of one another. If correctly answering Question 2 requires information obtained by solving Question 1, a candidate who misses the first is penalized twice for a single error. 
  6. Turning language tests into word puzzles: Asking students to unscramble anagrams or solve cryptic riddles tests puzzle-solving abilities and spatial reasoning, not linguistic proficiency.
  7. Unrestricted open tasks with inadequate context: Prompts such as “Write an essay about envy” force students to rely on instant imagination and philosophical depth rather than demonstrating language control. Good prompts restrict the task, specify the audience, define the genre, and make the criteria explicit 

Designing a sound language assessment is neither a matter of blind intuition nor a job that belongs exclusively to large international testing institutions. By grounding our classroom assessments in clear specifications, balancing the four principles, and eliminating misleading item-writing traps, we turn evaluation into a powerful partner of learning.

If you are eager to deepen your knowledge on assessment, we strongly encourage you to explore the foundational literature in our references below.

References

ALDERSON, J. C.; CLAPHAM, C.; WALL, D. Language test construction and evaluation. Cambridge: Cambridge University Press, 1995.

BAXTER, A. Evaluating your students. London: Richmond Publishing, 1997.

GREEN, A. Exploring language assessment and testing: language in action. London: Routledge, 2014.

HARRIS, M.; McCANN, P. Assessment. Oxford: Heinemann, 1994.

HEATON, J. B. Writing English language tests. New ed. London: Longman, 1988.

HUGHES, A. Testing for language teachers. Cambridge: Cambridge University Press, 1989.

McNAMARA, T. Language testing. Oxford: Oxford University Press, 2000.

UNDERHILL, N. Testing spoken language: a handbook of oral testing techniques. Cambridge: Cambridge University Press, 1987.

About the writer

Jonas Ishikawa é um eterno aprendiz que atua na área da educação há mais de 9 anos. Já foi professor de inglês, formador de professores e hoje atua como coordenador pedagógico na D’Agosto Language School.

Deixe um comentário

O seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *

Posts Relacionados