In one line: Reliability is the consistency of a measure: it yields the same result when the thing measured has not changed.
Two psychologists watch the same recorded interview and rate the patient's anxiety from 0 to 10. One gives a 7, the other a 3.
The patient did not change between the two ratings. The measure did. When a score depends on who is measuring, or on the day, it lacks reliability, and nothing built on it can be trusted.
Remember: If the score moves when the person has not, the fault is in the measure.
Free preview
This is the free notes preview
You're reading the free notes. Aimnova Pro unlocks the full study experience — and you can try it with your first topic free to keep:
- FlashcardsLock in vocabulary and key terms with spaced repetition.
- Practice questionsAnswer exam-style questions and get instant AI marking.
- Mock exams & past-paper vaultSit full mocks and see exactly how examiners award marks.
- Personalised study planA daily plan built around your exam date and weak areas.
Key idea: Each check compares two sets of scores that ought to agree, and expresses the agreement as a number.
Checking reliability
Test-retest
The same people complete the measure twice, some time apart.
The two sets of scores are correlated. A strong positive correlation shows stability over time.
Inter-rater
Two raters score the same behaviour independently.
Their scores are correlated, or the percentage of agreements is counted. In qualitative research the same check is made on two people coding the same transcripts.
Internal consistency
The items of a questionnaire are compared with one another.
One way is to correlate the scores on one half of the items with the scores on the other half.
Over time · Between raters · Within the measure
A correlation of about +0.8 or higher is usually treated as acceptable. In the example above, the two psychologists would rate twenty recorded interviews each, and the two columns of ratings would be compared.
Exam tip: Say what is compared with what.
"The two raters' scores for the same interviews are correlated" shows that you know how reliability is established.
See how examiners mark answers
Answer exam-style questions with model answers. Learn exactly what earns marks and what doesn't.
Key idea: Unreliability comes from the tool, from the people using it, and from the conditions of use.
Three sources and their remedies
The tool
Ambiguous items or undefined categories are read differently each time.
Remedy: operationalise precisely. "Anxious" becomes a list: trembling hands, rapid speech, avoiding eye contact.
The rater
Untrained raters apply their own standards.
Remedy: training on practice cases until agreement is high, and a check on agreement during the study.
The conditions
Differences in instructions, setting or time of day add noise.
Remedy: a standardised procedure, and enough items or observations for chance variation to average out.
Tool · Rater · Conditions
Some variation lies in the participant and is real. Mood and tiredness change from day to day, so a different score on a retest may reflect a change in the person. Test-retest reliability is therefore most meaningful for characteristics that are expected to be stable.
Behavioural categories: In an observation, inter-rater reliability rests on the categories.
Each one should be observable, defined in writing, and separate from the others.
Key idea: Reliability is necessary for a good measure. It is not enough.
Suppose the two psychologists are trained until they agree almost perfectly. If what they have agreed to rate is how fast the patient talks, their ratings are consistent, and they may still not be measuring anxiety.
Reliability of a measure
- The tool gives consistent scores
- Checked by retest, raters, items
- Needed before any conclusion is drawn
Reliability of a finding
- The result appears again in a new study
- Checked by replication
- Depends on a procedure that others can repeat
Remember: A finding that cannot be repeated is not reliable, however consistent the measure was.
Replication needs a procedure described in enough detail for others to follow.
Study smarter, not longer
Most students waste 40% of study time on topics they already know. Our AI tracks your progress and optimizes every minute.
How this is asked: Reliability is part of the concept of measurement, offered in the 15-mark question on an unseen study.
Ratings and observations made by people are where it matters most.
The pattern
- State how the variable was measured, and by whom.
- Ask whether the raters would agree, and what in the study suggests they would not.
- Ask whether the conditions were the same at each measurement.
- Show what unreliability does to the study's comparison.
- Apply a second concept, then judge the conclusion.
The trap: reliability and validity as one idea: "Unreliable" means inconsistent.
It is a different criticism from "measuring the wrong thing", and a study can deserve one without the other.
Discuss the following study with reference to two or more of the following concepts: bias, causality, measurement and/or responsibility.
A clinic wanted to know whether a new group therapy reduces anxiety. At the start and end of the ten-week programme, each of the 36 patients was interviewed for fifteen minutes by one of four therapists, who then gave an anxiety rating from 0 to 10. No guidance was given on how to arrive at the rating, and patients were not always interviewed by the same therapist on both occasions. The average rating fell from 6.8 to 5.1. The clinic concluded that the therapy reduces anxiety.
Model answer plan
See the mark-by-mark plan — for / against / judgement, with marking guidance — in study mode.