aimnova.
DashboardMy LearningPaper MasteryStudy Plan

Aimnova site navigation

Stay in the loop

Get the latest study resources and updates

New features, study tips and exam insights — straight to your inbox.

IB Diploma

  • IB Past Papers
  • IB Study Notes
  • IB Question Bank
  • IB Mock Exams
  • IB Revision

IB Subjects

  • IB Math AA
  • IB Math AI
  • IB Economics
  • IB Business Management
  • IB Physics
  • IB Biology
  • View all IB subjects→

IB Past Papers

  • IB Math AA HL Past Papers
  • IB Math AA SL Past Papers
  • IB Math AI HL Past Papers
  • IB Math AI SL Past Papers
  • IB Economics HL Past Papers
  • IB Economics SL Past Papers
  • IB ESS Past Papers
  • View all past papers→

Study Resources

  • Study Notes
  • Question Bank
  • Mock Exams
  • Flashcards
  • Revision Guide
  • Exam Skills
  • Command Terms
  • Grade Calculator
  • Exam Timetable 2026

Aimnova

  • Features
  • Pricing
  • For Teachers
  • For Schools
  • For Parents
  • About Us
  • Blog
  • Contact
aimnova.

AI-powered study platform for smarter revision, past-paper analysis and examiner-style feedback.

TermsPrivacyCookies·© 2026 Aimnova. All rights reserved.4d879c2

Aimnova is not affiliated with or endorsed by the International Baccalaureate Organization (IB).

NotesPsychology HLTopic 1.4Reliability
Back to Psychology HL Topics
1.4.27 min read

Reliability (Psychology HL)

IB Psychology • Unit 1

AI-powered feedback

Stop guessing — know where you lost marks

Get instant, examiner-style feedback on every answer. See exactly how to improve and what the markscheme expects.

Try It Free

Contents

  • What reliability means
  • Three checks of reliability
  • What makes a measure unreliable
  • Consistent is not the same as correct
  • Exam-style question
In one line: Reliability is the consistency of a measure: it yields the same result when the thing measured has not changed.

Two psychologists watch the same recorded interview and rate the patient's anxiety from 0 to 10. One gives a 7, the other a 3.

The patient did not change between the two ratings. The measure did. When a score depends on who is measuring, or on the day, it lacks reliability, and nothing built on it can be trusted.

Remember: If the score moves when the person has not, the fault is in the measure.

Free preview

This is the free notes preview

You're reading the free notes. Aimnova Pro unlocks the full study experience — and you can try it with your first topic free to keep:

  • FlashcardsLock in vocabulary and key terms with spaced repetition.
  • Practice questionsAnswer exam-style questions and get instant AI marking.
  • Mock exams & past-paper vaultSit full mocks and see exactly how examiners award marks.
  • Personalised study planA daily plan built around your exam date and weak areas.
Start Studying Free Full access to Aimnova Pro · cancel anytime
Key idea: Each check compares two sets of scores that ought to agree, and expresses the agreement as a number.

Checking reliability

1

Test-retest

The same people complete the measure twice, some time apart.

The two sets of scores are correlated. A strong positive correlation shows stability over time.

2

Inter-rater

Two raters score the same behaviour independently.

Their scores are correlated, or the percentage of agreements is counted. In qualitative research the same check is made on two people coding the same transcripts.

3

Internal consistency

The items of a questionnaire are compared with one another.

One way is to correlate the scores on one half of the items with the scores on the other half.

Over time · Between raters · Within the measure

A correlation of about +0.8 or higher is usually treated as acceptable. In the example above, the two psychologists would rate twenty recorded interviews each, and the two columns of ratings would be compared.

Exam tip: Say what is compared with what.

"The two raters' scores for the same interviews are correlated" shows that you know how reliability is established.

See how examiners mark answers

Answer exam-style questions with model answers. Learn exactly what earns marks and what doesn't.

Try Exam Vault FreeYour first topic is free to keep • No credit card required
Key idea: Unreliability comes from the tool, from the people using it, and from the conditions of use.

Three sources and their remedies

1

The tool

Ambiguous items or undefined categories are read differently each time.

Remedy: operationalise precisely. "Anxious" becomes a list: trembling hands, rapid speech, avoiding eye contact.

2

The rater

Untrained raters apply their own standards.

Remedy: training on practice cases until agreement is high, and a check on agreement during the study.

3

The conditions

Differences in instructions, setting or time of day add noise.

Remedy: a standardised procedure, and enough items or observations for chance variation to average out.

Tool · Rater · Conditions

Some variation lies in the participant and is real. Mood and tiredness change from day to day, so a different score on a retest may reflect a change in the person. Test-retest reliability is therefore most meaningful for characteristics that are expected to be stable.

Behavioural categories: In an observation, inter-rater reliability rests on the categories.

Each one should be observable, defined in writing, and separate from the others.
Key idea: Reliability is necessary for a good measure. It is not enough.

Suppose the two psychologists are trained until they agree almost perfectly. If what they have agreed to rate is how fast the patient talks, their ratings are consistent, and they may still not be measuring anxiety.

Reliability of a measure

  • The tool gives consistent scores
  • Checked by retest, raters, items
  • Needed before any conclusion is drawn

Reliability of a finding

  • The result appears again in a new study
  • Checked by replication
  • Depends on a procedure that others can repeat
Remember: A finding that cannot be repeated is not reliable, however consistent the measure was.

Replication needs a procedure described in enough detail for others to follow.

Study smarter, not longer

Most students waste 40% of study time on topics they already know. Our AI tracks your progress and optimizes every minute.

Try Smart Study FreeYour first topic is free to keep • No credit card required
How this is asked: Reliability is part of the concept of measurement, offered in the 15-mark question on an unseen study.

Ratings and observations made by people are where it matters most.

The pattern

  • State how the variable was measured, and by whom.
  • Ask whether the raters would agree, and what in the study suggests they would not.
  • Ask whether the conditions were the same at each measurement.
  • Show what unreliability does to the study's comparison.
  • Apply a second concept, then judge the conclusion.
The trap: reliability and validity as one idea: "Unreliable" means inconsistent.

It is a different criticism from "measuring the wrong thing", and a study can deserve one without the other.
IB-style questionDiscuss[15 marks]

Discuss the following study with reference to two or more of the following concepts: bias, causality, measurement and/or responsibility.

A clinic wanted to know whether a new group therapy reduces anxiety. At the start and end of the ten-week programme, each of the 36 patients was interviewed for fifteen minutes by one of four therapists, who then gave an anxiety rating from 0 to 10. No guidance was given on how to arrive at the rating, and patients were not always interviewed by the same therapist on both occasions. The average rating fell from 6.8 to 5.1. The clinic concluded that the therapy reduces anxiety.

Model answer plan

See the mark-by-mark plan — for / against / judgement, with marking guidance — in study mode.

Claim your free topic

Try an IB Exam Question — Free AI Feedback

Test yourself on Reliability. Write your answer and get instant AI feedback — just like a real IB examiner.

what is meant by inter-rater reliability. [2 marks]

Related Psychology HL Topics

Continue learning with these related topics from the same unit:

1.1.1Types of bias
1.1.2Cultural bias
1.1.3Gender bias
1.1.4Reducing bias
View all Psychology HL topics

Improve your exam technique

Command terms, paper structure, and mark-scheme tips for Psychology HL

Previous
1.4.1Measuring behaviour
Next
Validity1.4.3

13 questions to test your understanding

Reading is just the start. Students who tested themselves scored 82% on average — try IB-style questions with AI feedback.

Start FreeView All Psychology HL Topics