How the test works, in plain English

The words that appear on school reports, explained without jargon

These are the terms you will meet on a school report and in a Point of Progress report, written the way we would explain them to a parent rather than the way an assessment manual defines them.

Adaptive test

A test that changes as the child answers. A correct answer makes the next question harder, a wrong one makes it easier. This lets a short test measure a wide range of abilities, and it is why children get roughly a third of the later questions wrong - the test has found their edge and is staying there.

Score band, or confidence interval

The range the child's true ability most likely sits in. Reported because a single number implies precision no test of reasonable length has. If a score changes by less than the width of the band, treat it as unchanged.

Standard error

The technical name for the width of that band. A smaller standard error means a more precise estimate, and it usually comes from asking more questions.

RIT score

A scale used by some adaptive assessments, most notably NWEA's MAP Growth. Its useful property is that it is equal-interval: the gap between 200 and 210 means the same amount of learning as the gap between 230 and 240, which ordinary test percentages do not. It is not out of 100 and there is no pass mark.

We use our own scale with the same property. We do not report RIT scores, because they belong to NWEA and pretending otherwise would be dishonest.

Percentile

Where a child sits relative to others of the same age. The 70th percentile means about 70 out of 100 scored the same or lower. It is a comparison, not a measure of what a child knows.

Grade-level equivalent

A statement like "reading at a grade 6 level". Intuitive and frequently misleading: it does not mean the child could do grade 6 work, only that their score resembles the average score of children in grade 6 on this test.

Domain, cluster, skill

Layers of detail in a report. A domain is a large area such as number and operations. A cluster is a group within it, such as fractions. A skill is the specific thing that went wrong, such as comparing fractions with unlike denominators. The skill layer is where useful action lives; the domain layer is where scores are reliable enough to report. Both matter, for different reasons.

Item

A single question, including its options and the correct answer. Test designers say "item" because a question is only part of it.

Distractor

A wrong option. Good ones are not random - each represents a specific mistake, so which wrong answer a child chose tells you something. That is why a report that shows the chosen answer is worth more than one that just says "incorrect".

Item difficulty and calibration

Every question carries an estimated difficulty, which is what lets an adaptive test choose the next one. Calibration is checking those estimates against how real children actually answered, and adjusting them. An uncalibrated bank works, but its scores are less trustworthy than a calibrated one - and any honest report should say which it is.

Growth

The change between two sittings. Only meaningful if the gap between the scores is larger than the bands around them, and only fair if the two tests are on the same scale.

Formative and summative

Formative assessment is meant to guide what happens next - a diagnostic. Summative assessment is meant to record what was achieved - an exam. A practice test used to decide what to work on is formative, whatever it looks like.

Norm-referenced and criterion-referenced

Norm-referenced compares a child to other children, which is where percentiles come from. Criterion-referenced compares a child to a fixed standard - can they do long division or not. Most school reports mix both, which is a common source of confusion.