Education students in Nigeria usually build their own questionnaire or achievement test, so the panel will ask two questions: how do you know it measures what you say it measures (validity), and how do you know it gives consistent results (reliability)? Answer both in Chapter Three with evidence, not assertion. The six steps below show what that evidence is and how to report it.
Step 1: Know What the Panel Is Asking For
Validity is about whether the instrument measures the thing you intend to measure. For a researcher-made instrument in a project, the evidence is usually face validity and content validity, judged by experts. Reliability is about consistency: whether the instrument would give similar results if used again under similar conditions. The evidence is a reliability coefficient computed from a pilot test. A common student error is to write “the instrument was validated by my supervisor” and stop. That sentence names a person, not a method, and it contains no evidence. A panel wants to read who judged it, what they were given, what they were asked, what changed afterwards and what coefficient came out of the pilot.
If you are using a published instrument instead of building your own, the work is lighter: you cite the validation evidence already published and report reliability for your own sample. Our guide to which validated instrument to use for an education project covers that route. The rest of this guide is for the researcher-made instrument.
Step 2: Establish Face Validity
Face validity asks a simple question: does the instrument look, to a sensible reader, like it measures the topic? Give the draft to your supervisor and two or three colleagues. Ask them to read each item and tell you whether it is clear, whether any word is ambiguous for your respondents (for example, a word a Primary Six pupil will not know) and whether the layout, instructions and response options are easy to follow. Record who you asked and what they changed. This takes a few days and costs nothing, and it removes the embarrassing errors before the experts see the draft.
Step 3: Establish Content Validity With an Expert Panel
Content validity asks whether the items cover the whole of the topic and nothing outside it. The standard approach is judgement by people who know the subject.
- Choose the panel. Pick lecturers or practitioners with relevant expertise: someone from measurement and evaluation, someone from your subject area and, where relevant, an experienced classroom teacher. Agree the number and the names with your supervisor, and give the reasons for choosing them in Chapter Three.
- Give them a packet. The packet should contain the title, objectives and research questions, a one-paragraph description of the respondents, the draft instrument and a rating form. For an achievement test, add the table of specification (Step 4).
- Ask a specific question. A good form asks each expert to rate every item for relevance to the objective it is meant to serve and for clarity, and to add comments on missing content.
- Calculate agreement. One widely used method was developed by C. H. Lawshe (1975) in Personnel Psychology. Each expert says whether the item is “essential”, “useful but not essential” or “not necessary”, and the content validity ratio is CVR = (ne − N/2) / (N/2), where ne is the number of experts who rated the item essential and N is the total number of experts. The ratio runs from −1 to +1, and Lawshe published a table of critical values; Wilson, Pan and Schumsky (2012) later recalculated them. Check the table you cite before you set a cut-off.
- Decide and record. Keep items that meet your stated rule, revise or drop the rest, and write down every change.
Here is an illustrative calculation. Suppose eight experts rate one item and seven say it is essential. CVR = (7 − 8/2) / (8/2) = (7 − 4) / 4 = 0.75. If only five of the eight say essential, CVR = (5 − 4) / 4 = 0.25, and that item would be revised or dropped under any sensible rule. The numbers here are fictitious and show the arithmetic only.
An alternative that many education departments accept is a simple agreement index: the proportion of experts who rate an item as relevant. Whichever method you use, name it, cite its source, state the cut-off you applied and say where the cut-off comes from.

Step 4: Build a Table of Specification for an Achievement Test
If your instrument is an achievement test, content validity is usually shown through a table of specification, also called a test blueprint. It lists the topics you taught, the levels of thinking you want to test and the number of items for each cell. The illustrative table below is for a 40-item test on four topics in a secondary school subject. The weights follow the share of teaching time and are the student’s own decision.
| Topic | Share of teaching time | Knowledge | Comprehension | Application | Total items |
|---|---|---|---|---|---|
| Topic A | 30% | 4 | 4 | 4 | 12 |
| Topic B | 25% | 3 | 4 | 3 | 10 |
| Topic C | 25% | 3 | 3 | 4 | 10 |
| Topic D | 20% | 2 | 3 | 3 | 8 |
| Total | 100% | 12 | 14 | 14 | 40 |
Check the arithmetic: 12 + 10 + 10 + 8 = 40 items, and 12 + 14 + 14 = 40 as well. Item counts per topic are near the teaching-time share (30 per cent of 40 is 12, 25 per cent is 10, 20 per cent is 8). The experts then judge whether every item matches the cell it was written for and whether the topics are fairly covered.
Step 5: Run a Pilot Test
A pilot test is a trial run of the instrument on a small group that resembles your real respondents but is not part of your final sample. It tells you whether the instructions work, how long the instrument takes and whether the items behave. It also gives you the data for the reliability coefficient.
- Choose a similar group outside the sample. A different school or class with the same level, subject and background as the real respondents works well. Do not include pilot respondents in the main study.
- Agree the size with your supervisor. There is no single national number, so state the number you used and justify it.
- Use the same conditions. Same instructions, same time limit, same way of administering it.
- Record problems. Items that many respondents skip, misread or ask about should be fixed. Write down what you changed.
The interview-based version of this step is covered in our guide to the interview and focus group guide for an education project, which includes the pilot round for qualitative tools.
Step 6: Calculate Reliability
Choose the coefficient that matches the kind of instrument.
| Instrument | Coefficient | Source to cite | When to use it |
|---|---|---|---|
| Likert questionnaire (items scored 1 to 5) | Cronbach’s alpha | Cronbach (1951), Psychometrika, 16(3), 297–334 | One construct per scale, items answered on the same scale |
| Achievement test with items scored right or wrong | Kuder–Richardson formula 20 (KR-20) | Kuder and Richardson (1937), Psychometrika, 2(3), 151–160 | Dichotomous items such as multiple choice |
| Any instrument given twice to the same group | Test–retest correlation | The measurement chapter of your research methods textbook | When stability over time is the question |
Two points about alpha. First, compute it separately for each scale or construct, not once for the whole questionnaire, because alpha assumes the items measure one thing. Second, the usual rule of thumb is about 0.70 as a minimum, but cite the source you rely on, and if yours comes out low, find out why before you pad the instrument; our article on what a good Cronbach’s alpha is and what to do if yours is low shows the four usual causes and the SPSS diagnostic.
For KR-20 the formula is KR-20 = (k / (k − 1)) × (1 − Σpq / s2), where k is the number of items, p is the proportion answering an item correctly, q is 1 − p, and s2 is the variance of total scores. An illustrative case: a 40-item test with Σpq = 8.4 and a total-score variance of 36 gives (40/39) × (1 − 8.4/36) = 1.0256 × (1 − 0.2333) = 1.0256 × 0.7667 = 0.786, which you would report as 0.79. The inputs are fictitious; SPSS and most statistics packages compute KR-20 for you when the items are coded 0 and 1.

How to Write the Chapter Three Paragraphs
Here is an illustrative pair of paragraphs for a researcher-made questionnaire. Replace every detail with your own.
“Validity of the instrument. The draft questionnaire was first read by the project supervisor and two colleagues for clarity and layout, and ambiguous wording was corrected. It was then given, with the study objectives, the research questions and a rating form, to a panel of experts drawn from measurement and evaluation and from the subject area. Each expert rated every item for relevance and clarity. Items rated as relevant by the agreed proportion of experts were retained, and items below it were revised or removed, as recorded in Appendix C.”
“Reliability of the instrument. The revised instrument was administered once to a pilot group outside the main sample but similar to it in level and background. Internal consistency was computed with Cronbach’s alpha for each section. The coefficients are reported in Table 3.2.”
The same two paragraphs belong in the methodology chapter whichever department you are in. Our guide to writing Chapter Three shows where they sit, and the annotated Chapter Three example shows a complete written version. Public health students can compare the process in the questionnaire design and validation guide for public health.
Mistakes That Cost Marks
- Naming a person instead of a method. Say who the experts were, what they rated and what changed.
- Piloting on your real sample. Pilot respondents must come from outside it.
- One alpha for a mixed questionnaire. Compute it per scale.
- Using alpha for a right-or-wrong test. Use KR-20 for dichotomous items.
- Confusing the two ideas. A reliable instrument can still measure the wrong thing, so reliability does not prove validity.
- Changing items after the pilot without saying so. Record every change and, if the change is large, re-check it with an expert.
- Copying the coefficient from a published study. Report the value from your own pilot or sample.
Tesify helps you structure and organise your project, and the text is 100% written by you. Start your project with Tesify — 9,000+ students have written 15,000+ chapters with Tesify.
Frequently asked questions
What is the difference between validity and reliability?
Validity is whether the instrument measures what it is meant to measure. Reliability is whether it gives consistent results. An instrument can be consistent and still measure the wrong thing.
How many experts do I need for content validity?
There is no single national number. Agree the number with your supervisor, choose experts with relevant knowledge and state your reasons in Chapter Three.
Is it enough for my supervisor to validate the instrument?
Usually not on its own. Panels expect a described process: who judged it, what they received, how they rated it and what changed. Your supervisor can be part of it.
What is face validity?
It is whether the instrument appears, to a sensible reader, to measure the topic. It is a first check on clarity and wording, done before expert review.
When should I use KR-20 instead of Cronbach’s alpha?
Use KR-20 for items scored right or wrong, such as multiple-choice tests. Use alpha for scales such as Likert items.
Can I pilot my instrument on my own class or sample?
No. Pilot on a similar group outside your main sample, so the pilot does not influence the study and respondents are not answering the same instrument twice.
What if my reliability coefficient is low?
Check for items that are reverse-worded and not recoded, two constructs mixed in one scale and too few items, then fix the instrument and re-pilot if the change is large.
Do I need validity and reliability for a published instrument?
You cite the published evidence for validity and report reliability for your own sample, because the coefficient depends on who answered.
Where do I put the evidence?
In Chapter Three under the instrument section, with the expert rating forms and the reliability output in the appendices.
