Option A

Standardised Testing

The established, data-driven benchmark for system-wide accountability.

Best for: Policymakers and administrators comparing performance across schools, districts, or entire education systems.

Option B

Continuous Assessment

The flexible, learner-centred alternative to high-stakes exams.

Best for: Educators who want ongoing insight into individual student progress and classroom teaching effectiveness.

How Each Model Defines and Measures Learning

Standardised testing refers to any assessment administered and scored under consistent, predetermined conditions — every student receives the same questions, the same time limit, and results are evaluated against a fixed rubric. High-stakes versions, such as state proficiency exams or college-entry tests, carry significant consequences for students, schools, and funding decisions.

Continuous assessment (sometimes called formative assessment or portfolio-based evaluation) distributes measurement across a course or school year. It encompasses quizzes, projects, classroom observations, self-reflection exercises, and teacher feedback gathered regularly rather than in a single sitting.

The core philosophical difference is this: standardised tests measure what a student can demonstrate on a particular day; continuous assessment attempts to capture how a student develops over time. Both are legitimate — and both have documented limitations that shape ongoing policy debates. For a broader look at how assessment questions fit into larger education research, see our overview of decades of early learning research.

CriterionStandardised TestingContinuous Assessment
Frequency Once or twice per year Ongoing throughout the term
Comparability High — consistent across schools Variable — depends on teacher/school
Feedback timeliness Delayed — results arrive weeks later Immediate or near-immediate
Curriculum impact Can narrow to tested subjects Supports broader subject coverage
Equity considerations Documented score gaps by income and race Grading inconsistency risks across schools
Administrative burden Lower for teachers; higher for test logistics Higher ongoing workload for teachers
Student anxiety Higher — single high-stakes event Generally lower — spread across time
Policy suitability Strong for system-level accountability Strong for classroom-level decisions

What the Research Actually Shows

A 2019 review published in Educational Research Review found that frequent low-stakes formative assessment — a cornerstone of continuous models — consistently improved student outcomes across subject areas and age groups when feedback was specific and timely. However, the same body of literature notes that quality of implementation matters enormously: poorly designed continuous assessment can be just as superficial as a poorly designed standardised test.

On the standardised side, research from the National Bureau of Economic Research and similar institutions has shown that accountability testing can raise average scores, particularly in foundational literacy and numeracy. The concern is what else changes: studies document narrowing of curriculum toward tested subjects, reduced time for arts, physical education, and social-emotional learning, and elevated anxiety — particularly among students in lower-income schools where consequences for low scores are most severe.

90%

Teachers reporting curriculum narrowing due to testing

A survey by the Center on Education Policy found the vast majority of U.S. elementary teachers reported reducing time on non-tested subjects after high-stakes accountability laws took effect.

~0.4

Average effect size of formative assessment on achievement

Meta-analyses cited in educational psychology literature consistently place the effect of well-implemented formative feedback in a moderate-to-strong range, measured as Cohen's d.

32

U.S. states using some form of performance assessment

As of reporting by FairTest, roughly two-thirds of U.S. states incorporate performance or portfolio tasks alongside standardised tests at certain grade levels.

Equity is a recurring theme. Standardised tests have documented score gaps correlated with race, income, and first language. Critics argue these gaps partly reflect test design rather than solely student capability. Continuous assessment can close some of those gaps — but it introduces its own equity risks, including grading inconsistency between teachers and schools, and heavier workloads that disadvantage under-resourced classrooms. Understanding how education stories frame these trade-offs is itself a skill; our guide to reading education coverage critically explores that in detail.

Several countries have moved toward hybrid systems in response to this evidence. England's GCSE examinations, for example, were redesigned in the 2010s to reduce coursework components after concerns about inconsistency — then subsequently partially restored after concerns about exam pressure. In the United States, states such as New Hampshire have piloted competency-based assessment systems that blend standardised benchmarks with locally designed tasks, an approach that aligns with broader shifts reshaping higher education credentials.

The research base for hybrid models is still developing. Early indicators from several pilots suggest that combining a common standardised measure with robust teacher-assessed components can maintain comparability while giving students more varied ways to demonstrate competence. The trade-off is administrative complexity and the need for significant teacher training — two resource demands that fall unevenly across school systems.

For adults re-entering education, assessment design also matters significantly. The growth of competency-based and portfolio credentials reflects continuous assessment principles applied outside traditional schooling — a trend explored in our overview of how adult education has changed.

A Note on How Evidence Is Reported

Studies comparing standardised testing and continuous assessment often measure different outcomes — test scores, graduation rates, self-reported well-being — making direct comparisons difficult. Effect sizes also vary significantly depending on grade level, subject area, and implementation quality. Readers should treat headline findings with appropriate caution and look for the specific context a study examined before drawing broad conclusions.

Share

Education Editorial Team · Contributor

Education Editorial Team is the collective byline for our editorial team and contributor network. Articles published under this byline or an editorial pen name are researched, written, and reviewed according to our editorial standards for clarity, consistency, and independence before publication.

The content on this site is for informational purposes only and is not a substitute for professional advice. Always consult a qualified professional for guidance specific to your situation.