References
d, examination-dominated traditions. Judgements are criterion-referenced against explicit performance standards rather than against the distribution of peer achievement; evidence is gathered from authentic or simulated occupational tasks rather than exclusively from written tests; and assessors are frequently practitioner-teachers exercising professional judgement rather than anonymous examiners applying mechanical mark schemes (Gonczi, 1994). These features confer considerable validity advantages, since the inference from observed performance to occupational capability is shorter and more direct than the inference from a written examination score. Yet they simultaneously generate acute reliability and comparability challenges, because distributed judgement by many assessors, across many sites, against verbally specified criteria, is inherently vulnerable to inconsistency, leniency, and divergent interpretation of standards (van der Vleuten &Schuwirth, 2005). The literature on competence-based education has repeatedly documented implementation pitfalls, including fragmented atomistic assessment, superficial tick- box evidence gathering, and the displacement of disciplinary knowledge by narrow task performance (Biemans et al., 2004; Wheelahan, 2007). It is in response to these vulnerabilities that moderation and allied quality assurance mechanisms have assumed such prominence. Moderation denotes the family of practices through which assessment judgements are reviewed, compared, calibrated, and where necessary adjusted, so that learners of equivalent capability receive equivalent results irrespective of who assessed them, where, and with what instruments (Bloxham, Hughes & Adie, 2016). In vocational systems it is typically embedded within larger architectures of internal verification, external verification, audit, provider registration, programme accreditation, and qualifications frameworks, which together are intended to warrant the claim that a certificate issued in one institution carries the same meaning as the same certificate issued in another (Cedefop, 2015). For engineering programmes, an additional layer of assurance derives from professional and international accreditation regimes that specify graduate attribute profiles for engineers, technologists, and technicians, thereby linking institutional assessment decisions to transnational benchmarks of occupational competence. The empirical record on how well these arrangements function is mixed. Studies of vocational assessment practice report substantial variability in assessor interpretation of standards, uneven www.iiardpub.org institutional capacity, and a tendency for quality assurance to drift toward documentary compliance rather than genuine calibration of judgement (Boahin and Hofman, 2014). In many developing systems, including large African TVET sectors, chronic underfunding, obsolete workshop equipment, weak industry linkages, and shortages of pedagogically prepared assessors compound these difficulties and widen the gap between the formal promise of competency-based certification and the realities of classroom and workshop practice (Okoye &Arimonu, 2016). At the same time, a rich scholarship on teacher judgement and social moderation suggests that consistency is achievable, not primarily through ever more detailed specification of criteria, but through sustained professional dialogue around exemplars of authentic student work and through assessment communities that develop shared tacit standards. The stakes of resolving these difficulties extend well beyond institutional administration. Where certification is dependable, employers reduce screening costs, learners gain portable signals of capability, articulation into higher technical study becomes negotiable, and governments can direct funding toward demonstrated performance; where it is not, the entire policy case for vocational pathways weakens, and the burden falls hardest on learners from disadvantaged backgrounds for whom a credible certificate is the principal instrument of labour-market entry. Engineering occupations intensify these stakes because incompetent practice endangers life and property, and because the sector's qualifications increasingly circulate across borders through regional labour mobility and international accreditation accords, making the comparability of judgement a matter of transnational as well as domestic concern. Assessment and its assurance are thus strategic infrastructure for skills systems, not technical appendages to curriculum policy. This review consolidates that dispersed literature for the specific context of engineering-oriented TVET. It examines the conceptual foundations of competence, the principles and instruments of outcome-referenced assessment, the psychometric and judgement-based perspectives on quality of evidence, the models and micro-practices of moderation, and the institutional and international frameworks of quality assurance within which these practices sit. It also attends to workplace- based assessment, recognition of prior learning, and the growing role of digital technologies in evidence capture and remote verification. Throughout, the review is oriented by a practical question of enduring policy significance: under what conditions can distributed, judgement- based certification of practical engineering competence be made consistent, credible, and fair, without collapsing into bureaucratic ritual or hollowing out the knowledge base of the occupations it serves (Bloxham, Hughes & Adie, 2016; Mulder, 2017)? 1.1 Background to the Review The competency-based movement in vocational education emerged from a confluence of behaviourist instructional design, manpower planning imperatives, and political demands that public training systems demonstrate labour-market relevance. Its early institutional expressions, including performance-based teacher education in the United States and subsequent national reforms in the United Kingdom and Australia, established the template of occupational standards, units of competency, and evidence-based certification that continues to structure TVET assessment today (Hodge, 2007). As the model diffused internationally, it was adapted to markedly different administrative traditions: English-speaking systems tended toward functional analysis of discrete workplace tasks, whereas continental European systems retained more holistic, occupation-centred conceptions in which knowledge, skill, and professional identity remain integrated (Brockmann, Clarke & Winch, 2008). www.iiardpub.org For engineering trades and technician education, the attraction of the competency-based template has been its promise of transparency to employers and its insistence that certification attest to demonstrable performance with tools, machines, materials, and systems rather than to seat time or examination recall. African and Asian TVET systems adopted the model partly under the influence of donor agencies and regional qualifications harmonisation initiatives, and partly from domestic frustration with theory-heavy curricula that produced graduates unable to perform basic workshop operations (Anane, 2013). Yet adoption has consistently outpaced the development of assessment capability. Standards documents have proliferated faster than assessor training, workshop infrastructure, and moderation systems, with the result that the certification signal in many systems remains noisier than the policy rhetoric concedes. Understanding why this gap persists, and what practice-level and system-level arrangements close it, constitutes the background problem space within which this review is situated. 1.2 Statement of the Problem Despite decades of reform investment, a persistent and well-documented disjuncture remains between the formal architecture of competency-based certification in engineering-oriented TVET and the actual dependability of the judgements that architecture is supposed to guarantee. Assessors working in different institutions, and frequently within the same institution, interpret identical performance criteria differently; practical tasks of nominally equivalent demand vary widely in complexity, resourcing, and supervision; and the documentary apparatus of checklists and evidence portfolios can be completed in ways that satisfy auditors while revealing little about genuine occupational capability. The consequence is an erosion of the currency of qualifications, manifested in employer scepticism, redundant in-house retesting of certified graduates, and the marginalisation of vocational pathways relative to academic ones. The problem is compounded by structural conditions specific to engineering provision. Authentic assessment of engineering competence requires functioning workshops, consumable materials, calibrated instruments, industrial placements, and assessors who combine trade mastery with assessment literacy, all of which are unevenly distributed and chronically scarce in resource- constrained systems. Moderation, where it exists, is often reduced to retrospective paperwork checking rather than substantive calibration of judgement, and external verification frequently samples documentation rather than performance. Meanwhile, the scholarly literature relevant to these difficulties is fragmented across vocational pedagogy, educational measurement, higher education assessment, and quality assurance studies, with little synthesis directed specifically at engineering trades and technician programmes. There is, accordingly, a need for an integrative review that draws these strands together, identifies what is known about securing consistent and valid competence judgements, and exposes the conditions under which prevailing quality assurance arrangements succeed or fail. 1.3 Significance of the Review The significance of this review lies first in its integrative contribution. By assembling conceptual, empirical, and policy-oriented scholarship on competence, judgement, calibration, and institutional assurance into a single analytical account focused on engineering-oriented vocational provision, it offers researchers a consolidated map of a literature that is otherwise scattered across disciplinary silos. This consolidation clarifies where evidence is robust, where it is thin, and where claims circulating in policy documents rest on assumption rather than www.iiardpub.org demonstration, thereby providing a foundation for more targeted empirical inquiry into assessor cognition, moderation efficacy, and the costs and benefits of alternative verification regimes. Second, the review carries direct practical significance for institutions and practitioners. Heads of department, instructors, workshop technicians, and internal verifiers require defensible guidance on designing practical tasks, combining evidence types, documenting judgements, and organising calibration meetings under real resource constraints. The synthesis presented here distils principles that can inform assessor development programmes, institutional assessment policies, and the design of moderation cycles that genuinely improve consistency rather than merely generating audit trails. Third, the review speaks to system-level actors. Qualifications authorities, regulators, accreditation bodies, and ministries face recurring design choices concerning the granularity of standards, the balance between internal and external verification, the role of industry assessors, and the recognition of workplace and prior learning. The analysis of comparative experience offered here illuminates the trade-offs embedded in those choices. Finally, by foregrounding conditions in resource-constrained systems, the review contributes to debates on how credible certification can support skills development, graduate employability, and public confidence in vocational pathways where they are needed most. 1.4 Aim, Objectives and Scope of the Review The aim of this review is to provide a critical, integrative synthesis of scholarship and documented practice concerning how learner competence is assessed, moderated, and quality assured within engineering-oriented technical and vocational education, and to derive from that synthesis the conditions under which certification judgements can be made consistent, valid, and credible. In pursuit of this aim, the review addresses six specific objectives. The first is to clarify the conceptual foundations of competence and competency-based education, including the principal traditions through which competence has been theorised and operationalised. The second is to examine the principles, methods, and instruments of competency-based assessment as applied to practical, cognitive, and integrated engineering performance. The third is to evaluate the quality of competence judgements through the lenses of validity, reliability, authenticity, and fairness. The fourth is to analyse moderation as a set of models, processes, and professional practices for calibrating assessor judgement. The fifth is to situate assessment and moderation within wider quality assurance architectures, including verification systems, qualifications frameworks, and professional accreditation regimes relevant to engineering. The sixth is to identify persistent challenges and promising directions, including workplace-based evidence, recognition of prior learning, and digital innovation. The scope of the review encompasses certificate, diploma, and technician-level engineering provision within public and private TVET institutions, together with associated workplace learning components. It draws on literature from anglophone and continental European systems, Australasia, and developing regions, particularly sub-Saharan Africa. Purely academic university engineering degrees are considered only where their accreditation and assessment scholarship illuminates vocational practice. The review is conceptual and narrative in method, organised thematically rather than as a systematic meta-analysis. 2. Conceptual Foundations of Competence and Competency-Based Education Any serious analysis of assessment in vocational engineering education must begin with the contested concept of competence itself, because what assessors are asked to judge depends www.iiardpub.org entirely on what competence is taken to be. The scholarly literature distinguishes several traditions. The behaviourist tradition, dominant in early Anglophone reforms, treats competence as the observable performance of discrete workplace tasks, identified through functional analysis of occupational roles and codified as units and elements of competency with associated performance criteria (Hodge, 2007). Its appeal lies in transparency and apparent objectivity: if the standard says the learner must terminate a cable, balance a rotor, or weld a butt joint to specification, then assessment is a matter of observing whether the behaviour occurs under defined conditions. Its weakness, extensively documented, is atomism. Occupational capability is decomposed into long inventories of fragmentary tasks, the integrative judgement that characterises skilled work disappears from the specification, and underpinning knowledge is reduced to whatever is strictly required to perform the listed behaviours (Hodge & Harris, 2012). A second, generic-attribute tradition conceives competence as clusters of transferable abilities, such as problem solving, communication, and teamwork, that underpin performance across contexts. While influential in graduate-attribute movements, this tradition has been criticised for vagueness and for underestimating the domain-specificity of expertise. The third, integrated or holistic tradition, articulated influentially in Australian professional education debates, defines competence as the complex combination of knowledge, skills, attitudes, and values deployed in judgement-laden performance within authentic occupational situations (Gonczi, 1994; Hager, Gonczi& Athanasou, 1994). On this view, competence is not the sum of atomised behaviours but a relational capacity exercised in context, and assessment must therefore sample integrated performance on whole tasks rather than checking off isolated elements. Eraut (1994) deepened this account by analysing the forms of professional knowledge, including tacit and procedural knowledge developed through practice, that resist codification in written standards yet are central to capable performance, a point of particular force in engineering trades where diagnostic judgement and craft knowledge are decisive. Comparative European scholarship adds a further dimension by demonstrating that competence is institutionally as well as theoretically constructed. Brockmann, Clarke and Winch (2008) show that the English skills-based conception, the German occupation-centred conception organised around the Beruf, and the French knowledge-integrated conception embody divergent assumptions about the relationship between education, work, and citizenship, with profound consequences for assessment. Where the English model certifies fragmentary outcomes assessable in any setting, the German model certifies the whole person as a member of an occupation, assessed through extended integrated examinations co-governed by social partners. Mulder, Weigel and Collins (2007) similarly document how member states operationalise competence differently within nominally convergent European policy frameworks, cautioning against the assumption that shared vocabulary implies shared practice. The Dutch research programme on comprehensive competence-based vocational education offers the most developed design framework. Wesselink et al. (2007) and Sturing et al. (2011) specify principles distinguishing genuinely competence-based programmes from relabelled traditional ones: curricula organised around core occupational problems, learning situated in authentic contexts, assessment integrated with learning and based on multiple methods, and growing learner self-responsibility. Crucially, this work treats competence-based education as a matter of degree along multiple dimensions rather than a binary status, providing institutions with an instrument for self-evaluation and staged development. Biemans et al. (2004) had earlier warned of pitfalls attending implementation, including conceptual confusion among teachers, assessment systems lagging behind curricular rhetoric, and the risk that schools adopt the www.iiardpub.org language of competence while preserving conventional testing, warnings that subsequent implementation studies have repeatedly vindicated. The conceptual debate is not merely academic, because each conception carries a distinctive assessment logic and a distinctive failure mode. Behaviourist specification yields checklist observation that is administratively tractable but hostile to integration; generic-attribute models yield rubrics so abstract that assessors cannot apply them consistently; holistic models demand rich judgement that is valid but expensive and dependent on assessor expertise. Sociological critique sharpens the stakes. Wheelahan (2007), drawing on Bernsteinian theory, argues that competency-based training, by tying curriculum tightly to workplace tasks, denies vocational learners access to the disciplinary knowledge that enables participation in society's conversations and occupational progression, an argument with particular bite in engineering, where mathematics and engineering science constitute precisely such powerful knowledge. Allais (2014) extends the critique to outcomes-led qualifications frameworks, contending that specifying outcomes independently of curriculum and institution overestimates what written standards can communicate and underestimates the institutional fabric that gives qualifications meaning. For engineering-oriented TVET, the synthesis emerging from this literature is reasonably clear. Defensible programmes require occupational standards as boundary objects linking education and industry, but those standards must specify integrated performances of professional significance rather than exhaustive task inventories; they must explicitly incorporate underpinning engineering knowledge rather than leaving it implicit; and they must be interpreted by assessor communities that understand standards as requiring judgement rather than mechanical application. The conception of competence adopted, in other words, is the first and most consequential quality assurance decision a system makes, because every downstream instrument, judgement, and moderation practice inherits its strengths and pathologies (Mulder, 2017). The sections that follow trace those inheritances through assessment design, evidence quality, and calibration practice. 3. Principles, Methods and Instruments of Competency-Based Assessment in Engineering Programmes Competency-based assessment rests on a small set of organising principles from which its characteristic methods derive. Judgement is criterion-referenced: the question is whether the learner meets the standard, not how the learner ranks against peers. Evidence must satisfy recognised rules, conventionally summarised as validity, sufficiency, authenticity, currency, and consistency, meaning that it must relate to the standard, cover its scope, be the learner's own work, be recent, and be demonstrable on more than one occasion (Gillis & Bateman, 1999). Assessment should, wherever feasible, sample integrated performance in conditions approximating real work, and it should serve learning as well as certification, supplying the feedback through which learners close the gap between current and required performance (Sadler, 1989). Finally, decisions should rest on multiple sources of evidence combined programmatically, since no single instrument can warrant an inference as consequential as occupational competence (van der Vleuten, 1996). Miller's (1990) pyramid, though developed for clinical education, supplies a widely used architecture for matching instruments to inference. Written and oral questioning assess what the learner knows and knows how to apply; structured practical tests and simulations assess whether the learner can show how under controlled conditions; and workplace observation assesses what the learner actually does in authentic practice. Engineering-oriented TVET deploys instruments www.iiardpub.org across all four levels. Knowledge of engineering science, materials, standards, and safety regulation is assessed through written tests and structured oral questioning, the latter being particularly valuable for probing the reasoning behind practical decisions. Demonstration of skill is assessed through practical tests in workshops and laboratories, ranging from timed standard tasks, such as machining a component to tolerance or wiring a distribution board to code, to fault-finding exercises on deliberately defected systems that elicit diagnostic reasoning. Integrated capability is assessed through projects requiring design, fabrication, testing, and documentation, and through capstone tasks that mirror the integrated examinations of apprenticeship traditions. Authentic performance is assessed through structured observation during industrial attachment, supported by logbooks and supervisor reports. The authenticity of these tasks is not an incidental virtue but a validity requirement. Gulikers, Bastiaens and Kirschner (2004) provide a five-dimensional framework in which authenticity is a property of the task, the physical context, the social context, the assessment result or form, and the criteria, and they emphasise that authenticity is relative to the criterion situation of professional practice. A practical test conducted on obsolete equipment, individually, under examination silence, may be inauthentic in physical and social context even if the task itself is occupationally derived, since much engineering work is collaborative and conducted amid the noise and contingency of real workshops. Constructive alignment supplies the complementary design principle: intended outcomes, learning activities, and assessment tasks must form a coherent system, such that learners cannot achieve the assessment without engaging in the learning the outcomes intend (Biggs, 1996). Misalignment, in which competence-rich outcomes are taught through demonstration but assessed through recall tests, remains among the most commonly observed defects in vocational engineering programmes. Portfolios occupy a distinctive position in the competency-based repertoire because they aggregate evidence over time and place responsibility for evidence assembly partly on the learner. A well-governed portfolio in an engineering programme combines artefacts such as fabricated components, technical drawings, test reports, and photographic or video records of process, with attestations from instructors and workplace supervisors and reflective commentary connecting evidence to criteria. The literature cautions, however, that portfolios degenerate quickly into compliance scrapbooks unless institutions specify evidence quality rules, train learners in evidence selection, and subject portfolio judgements to the same moderation as other instruments (Baartman et al., 2006). Self-assessment and peer assessment, reviewed extensively by Dochy, Segers and Sluijsmans (1999), contribute both formatively and to the development of evaluative judgement itself; Boud (2000) argues that the capacity to appraise one's own work against standards is a core outcome of education for a learning society, since certified engineers and technicians must sustain judgements about the adequacy of their own work long after instructors disappear. Boud and Falchikov (2006) accordingly urge that assessment design be evaluated not only for certification accuracy but for its contribution to learning beyond graduation. Engineering education research reinforces and extends these principles. The profession-defined attribute sets embedded in accreditation criteria encompass not only technical analysis and design but communication, teamwork, ethics, and lifelong learning, and the assessment of these professional capabilities has proved persistently difficult; Shuman, Besterfield-Sacre and McGourty (2005) review the instrument repertoire, including rubrics, behavioural observation, and portfolio evidence, while documenting the measurement challenges each entails. Felder and Brent (2003) translate outcome-based accreditation requirements into concrete course design www.iiardpub.org guidance, illustrating how programme-level attributes are decomposed into assessable course outcomes without losing integration. For technician and trade-level provision the lesson is parallel: instruments must be deliberately mapped to standards, and the map must be auditable, so that certification decisions can be traced to identified evidence. The cumulative message of this literature is that competency-based assessment is less a particular technique than a designed system of multiple, aligned, authentic instruments whose individual weaknesses are compensated programmatically, and whose dependability ultimately rests on the quality of human judgement examined in the next section (van der Vleuten &Schuwirth, 2005). A further design question with substantial practical consequences concerns the reporting metric of competence decisions. Classical competency-based systems report binary outcomes, competent or not yet competent, on the argument that occupational standards define thresholds of safe, employable performance rather than gradations of merit. Employers and learners, however, frequently demand differentiation, and many systems have introduced graded competency, in which threshold attainment is supplemented by merit and distinction bands defined through quality dimensions such as autonomy, efficiency, finish, and problem-solving sophistication. Grading raises the stakes of judgement consistency, because band boundaries are less determinate than competence thresholds and therefore more dependent on shared assessor norms; it correspondingly increases the load on exemplification and moderation (Gillis & Bateman, 1999). Rubric construction becomes the pivotal craft. Effective engineering rubrics describe qualitative differences in integrated performance, for example the distinction between a learner who diagnoses a fault through systematic isolation and one who succeeds through trial and error, rather than multiplying countable sub-criteria, and they are validated empirically by testing whether trained assessors can apply them consistently to benchmark performances. Time allocation, attempt policies, and error tolerance must likewise be standardised across sites, since identical tasks performed under different temporal and material conditions are not comparable evidence. Equally important is the policy architecture surrounding marginal evidence: rules governing resubmission, supplementary questioning to resolve doubtful observations, compensation across components, and the documentation required when assessors exercise discretion. Where such rules are explicit, assessor discretion operates within a defensible frame; where they are tacit, discretion accumulates silently into inconsistency. Finally, feedback design deserves the same rigour as instrument design. Criterion-referenced systems possess a natural feedback advantage, since performance gaps can be articulated against public standards, but the advantage is realised only when feedback is specific, actionable, and scheduled to permit improvement before terminal decisions, conditions that connect certification machinery back to its formative purpose (Sadler, 1989). 4. Validity, Reliability and Fairness in Competence Judgement The defensibility of competency-based certification turns on the classical, though continually reconceptualised, qualities of validity, reliability, and fairness. Contemporary validity theory treats validity not as a property of instruments but of interpretations: what must be defended is the inference from observed evidence to the claim that the learner is occupationally competent, together with the decisions that follow from that claim (Messick, 1995). Kane (2013) operationalises this view as an argument-based approach in which the chain of inference from scoring, through generalisation and extrapolation, to decision is made explicit and each link supported by evidence. Applied to engineering TVET, the framework is illuminating. Scoring inferences are threatened when observation checklists are completed retrospectively or when www.iiardpub.org criteria are interpreted idiosyncratically; generalisation is threatened when a single practical task is taken to warrant competence across a whole unit, since performance is notoriously task- specific; extrapolation is threatened when workshop conditions diverge sharply from workplace conditions, as when learners are assessed on equipment generations older than industry standard; and decision inferences are threatened when borderline judgements are made without clear policies on resits, compensation, and evidence sufficiency. Messick's (1995) unified framework also identifies the two cardinal threats of construct under- representation and construct-irrelevant variance, both endemic in vocational assessment. Under- representation occurs when the assessed sample fails to cover the competence construct, for instance when fault diagnosis, arguably the heart of maintenance competence, is omitted because defected equipment is unavailable, or when underpinning engineering knowledge goes untested on the assumption that performance implies understanding. Construct-irrelevant variance occurs when scores reflect factors extraneous to competence, such as the quality of workshop tooling, the legibility of handwriting in technical reports, or assessor impressions formed during instruction. Wolf (1995), in a foundational analysis of competence-based assessment, demonstrated that ever more detailed specification of performance criteria cannot eliminate judgement, because written standards are inherently indeterminate and acquire meaning only through exemplification and use; the pursuit of objectivity through specification yields documentation burdens without the hoped-for consistency. Reliability in distributed, judgement-based systems is best understood as the generalisability of decisions across tasks, occasions, and assessors. The empirical literature converges on the finding that task variability is the dominant source of unreliability, which implies that dependability is improved more by increasing the number and breadth of assessed tasks than by perfecting any single instrument (van der Vleuten, 1996). This insight grounds the programmatic principle that high-stakes decisions should aggregate many low-stakes observations. Assessor variability remains substantial nonetheless. Harlen's (2005) systematic review of teacher summative judgement found that reliability is conditional on training, exemplification, and moderation, while studies of marking in higher education reveal that assessors weight criteria differently, import tacit standards, and are influenced by holistic impressions even when using analytic rubrics (Shay, 2005). Ecclestone (2001) showed that even experienced assessors operating mature criterion systems rely on internalised, community-held norms rather than written criteria alone, a finding that reframes moderation as the cultivation of shared norms rather than the policing of rule application. Baartman et al. (2007) argue that competence assessment programmes should be evaluated against an expanded quality framework that supplements validity and reliability with criteria such as authenticity, transparency, fairness, educational consequences, and costs, evaluated at programme rather than instrument level. This reframing is consequential for engineering TVET because it legitimises combinations in which highly authentic but less standardised evidence, such as workplace observation, is balanced by more controlled evidence, such as structured practical tests, with overall dependability judged holistically. Fairness, on this account, encompasses comparability of task demand and resourcing across sites, freedom from bias in assessor judgement, reasonable adjustment for learners with disabilities, and transparency of criteria and appeal processes. In resource-diverse systems, fairness is threatened structurally: learners in poorly equipped institutions face harder conditions for demonstrating the same standard, a comparability problem that documentation-centred verification rarely detects (Halliday-Wynes & Misko, 2013). www.iiardpub.org Understanding the cognition of assessors clarifies both the sources of inconsistency and the levers for reducing it. Experienced assessors do not typically build judgements upward from criteria; they form rapid holistic appraisals grounded in internalised standards accumulated through occupational and assessment experience, and then consult criteria to check, articulate, and justify those appraisals (Shay, 2005; Ecclestone, 2001). This dual character of judgement explains familiar phenomena: two assessors can agree on a decision while citing different criteria, or apply identical rubrics to reach different decisions, because the operative standard resides in the assessor rather than the document. It also identifies the biases to which competence judgement is exposed. Halo effects allow strong performance in one component to colour judgement of others; instructor-assessors carry impressions formed during teaching into summative decisions; leniency and severity tendencies vary stably across individuals; and first impressions of confidence, fluency, or workshop demeanour can masquerade as evidence of capability. Sequence effects matter too, since a performance judged after several weak ones benefits by contrast. None of these biases is eliminated by exhortation; they are managed structurally, through second assessment of high-stakes decisions, separation where feasible of teaching and terminal assessing roles, anonymisation of written components, deliberate sampling of borderline cases for review, and above all through calibration activities that confront assessors with evidence of their own divergence (Harlen, 2005). Training designs that merely explain criteria show weak effects; designs that require assessors to judge benchmark performances, compare outcomes with peers and reference standards, and articulate the basis of discrepancies show durable gains in consistency. Wolf (1995) drew the enduring conclusion that assessor expertise is the binding constraint of criterion-referenced systems: specification can scaffold judgement but never replace it, and systems that economise on the development of judgement purchase apparent rigour at the price of real dependability. Finally, the formative dimension of assessment quality deserves emphasis. Black and Wiliam's (1998) synthesis established that assessment integrated with instruction, rich in feedback and self-assessment, yields substantial learning gains, and Sadler (1989) explained the mechanism: learners improve when they hold a concept of the standard, can compare their performance to it, and can act to close the gap. In competency-based engineering programmes, where criteria are public and performance is observable, the conditions for powerful formative practice are unusually favourable, and systems that treat assessment purely as terminal verification squander this advantage. Quality of judgement, in sum, is not achievable through instrumentation alone; it is an emergent property of programme design, assessor expertise, and the calibration practices to which the review now turns (van der Vleuten &Schuwirth, 2005). 5. Moderation of Assessment: Models, Processes and Professional Judgement Moderation is the institutional answer to the inescapability of judgement. It comprises the practices through which assessment decisions made by different assessors, in different sites, on different tasks, are brought into alignment with a common standard, so that grades and competence decisions carry equivalent meaning wherever they are issued (Bloxham, Hughes & Adie, 2016). The literature distinguishes several models. Statistical moderation adjusts school- based results against a common external measure, preserving rank order within cohorts while aligning distributions across them; it is administratively efficient but ill-suited to binary competence decisions and opaque to the assessors whose judgements it adjusts. Expert or inspectorial moderation deploys external verifiers who sample judged work, observe assessment events, and confirm or amend decisions; it provides system-level assurance but, conducted www.iiardpub.org episodically and at distance, exerts limited influence on the everyday judgement of assessors. Consensus or social moderation, in which assessors meet to compare judgements of common samples of student work and negotiate agreement with reference to standards and exemplars, is slower and costlier but is the only model that directly develops the shared understandings on which consistent judgement depends (Klenowski & Wyatt-Smith, 2010). The Australian Queensland tradition of externally moderated, school-based assessment supplies the richest empirical record of social moderation at scale. Studies of teacher judgement within that system show that consistency is an achievement of communities rather than documents: standards written as criteria and descriptors are necessary but radically insufficient, acquiring stable meaning only as teachers use them in dialogue around actual student work (Adie, Klenowski & Wyatt-Smith, 2012). Connolly, Klenowski and Wyatt-Smith (2012) found that teachers regard moderation meetings as powerful professional learning, sharpening their grasp of standards and feeding forward into task design and feedback, while also documenting the social dynamics, including status hierarchies and conflict avoidance, that can distort consensus. Crisp (2017) examined the cognition of moderators reviewing teacher-assessed projects, showing that moderation judgement is itself a complex evaluative performance involving holistic appraisal checked against criteria, rather than mechanical re-marking. These findings transfer directly to vocational engineering contexts: a moderation meeting at which assessors jointly examine a sample of welded assemblies, machined components, wiring installations, or fault-diagnosis records, argue about borderline cases, and annotate exemplars, does more for consistency than any volume of written guidance. Higher education research adds a cautionary strand concerning the rituals that pass for moderation. Bloxham (2009) argues that conventional second marking and external examining in the United Kingdom rest on false assumptions about the precision of marks and consume resources without delivering the reliability they promise. Bloxham, Boyd and Orr (2011) demonstrate that assessors' espoused use of criteria diverges from their practice, with experienced markers forming holistic judgements that criteria then rationalise. Orr (2007) reframes moderation meetings as sites where marks, and indeed students, are socially constructed through negotiation, power, and institutional positioning, while Price (2005) draws on communities of practice theory to argue that assessment standards live in collegial networks and are transferred through participation rather than documentation. Sadler (2014) presses the critique furthest, contending that attempts to codify achievement standards exhaustively are futile and that calibration of assessors, through structured engagement with exemplars until their judgements converge, is the defensible route to comparability. The implication for TVET quality assurance is sharp: systems that equate moderation with signature trails and sampling percentages may certify procedural compliance while leaving judgement uncalibrated. Operationally, mature vocational systems embed moderation across the assessment lifecycle. Pre-assessment moderation reviews tasks, marking guides, and evidence requirements before use, checking alignment with the standard, appropriateness of demand, and clarity of criteria; this stage is especially valuable in engineering programmes, where task comparability depends on materials, tolerances, equipment condition, and time allocation. During-assessment moderation includes paired observation of practical performance and the recording of evidence, by photograph, video, or retained artefact, sufficient to permit later review. Post-assessment moderation samples judged work across assessors, sites, and the grade boundary most at risk, focusing deliberately on borderline and atypical cases rather than convenient ones. Beutel, Adie and Lloyd (2017) document the practical challenges institutions encounter, including time www.iiardpub.org scarcity, geographically dispersed delivery, and variable engagement, while Grainger, Adie and Weir (2016) highlight the particular problem of inducting sessional and industry-based assessors, whose occupational expertise is indispensable but whose assessment socialisation is often thin, into the judgement community. For engineering TVET, where part-time tradesperson assessors and workplace supervisors contribute substantial evidence, structured induction, exemplar banks, and inclusion in calibration meetings are therefore not refinements but preconditions of consistency. The accumulated evidence supports three conclusions. First, moderation should be conceived primarily as professional learning that happens to produce comparability, rather than as inspection that happens to involve teachers; systems reap consistency as a by-product of investing in assessor expertise (Klenowski & Wyatt-Smith, 2010). Second, exemplars of authentic student work, annotated to show how standards were applied, are the most powerful technology of calibration available, outperforming further elaboration of written criteria (Sadler, 2014). Third, moderation design must respect the economics of judgement: universal re-marking is unaffordable and unnecessary, while risk-based sampling concentrated on new assessors, new tasks, borderline decisions, and historically divergent sites yields the greatest consistency per unit of cost (Bloxham, Hughes & Adie, 2016). The outputs of moderation deserve as much design attention as its processes, because calibration that leaves no institutional trace evaporates with staff turnover. Productive moderation cycles generate three classes of artefact. The first is decisions: confirmations, adjustments, and, where evidence is insufficient, requirements for reassessment, recorded with reasons so that learners' results rest on an auditable basis and appeals can be adjudicated fairly. The second is exemplars: performances examined during moderation, annotated with the judgement reached and the reasoning behind it, which accumulate into the institutional reference library through which standards are transmitted to new assessors and stabilised over time (Adie, Klenowski & Wyatt- Smith, 2012). The third is improvement intelligence: moderation reliably surfaces defects upstream of judgement, ambiguous task instructions, criteria that assessors interpret divergently, tasks whose demand drifted from the standard, and gaps in evidence capture, and mature systems route these findings formally into task revision and assessor development rather than leaving them as meeting talk. Timing matters correspondingly. Moderation conducted only after results are issued can protect future cohorts but not present learners, whereas moderation scheduled before certification, on a risk-sampled basis, allows adjustment while it still matters; the most developed systems operate both, with rapid pre-certification sampling of borderline and novel cases and slower annual calibration addressing systemic drift. Facilitation, finally, determines whether meetings calibrate or merely socialise. Effective sessions are chaired to surface disagreement rather than suppress it, require independent judgement before discussion so that anchoring is visible, give junior and industry-based assessors a protected voice against status hierarchies, and close with explicit articulation of the agreed standard in terms of the evidence examined (Connolly, Klenowski & Wyatt-Smith, 2012; Beutel, Adie & Lloyd, 2017). Treated this way, moderation becomes the connective tissue of the assessment system, linking judgement, task design, assessor learning, and institutional memory into a single quality loop. 6. Quality Assurance Architectures: Verification, Qualifications Frameworks and Engineering Accreditation Moderation operates within larger institutional architectures of quality assurance whose design profoundly conditions its effectiveness. Quality itself is a contested notion: Harvey and Green www.iiardpub.org (1993) distinguish quality as exception, as consistency, as fitness for purpose, as value for money, and as transformation, and vocational systems oscillate among these conceptions, with regulators emphasising consistency and compliance while educators emphasise transformation of learners. The dominant TVET architecture comprises nested layers. At institutional level, internal verification encompasses the approval of assessment instruments before use, the sampling of assessment decisions, the standardisation of assessor judgement, and the maintenance of records supporting certification claims. At system level, external verification or external moderation, conducted by awarding bodies or qualifications authorities, samples institutional practice, confirms or overturns decisions, and licenses institutions to certify. Surrounding both are provider registration regimes, programme accreditation, audit cycles, and learner appeal mechanisms. Cedefop (2015) maps these certification assurance arrangements across European vocational systems, documenting wide variation in the balance struck between trust in providers and external control, and emphasising that the credibility of certification depends on assurance reaching the assessment act itself rather than terminating at documentation. Qualifications frameworks constitute the most ambitious layer of this architecture. By specifying levels defined through learning outcomes, national and regional frameworks promise transparency, transferability, and parity of esteem between vocational and academic awards. The critical literature, however, urges caution about what frameworks can deliver. Allais (2010), reviewing implementation across sixteen countries, found that outcomes-led frameworks frequently absorbed enormous reform energy while producing modest improvement in provision, particularly where frameworks were introduced ahead of, or instead of, investment in institutions and teachers. Allais (2014) develops the theoretical argument: learning outcomes cannot bear the communicative weight placed upon them, because their meaning is parasitic on curricula, assessment traditions, and institutional reputations that frameworks neither create nor control. For assessment practice the implication is that level descriptors and unit standards function as coordination devices among parties who already share understanding, not as self-executing specifications; quality assurance regimes that audit alignment of paperwork with outcomes, while neglecting the cultivation of shared judgement, mistake the map for the territory. The African experience is instructive here: ambitious frameworks and competency-based reforms have repeatedly collided with under-resourced delivery systems, and scholars of Nigerian TVET document persistent constraints of funding, equipment, staffing, and societal esteem that no specification architecture can offset (Okoye &Arimonu, 2016). Engineering provision adds a distinctive accreditation layer linking institutional assessment to professional and international recognition. Programme accreditation by professional engineering bodies evaluates whether curricula, resources, and assessment systems produce graduates meeting specified attribute profiles, and international accords align these profiles across jurisdictions for professional engineers, technologists, and technicians, the lattermost categories being directly relevant to TVET (International Engineering Alliance, 2013). Patil and Codner (2007) review the global spread of outcomes-based engineering accreditation and the movement toward mutual recognition, while Augusti (2007) describes the European framework through which engineering programmes are accredited against shared standards while respecting national diversity. Accreditation regimes of this kind exert powerful backwash on assessment: programmes must demonstrate, with evidence, that each graduate attribute is taught and assessed, driving the construction of outcome-assessment matrices, capstone integrative tasks, and longitudinal evidence portfolios. Empirical work on engineering competencies supports the www.iiardpub.org attribute profiles these regimes encode: Passow (2012) found that practising graduates rate teamwork, communication, data analysis, and problem solving among the most important competencies in their work, and Passow and Passow (2017), synthesising a large literature, conclude that engineering practice demands integrated deployment of technical and professional competencies, vindicating assessment designs that sample whole performances. Male, Bush and Chapman (2011) similarly identify generic engineering competencies, including self- management and practical engineering capability, whose assessment requires evidence beyond written examination. The design lesson emerging from this layer of the literature concerns the relationship between compliance and trust. Heavily proceduralised assurance regimes generate documentation, standardise artefacts, and deter egregious malpractice, but they also consume assessor time, encourage performative compliance, and can crowd out the dialogic practices that actually calibrate judgement. Conversely, high-trust regimes presuppose the assessor professionalism they are excused from building. Mature systems therefore sequence their architecture developmentally: tight external control and intensive exemplification while assessor communities are immature, with progressive delegation of authority to institutions that demonstrate calibrated judgement, retaining risk-based external sampling as a permanent backstop (Cedefop, 2015). For engineering TVET in developing systems, the priority implied is unambiguous: investment in assessor capability, workshop infrastructure, and industry-linked verification yields more certification credibility than further elaboration of frameworks and forms (Allais, 2010; Okoye &Arimonu, 2016). Beneath the formal architecture, the literature consistently points to institutional quality culture as the variable that determines whether assurance mechanisms function as intended or merely as ritual. The same verification procedure operates differently in an institution where assessment is regarded as a collective professional responsibility, openly discussed and routinely improved, than in one where it is experienced as individual exposure to inspection; in the latter, sampling provokes defensive documentation, divergence is concealed rather than examined, and the information on which improvement depends never surfaces (Harvey & Green, 1993). Building the former culture is partly a leadership task, requiring heads of programme to schedule and protect time for assessment work, to model openness about their own judgements, and to decouple moderation findings from punitive personnel consequences so that error becomes discussable. It is partly structural, achieved through annual programme self-evaluation that examines pass-rate patterns, moderation adjustment rates, appeal outcomes, employer feedback, and destination data as evidence about assessment health, feeding documented action plans whose implementation is itself reviewed. And it is partly relational, sustained through stable assessment teams in which trust accumulates and through external relationships, with verifiers, partner employers, and peer institutions, conducted as professional dialogue rather than adversarial audit. Cross-institutional arrangements extend the same logic outward: networks of providers offering the same qualifications can exchange moderation samples, conduct reciprocal verification visits, and maintain shared exemplar banks, generating comparability laterally at modest cost and reducing dependence on thin central inspection capacity (Cedefop, 2015). Regulators influence culture more than they commonly acknowledge, since assurance regimes that reward demonstrated self-correction cultivate candour, while regimes that punish every disclosed weakness teach concealment. The design maxim that emerges is that every external requirement should be evaluated for the institutional behaviour it actually incentivises, not merely the assurance it nominally provides. www.iiardpub.org 7. Workplace-Based Assessment, Recognition of Prior Learning and Industry Engagement Because the criterion situation for vocational competence is the workplace, evidence generated in workplaces carries unique extrapolation validity, and competency-based systems have progressively formalised its capture. The theoretical foundation lies in workplace learning scholarship: Billett (2001) demonstrates that workplaces are structured learning environments whose affordances, the sequencing of tasks, access to guidance, and opportunities to participate, shape what learners can come to do, which implies that workplace assessment must attend to opportunity as well as performance, since a learner denied access to commissioning tasks cannot evidence commissioning competence. Trevelyan (2010), reconstructing engineering practice from field studies, shows that real engineering work is pervasively social and coordinative, consisting substantially of technical collaboration, informal teaching, and the negotiation of requirements, findings that expose the construct under-representation of assessment regimes confined to solitary bench tasks. Workplace evidence in engineering TVET typically comprises structured observation against criteria by trained workplace assessors, supervisor attestations, learner logbooks recording tasks and conditions, artefacts and records of work completed, and professional discussions in which learners account for decisions; triangulation across these sources, with institutional assessors retaining decision authority, guards against both supervisor leniency and idiosyncratic site practices. The integrity risks of workplace evidence are real and are best managed through the moderation apparatus already described: induction and calibration of workplace assessors, sampling of site judgements by institutional verifiers, evidence rules requiring corroboration, and rotation of learners across work areas to broaden the task sample. Boahin and Hofman (2013) examined competency-based training in Ghanaian polytechnics and found that the acquisition of employability skills depended on disciplinary context and on the depth of industry involvement, underlining that partnership quality, not merely placement existence, determines the value of workplace components. The wider development literature reinforces this conditionality: Tukundane et al. (2015) show that vocational programmes for marginalised youth improve livelihoods only when training quality, guidance, and labour-market linkage are present, while McGrath and Powell (2016) argue for reorienting VET around human development and capability expansion rather than narrow employability metrics, a reorientation with assessment consequences, since what is certified signals what the system values. Recognition of prior learning extends competency-based logic to its natural conclusion: if certification attests to competence rather than attendance, then competence acquired through work, informal practice, or non-formal training deserves assessment on equal terms. RPL processes in vocational engineering typically combine portfolio evidence of past work, employer verification, challenge testing on practical tasks, and professional discussion. The scholarly literature, however, documents persistent tensions. Andersson and Harris (2006) assemble theoretical perspectives showing that RPL is never a neutral mirror of existing competence: candidates must translate practical, situated knowledge into the codified categories of standards documents, a translation that disadvantages precisely the experienced workers RPL is meant to serve. Cooper and Harris (2013) analyse the knowledge question directly, showing through case studies that the alignment demanded between experiential knowledge and formal curriculum knowledge is epistemologically problematic, and that successful RPL requires mediation, navigational guidance, and assessment instruments tolerant of non-standard evidence. For engineering trades in economies with large informal sectors, these findings matter enormously: master craftspeople with decades of capability may fail RPL processes calibrated to schooled www.iiardpub.org literacies, while well-designed processes, using practical demonstration and oral examination as primary instruments, can certify them validly. Quality assurance of RPL therefore requires the same moderation disciplines as mainstream assessment, with particular attention to evidence authenticity and to fairness across candidates of differing documentary resources. Industry engagement binds these strands together. Employers contribute to standards development, supply authentic assessment contexts, second assessors with current occupational expertise, and confer labour-market recognition on resulting certificates. Yet engagement is asymmetric and fragile: small enterprises lack the capacity to host structured assessment, industry assessors require pedagogical induction, and commercial pressures can subordinate learning to production. The literature on engineering competencies suggests a constructive division of labour in which industry defines and validates the performance demands of practice, documented in studies such as Passow (2012), while educational institutions contribute assessment expertise, calibration infrastructure, and the underpinning knowledge component that workplaces under-teach. Systems that institutionalise this division, through joint assessment panels, shared exemplar banks, and reciprocal verification, convert industry engagement from rhetorical aspiration into a functioning component of quality assurance (Billett, 2001; Boahin& Hofman, 2013). 8. Digital Technologies in Competence Assessment, Evidence Management and Moderation Digital technologies have moved from the periphery to the core of vocational assessment practice, reshaping how evidence is generated, captured, stored, judged, and verified, and creating new possibilities, alongside new risks, for the quality assurance of distributed judgement. The most consequential development for competency-based systems is the electronic portfolio. Where paper portfolios fragmented evidence across binders vulnerable to loss and difficult to verify, e-portfolio platforms aggregate multimedia evidence, photographs of fabricated components at successive stages, video of practical performance, scanned drawings, digital test data, supervisor attestations, and reflective commentary, indexed against units and criteria, time-stamped, and accessible simultaneously to learners, assessors, internal verifiers, and external moderators. Joyes, Gray and Hartnell-Young (2010) synthesise the conditions of effective e-portfolio practice, emphasising that benefits depend on clarity of purpose, integration into pedagogy rather than bolt-on compliance, and institutional support for the new literacies demanded of learners and staff; their threshold-concept framing explains why e-portfolio initiatives so frequently stall when treated as repositories rather than as environments for evidencing and reflecting on learning. For engineering programmes, the affordances are tangible. Video evidence of a learner executing a pipe weld, an alignment procedure, or a switchboard termination preserves process information that retained artefacts alone cannot show, including safety behaviour, tool handling, and sequencing, and permits judgement and moderation to be separated in time and place from performance. This evidentiary shift transforms moderation economics. Where verification once required physical travel to inspect artefacts and observe assessments, remote sampling of multimedia evidence enables external verifiers to review judgements across dispersed and rural providers at radically lower cost, and enables consensus moderation meetings to convene online around shared digital exemplars. The calibration literature's central technology, the annotated exemplar, becomes scalable: institutions can maintain growing banks of judged, anonymised performances at each standard, available to every assessor and especially valuable for inducting the part-time industry assessors whose socialisation into judgement communities is otherwise thin. Remote www.iiardpub.org and technology-mediated delivery also extends assessment reach toward learners whom conventional provision serves poorly. Frempong, Ifenatuora and Ofori (2020) examine artificial- intelligence-supported conversational systems for education delivery in remote and underserved regions, illustrating a broader trajectory in which interactive digital tools carry instructional and assessment functions to settings lacking specialist staff; within vocational engineering, analogous tools can administer underpinning-knowledge questioning, provide structured feedback, and triage learner readiness for scarce, expensive practical assessment slots, provided that final competence decisions on practical performance remain anchored in human judgement of authentic evidence. Simulation constitutes a second major strand. Software simulators for programmable logic controllers, electrical circuit design, computer numerical control machining, welding, and hydraulic and pneumatic systems allow learners to practise and be assessed on procedures that would otherwise be constrained by equipment scarcity, consumable cost, and safety risk. Within Miller's hierarchy, simulation strengthens assessment at the shows-how level: it standardises task conditions, captures rich performance data automatically, permits deliberate insertion of faults for diagnostic assessment, and removes the resource inequities that threaten fairness across institutions. Its validity limits must, however, be respected. Simulated performance under- represents the embodied, material, and contingent dimensions of workshop competence, the feel of a cutting tool, the behaviour of a real arc, the improvisation demanded by worn equipment, so that extrapolation from simulator scores to occupational capability requires corroboration through real-equipment and workplace evidence (G