Methods
How this works, and what it cannot tell you.
Every number this app produces is an inference about a person, so the reasoning behind it is written down here rather than kept in the code. Read this before you use a readiness band to make a decision about someone.
Who it is written for
Early-career individuals — new to the workforce or returning after a significant break with limited experience.
This is the first thing that decides whether a result means anything. Every level, item, and entry bar below is calibrated for this population — applied to experienced workers the bar reads as low, and applied to people well outside it the bar reads as unfair. Neither is the instrument working.
01Content comes from the occupations, not from a template
Module purposes, learning outcomes, rubric anchors, and the creditable option in every judgment item are generated from leveled behavior statements written for the specific occupations inside a sector. The scenario's setting, tools, and the people in the room come from O*NET work context for those same occupations.
The alternative — a generic item with the sector name pasted in — produces construct-irrelevant variance: you end up measuring how well someone reads generic business English. [Messick, S.]
02Distractors are built to be tempting
Each judgment item draws its options from four places: the target behavior at the level being measured, the same skill one level down (partial credit), an adjacent cluster (plausible, off-target), and a hand-written anti-pattern — the shortcut that feels efficient and creates the follow-on problem.
An item with three obviously wrong options measures reading speed. [Motowidlo, S. J., Dunnette, M. D., & Carter, G. W.] [McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P.]
03Levels are attained, not averaged
Evidence is level-anchored: an item written at Proficient tells you about Proficient and nothing else. A level is reached by earning 70% of that level's evidence with every sufficiently-evidenced level beneath it at 60% or better. A strong answer high up does not paper over a failed one lower down.
A level carrying under 2.0 weighted evidence points cannot be claimed and cannot block the levels above it — one unlucky item should not sink a cluster that holds everywhere else. Constructs resting on thin evidence are labeled provisional in the report rather than reported as a confident finding. [AERA, APA, & NCME]
04Design follows the content
Procedural work with one right answer is taught and tested differently from judgment work with a defensible range. Twelve patterns, each matched to what it can actually reach:
- Sector orientation. Setting, employers, entry roles, and the conditions the work happens under.
- Guided practice. Procedural work with a correct method — documentation, setup, handoffs, tools.
- Task simulation. Core job tasks where sequence, timing, and decision points matter.
- Case analysis. Judgment content — root cause, tradeoffs, competing priorities, risk.
- Role-play. Interaction-heavy skills — customer service, escalation, difficult news, handoffs.
- Judgment lab. Durable skills where the right move depends on reading the situation.
- AI-in-the-loop studio. The AI work already happening in this sector's occupations.
- Anchored reflection. Self-management and character content, where self-report needs a check.
- Retrieval drill. Safety rules, terminology, tolerances, codes — content that must be recalled cold.
- Team sprint. Creativity, leadership, and collaboration, which only show up under a shared deadline.
- Portfolio build. Capstone evidence — the thing that proves the rest of the program happened.
- Interview simulation. The last mile — translating what you can do into what you can say you can do.
Time is split across five strands:
- 30% Sector core — The work itself — tasks, tools, settings, and standards of this sector's entry roles.
- 33% Durable skills — The ten domains, weighted by what this sector's occupations actually demand.
- 17% AI at work — The AI use already documented in these occupations, and the judgment it requires.
- 11% Work readiness — Safety, compliance, workplace norms, and turning evidence into a hire.
- 9% Capstone — One artifact and one interview that carry the weight of the whole program.
05Every method has a limit, and the limits are published
Each question in an assessment carries this alongside its scoring key, where the person assigning it can read it:
sjt
Low-fidelity simulation. A written work situation with response options that differ in effectiveness rather than in correctness.
Cannot tell you: Measures knowledge of effective action, which is not the same as taking it under real pressure. Pair with observed performance before a high-stakes decision.
best-worst
Forced-choice discrimination. The respondent identifies the strongest and weakest of several defensible-looking behaviors.
Cannot tell you: Ranking ability can outrun performing ability. Reads discrimination, not execution.
mcq
Retrieval of sector knowledge, cued by a work context rather than a definition.
Cannot tell you: Recognition among options is easier than free recall on the job. Treat as a floor, not a ceiling.
sequence
Ordering task. The respondent reconstructs the order of a procedure.
Cannot tell you: Some real procedures allow more than one defensible order. Where they do, the scoring key over-constrains and should be reviewed.
triage
Prioritisation under a stated constraint. Ranked by consequence of delay.
Cannot tell you: Removing time pressure removes most of the difficulty. The paper version is easier than the shift.
verify
Multi-select audit with a penalty for wrong selections, so guessing everything scores worse than selecting carefully.
Cannot tell you: Selecting the right checks is not running them. Pair with an observed AI-assisted task.
constructed
Short constructed response scored against a published checklist of what the response must do.
Cannot tell you: Self-marked in this build, which inflates scores. The scorer weights self-marked evidence below observed evidence; an instructor override should replace it in a live cohort.
rubric
Behaviorally anchored rating scale. Each criterion's four levels are anchored in occupation-specific behavior statements, not in adjectives.
Cannot tell you: Rater-dependent. Anchors reduce drift but do not remove it; inter-rater agreement should be checked before the scores drive decisions.
artifact
Work sample. A finished product judged against criteria published before the work began.
Cannot tell you: Authorship is hard to verify outside a supervised setting, and one artifact is one occasion.
selfcheck
Anchored self-rating. Reported beside the measured level, never combined with it.
Cannot tell you: Self-assessment correlates weakly with performance and worst among the lowest performers. It carries zero weight in every score this instrument produces.
06Self-ratings are reported, never scored
Learners rate themselves against the same four levels. Those ratings carry zero weight in every score the app produces, and appear beside the measured level so the gap is visible. Self-assessment correlates weakly with measured performance, and worst among the people furthest from the bar. [Dunning, D., Heath, C., & Suls, J. M.] The gap is worth collecting because it is the coaching conversation, not because it is evidence.
07Prior credentials shorten the program only under three rules
Turning a credential into shortened seat time is a transfer inference. The credential says someone did something, somewhere, at some point; the claim being made is that they can do a related thing, in this sector, now. That gap needs its own warrant. [Kane, M. T.]
- Only signed, leveled credentials shorten anything. A résumé line or an unsigned badge personalises the material and buys no time.
- Four constructs can never be credited away — core, character, communication, output-verification. These are the readiness gates. A program that lets someone skip its gates on the strength of a prior badge is not measuring readiness.
- Total shortening is capped at 35%. Past a point, what is left is not an assessment of this sector's work.
A shortened module becomes a challenge check rather than disappearing: the two highest-level items remain, so the transferred claim is still tested here.
08What the instruction covers, and what it leaves to you
Durable skills and AI at work are taught here in full: the concept, the four levels as a ladder, a worked example with the reasoning shown, and the tempting mistake named. Those are teachable in prose, and a participant working alone learns them from this app.
Technical technique is not. The competency data describes what good looks like at four levels and contains no procedures — a statement reading “follow manufacturer protocols and facility checklists” points at a procedure without containing it, and searching a sector's entire bank for procedural markers returns nothing. That is not a gap in the data; occupational frameworks describe work, not curricula.
So the sector core strand teaches the judgment around a technical task — the standard it must meet, what gets verified, what the handoff needs, where it goes wrong — and measures whether someone has the technique. Building the technique happens in a shop, a lab, or a placement. Every core module says so, and an organization can attach its own training to a sector so participants are pointed at it rather than left guessing.
A program adopting this instrument expecting technical training would find the gap halfway through a cohort, which is the worst moment to find it. [Messick, S.]
09Time is a target, not a dose
The fifteen-to-twenty-hour figure is a planning target: what a learner with no prior credentialed evidence usually needs to produce enough evidence to be placed at a level. It is not a duration anyone is required to sit through, and sitting through it does not on its own guarantee a decision can be made.
It covers this program only. Technical training runs alongside it and is not counted in the figure, so a program scheduling a cohort should add its own shop, lab, or placement hours on top rather than treating fifteen to twenty hours as the whole intervention.
It is also time in this program only. Of a typical eighteen hours, about eleven are self-contained — durable skills, AI at work, work readiness — and about six build the judgment around a sector's technical work and assess it. The technique itself is learned in a shop, a lab, or a placement, and those hours are additional. A program that read the target as “eighteen hours and my welders are ready” would be reading it wrong.
What is fixed is the standard of evidence. What varies is the time to reach it. Three mechanisms make that real rather than rhetorical:
- Prior credentials shorten the plan before it starts, under the three rules in section 07.
- Settled constructs stop being asked about. Once a construct's evidence is clear of the attainment threshold by a margin — not sitting on it — further items on it cannot change the result, and the learner is told they can skip them.
- Unsettled constructs get extension items. Where evidence is thin or sits inside the band where one more item could flip the level, the program adds items and runs long. Stopping at the target hour there would mean reporting a level the evidence cannot carry.
The rule underneath is sequential: stop when the decision is stable, continue while it is not. [Kane, M. T.] [AERA, APA, & NCME] After two extension rounds a construct is flagged for an observed task rather than extended again — at that point more items are not the missing ingredient.
A learner may submit with areas unsettled. The report says which, and marks those levels provisional rather than quietly presenting them as findings.
10Reporting is de-identified, and suppression is visible
The reporting view carries no names, no email addresses, and no per-person rows. Any figure resting on fewer than five submitted attempts is suppressed, because a distribution over three people in a cohort of three is a name with extra steps. Suppressed cells are labeled as suppressed, so an empty chart is never read as a finding.
Timing is wall-clock between answers, capped per item. It measures pace, not effort. Read it for pacing problems; do not read it as how hard someone tried.
11What holds it to a standard
Subject matter experts validate this app in production. Sector content, the anti-pattern library, the crosswalks, and every question format are reviewed by people who do the work, and that review continues against live use rather than stopping at launch.
Alongside it: every question carries its scoring key, its provenance, and its published limits to the person assigning it; levels resting on thin evidence are marked provisional rather than reported as findings; and the measurement model is set out in full above rather than held as method.
12What this instrument cannot do
- Item statistics need real responses. Difficulty, discrimination, and distractor performance cannot be established from generated data — a simulation rediscovers whatever response model it was given. These come from a live cohort.
- Inter-rater agreement is unmeasured. Anchors reduce rater drift. They do not remove it. [Smith, P. C., & Kendall, L. M.]
- Self-marked written responses inflate. The scorer weights them below observed evidence; an instructor override should replace them in a live cohort.
- Knowing the right action is not taking it under pressure. Pair judgment items with observed performance before any decision that affects someone. [Roth, P. L., Bobko, P., & McFarland, L. A.]
- The sector crosswalks are unofficial. Prefix rules mapping occupations to sectors, written for this app.
References
Standard sources in personnel assessment and learning science.
- Motowidlo, S. J., Dunnette, M. D., & Carter, G. W. (1990). An alternative selection procedure: The low-fidelity simulation. Journal of Applied Psychology, 75(6).Written descriptions of work situations, with response options rated for effectiveness, predict job performance without the cost of a full simulation.
- McDaniel, M. A., Morgeson, F. P., Finnegan, E. B., Campion, M. A., & Braverman, E. P. (2001). Use of situational judgment tests to predict job performance: A clarification of the literature. Journal of Applied Psychology, 86(4).Situational judgment tests show useful criterion validity and add incremental prediction over cognitive ability and personality measures.
- Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3).Retrieving information from memory produces better long-term retention than restudying the same material for the same time.
- Bjork, E. L., & Bjork, R. A. (2011). Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning. In Psychology and the Real World.Spacing and effortful retrieval slow apparent progress during practice while improving durable performance.
- Sweller, J., van Merriënboer, J. J. G., & Paas, F. (1998). Cognitive architecture and instructional design. Educational Psychology Review, 10(3).Studying worked examples before independent problem-solving reduces extraneous load and improves acquisition for novices.
- Smith, P. C., & Kendall, L. M. (1963). Retranslation of expectations: An approach to the construction of unambiguous anchors for rating scales. Journal of Applied Psychology, 47(2).Rating scales anchored in concrete observed behaviors produce more consistent judgments than scales anchored in adjectives.
- Roth, P. L., Bobko, P., & McFarland, L. A. (2005). A meta-analysis of work sample test validity. Personnel Psychology, 58(4).Work samples — performing an actual task from the job — are among the stronger predictors of job performance.
- Campion, M. A., Palmer, D. K., & Campion, J. E. (1997). A review of structure in the selection interview. Personnel Psychology, 50(3).Structured interviews — consistent questions, anchored rating scales, multiple raters — substantially outperform unstructured ones.
- Dunning, D., Heath, C., & Suls, J. M. (2004). Flawed self-assessment: Implications for health, education, and the workplace. Psychological Science in the Public Interest, 5(3).Self-assessments of skill correlate weakly with measured performance, and least well among the lowest performers.
- Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1).Feedback about the task and the process behind it changes performance; feedback about the person generally does not.
- Messick, S. (1995). Validity of psychological assessment. American Psychologist, 50(9).Validity is a property of score interpretation, not of a test. Construct under-representation and construct-irrelevant variance are the two central threats.
- Kane, M. T. (2013). Validating the interpretations and uses of test scores. Journal of Educational Measurement, 50(1).A score's use has to be justified as an argument: from observation, to generalisation, to extrapolation, to the decision being made.
- Ericsson, K. A., Krampe, R. T., & Tesch-Römer, C. (1993). The role of deliberate practice in the acquisition of expert performance. Psychological Review, 100(3).Improvement depends on practice targeted at a specific weakness with immediate informative feedback, not on time on task.
- AERA, APA, & NCME (2014). Standards for Educational and Psychological Testing.Score reports should state what evidence supports each claim, and should not report a result the evidence cannot carry.
- America Succeeds. Pathsmith™ Durable Skills Framework — ten clusters, four performance levels. Cluster-level content in content/rubric.clusters.json.Defines the durable skill clusters and the four-level performance scale this instrument reports against.
- NSX Competency Framework (2026-07-10) — leveled core competency, durable skill, and AI-at-work statements for 1,016 O*NET occupations.Supplies the occupation-specific behavioral statements each item's creditable options and rubric anchors are written from.
- O*NET 25.1 occupation data (onetcenter.org) — tasks, skills, knowledge, technology, work context, and Job Zones.Supplies the work setting, tools, and task content that make each scenario specific to the sector rather than generic.
The shorter version of this page · The durable skills framework