Harmony HR
1 How Recruitment Works2 Role, Criteria, Plan3 Funnel, Statuses
4 Screening5 Interview Process6 What We Assess7 Motivation and Signal8 Assessment, Levels, AI9 Decisions and Feedback10 Candidate Experience
11 Offer Management12 Preboarding
13 Metrics and Analytics14 Hiring Economics
Home/Chapter 6

Chapter 6. What We Assess

pp. 98–116
Candidate evidence model
Candidate evidence modelA structure for separating signals, facts, and interpretation during assessment.
Candidate evidence chain
Evidence chainHow observations become evidence and then turn into a defensible assessment decision.
Autonomy
BlockContent
Plain definitionThe ability to move work forward without constant instructions, whilst maintaining alignment with goals and constraints.
Why it mattersAutonomy reduces management overhead and is especially important where tasks are ambiguous or the manager cannot oversee every step.
Where it showsIn ambiguous tasks, independent prioritisation, requesting context, making interim decisions and escalating.
What to ask"Tell me about a task where you were given a goal but no detailed plan. How did you clarify expectations, choose the first steps and know when to escalate?"
Strong signalsClarifies the goal and constraints, proposes a plan, iterates, reports progress and knows the limits of their autonomy.
Weak signalsWaits for detailed instructions, does lots of work without validating the approach, confuses autonomy with isolation.
Red flagsIgnores alignment, makes irreversible decisions without the necessary approval, hides lack of progress.
Scale 1 / 3 / 51 = stops without instructions or acts randomly; 3 = handles clear tasks independently but gets lost in uncertainty; 5 = turns an ambiguous goal into a plan, validates assumptions, pushes work forward and escalates in good time.
Assessment errorsRewarding people who "don't ask questions" when they may simply be working without alignment and creating risk.
Mini caseSituation: manager asks to "improve client onboarding"; risk: the candidate starts with the wrong problem; how to test: ask for the first 48 hours of actions; resolution: a strong answer starts with the goal, metrics, stakeholders and quick fact-checking.
Judgement
BlockContent
Plain definitionThe ability to make sound decisions with incomplete information, considering impact, risk, trade-offs and deadlines.
Why it mattersIn real work, complete data is rarely available; the quality of decision-making determines which risks the team takes consciously.
Where it showsIn prioritisation, escalation, hiring decisions, product choices, exceptions, incident response and resource allocation.
What to ask"Tell me about a decision where you lacked sufficient data. What options did you consider, what risks did you accept and what would you revisit now?"
Strong signalsNames the decision criteria, alternatives, assumptions, reversibility, worst-case risk and the trigger for review.
Weak signalsDecides by habit or the loudest opinion, conflates urgent with important, cannot see second-order consequences.
BlockContent
Red flagsConfidently takes high-risk decisions without assessment facts, ignores contrary data, rationalises failure after the fact.
Scale 1 / 3 / 51 = chooses on whim or pressure; 3 = compares obvious options and risks; 5 = clearly states criteria, trade-offs, unknowns, worst-case protection and the review plan.
Assessment errorsConfusing decision maturity with confidence, seniority, a loud opinion or the answer matching the interviewer's views.
Mini caseSituation: client asks for a non-standard discount to close quickly; risk: margin and precedent; how to test: request a decision memo; resolution: assess criteria, alternatives, the approval route and the consequences.
Collaboration
BlockContent
Plain definitionThe ability to achieve results with other people through clear roles, trust, information sharing and shared responsibility.
Why it mattersMost work outcomes are created across functions; poor collaboration turns even strong specialists into a bottleneck.
Where it showsIn cross-functional projects, context handover, joint planning, feedback, stakeholder management and shared metrics.
What to ask"Tell me about a project where success depended on another team. How did you agree on roles, expectations and resolve disagreements?"
Strong signalsClarifies the shared goal, roles, dependencies and the right to decide, keeps others in the loop, acknowledges the team's contribution.
Weak signalsWorks through personal back-channels without transparency, complains about other functions, does not manage dependencies.
Red flagsSabotages shared decisions, takes credit for others' work, creates coalitions instead of solving the problem.
Scale 1 / 3 / 51 = works in isolation and blames other teams; 3 = collaborates normally with clear roles; 5 = builds alignment themselves, clarifies responsibility, reduces friction and helps the group reach the result.
Assessment errorsConfusing collaboration with friendliness, conflict-free behaviour or a willingness to help everyone at the expense of their own commitments.
Mini caseSituation: sales and product disagree about an urgent feature for a client; risk: conflicting goals; how to test: ask how the candidate would organise the discussion; resolution: a strong answer connects the facts, impact, the decision owner and the next step.
Conflict maturity
BlockContent
Plain definitionThe ability to deal with disagreement directly, respectfully and productively, without resorting to avoidance or personal battle.
Why it mattersConflict is inevitable in complex work; conflict maturity determines whether the decision improves or trust breaks down.
Where it showsIn disagreement, feedback, performance conversations, stakeholder conflicts, prioritisation and customer escalations.
What to ask"Think of a strong workplace disagreement. What was the other party's position, how did you verify the facts and how did you reach a resolution?"
Strong signalsCan honestly describe the opponent's position, separates people from the problem, seeks criteria, records the decision and next action.
Weak signalsAvoids difficult conversations, concedes without discussion, argues through status, only recalls the conflict as someone else's fault.
Red flagsMakes it personal, takes revenge after a decision, publicly undermines agreements, uses pressure instead of arguments.
Scale 1 / 3 / 51 = avoids conflict or turns it into a personal battle; 3 = discusses disagreement but needs moderation; 5 = creates a safe structure for the conversation, holds to criteria and drives through to a resolution.
Assessment errorsAssuming someone is mature because they never argue; the absence of conflict can mean avoidance, not maturity.
Mini caseSituation: the candidate considers a manager's decision wrong; risk: silent disagreement or a public argument; how to test: ask them to describe the conversation; resolution: assess the facts, the timeline, respect and commitment after the decision.
Systems thinking
BlockContent
Plain definitionThe ability to see connections, root causes, consequences and the feedback loop, not just the isolated symptom.
Why it mattersWithout systems thinking, a person treats symptoms, creates local optimisations and pushes the problem to another part of the process.
Where it showsIn process improvement, incident analysis, org design, funnel diagnostics, product decisions and policy changes.
What to ask"Tell me about a recurring problem you addressed. How did you find the systemic root cause, not just fix one case?"
Strong signalsAnalyses inputs, context handover, incentives, constraints, metrics, unintended consequences and the feedback loop.
Weak signalsSees only the immediate cause, proposes more control without changing the process, does not check side effects.
BlockContent
Red flagsCreates elaborate diagrams without practical action, ignores people and incentives, protects the system even when assessment facts show failure.
Scale 1 / 3 / 51 = reacts to a symptom with a one-off action; 3 = sees root causes and proposes process improvement; 5 = analyses connections, incentives, risks, metrics and changes the system so the problem does not recur.
Assessment errorsConfusing systems thinking with complex jargon, elaborate diagrams or abstract reasoning without a verifiable result.
Mini caseSituation: candidates regularly disappear after interviews; risk: the team blames the market; how to test: ask for a funnel diagnosis; resolution: a strong answer examines the timeline, communication, offer fit, interviewer behaviour and data quality.
Reliability
BlockContent
Plain definitionThe ability to consistently deliver on commitments with predictable quality and to flag deviations in advance.
Why it mattersReliability builds trust in the person and the process; without it, the team must constantly re-check the work.
Where it showsIn deadlines, recurring tasks, documentation, data hygiene, operational processes, commitments to clients and compliance checks.
What to ask"What recurring work commitments do you typically have each week or month? How do you ensure they are completed on time and without loss of quality?"
Strong signalsUses a tracking system, keeps commitments, flags risk in advance, knows their capacity limits, documents what matters.
Weak signalsRelies on memory, frequently explains delays with urgent tasks, does not record agreements in writing.
Red flagsRegularly breaks promises without warning, hides status, fakes progress or quality.
Scale 1 / 3 / 51 = commitments are unpredictable and need constant monitoring; 3 = delivers core commitments under normal load; 5 = consistently manages commitments, capacity and risks even through change.
Assessment errorsConfusing reliability with the absence of errors; a reliable person can make mistakes but makes them visible and manageable.
Mini caseSituation: the role has many routine compliance tasks; risk: one missed check is costly; how to test: ask about their self-monitoring system; resolution: assess routines, reminders, backup, escalation and assessment facts.
Adaptability
BlockContent
Plain definitionThe ability to change approach when new data, conditions or constraints emerge, without losing sight of the goal or quality of work.
Why it mattersMarkets, clients, priorities and processes change; a strong employee does not stick to an old plan when the context has shifted.
Where it showsIn changing priorities, a new manager, customer escalations, market shifts, reorgs, new tools and strategy changes.
What to ask"Tell me about a situation when your original plan stopped working because conditions changed. How did you recognise it and what did you adjust?"
Strong signalsNotices new data, revisits assumptions, preserves the goal, involves stakeholders and explains the plan change.
Weak signalsCling to the old approach, changes everything without criteria, sees change only as an external disruption.
Red flagsSabotages new rules, creates chaos under the guise of flexibility, ignores mandatory constraints.
Scale 1 / 3 / 51 = resists change or loses the outcome; 3 = adapts after being told explicitly; 5 = notices the shift themselves, rebuilds the plan, explains trade-offs and stabilises the work of the team or client.
Assessment errorsAssuming someone is adaptable because they always agree; adaptability requires criteria, not endless acquiescence.
Mini caseSituation: a key client changes requirements a day before launch; risk: the team rushes to redo everything; how to test: ask for a response plan; resolution: a strong answer clarifies the impact, must-have criteria, risk, the decision owner and communication.
Stress tolerance
BlockContent
Plain definitionThe ability to maintain work quality, communication and decision maturity under pressure, without normalising chronic overload.
Why it mattersIn stressful situations, mistakes, blunt messages and tunnel vision become costlier; the role requires resilience without glorifying burnout.
Where it showsIn incidents, peak workloads, client escalations, tight deadlines, uncertainty, public criticism and high-stakes decisions.
What to ask"Describe a period of intense work pressure. How did you prioritise, what did you delegate or escalate and how did you control the quality of decisions?"
Strong signalsNarrows focus to critical tasks, states constraints, asks for help in good time, maintains baseline quality and recovery.
Weak signalsWorks only through overtime, becomes sharp, loses priorities, does not report capacity limits.
BlockContent
Red flagsBoasts about constant burnout, snaps at people, hides mistakes under pressure, takes risky shortcuts.
Scale 1 / 3 / 51 = under pressure loses communication, quality or ethics; 3 = handles clear stress with support; 5 = preserves decision maturity, prioritises, escalates constraints and helps the system exit overload.
Assessment errorsConfusing stress tolerance with willingness to put up with poor management, chronic overtime or a toxic environment.
Mini caseSituation: a work incident and an angry client at the same time; risk: chaotic promises; how to test: ask for the first 30 minutes of actions; resolution: assess triage, the owner, communication rhythm, facts and recovery.
Ethical behaviour
BlockContent
Plain definitionThe ability to act honestly, follow the rules and protect trust, even when a quick result can be achieved by cutting corners.
Why it mattersEthical breaches create legal, financial, reputational and team risks that often cost more than the short-term gain.
Where it showsIn handling data, client promises, hiring decisions, vendor selection, reporting, expenses, compliance and conflicts of interest.
What to ask"Tell me about a situation where you were expected to deliver a result in a way that felt wrong or risky to you. What did you do?"
Strong signalsStates the principle or rule, verifies the facts, escalates appropriately, proposes a safe alternative, prepared to lose short-term upside.
Weak signalsReasons only about whether they will be caught, does not know the current rules, justifies grey areas with business pressure.
Red flagsWilling to distort data, promise the client the impossible, bypass compliance, use confidential information for unauthorised purposes.
Scale 1 / 3 / 51 = chooses the result even if it requires breaking the rules; 3 = follows explicit rules but gets lost in grey areas; 5 = foresees ethical risk, asks questions, escalates and finds a workable, legitimate option.
Assessment errorsAsking abstract questions like "are you an honest person?" instead of concrete situations with pressure, incentives and consequences.
Mini caseSituation: a manager asks to "improve" a report before a board meeting; risk: data distortion; how to test: ask for the response and escalation route; resolution: a strong answer clarifies the facts, preserves data integrity and proposes the correct wording.
Role ContextWhich Behavioural Skills Become CriticalWhy
Junior / traineeAbility to learn, responsibility, reliabilityMistakes are expected but growth and discipline matter
Senior ICJudgement, systems thinking, communicationThe person influences complex decisions without direct authority
ManagerCommunication, conflict maturity, ethical behaviour, reliabilityThe manager creates the environment and makes decisions for others
Client-facingCommunication, adaptability, stress toleranceThe client environment changes and demands trust
Regulated / high-riskResponsibility, ethical behaviour, reliabilityA mistake can create legal, financial or safety risk

6. Scorecard, scales, weights and confidence

The scorecard is not about turning recruitment into arithmetic. Its purpose is to make candidate assessment reproducible: to pre-determine what matters for the role, how each criterion is tested, what assessment facts are considered sufficient and how confident the team is in the signal. A good scorecard protects the team from three common errors: each interviewer assesses different things, a strong impression replaces facts and the final decision is made before the assessment facts are discussed.

Universal scorecard template

Use one base format for all roles and adapt the criteria content to the role. In the template below, each row is a separate criterion that must be linked to a role task and tested by a specific method. If a criterion cannot be described through assessment facts, it cannot be used for decision-making.

CriterionTypeWeightDefinitionMandatory assessment factsMethodRating 1 / 3 / 5ConfidenceNotes
Criterion nameProfessional skill / Soft skill / Motivation / Signal reliability5-30%What the candidate must be able to do or demonstrate in the context of this roleSpecific facts, examples, artefacts, answers to follow-up questions, work output, candidate's roleStructured interview, work trial, case study, portfolio review, reference check, screening1 = insufficient for the role; 3 = working minimum; 5 = strong level for this roleHigh / Medium / Low and the reasonQuestions, risks, contradictions, missing assessment facts

Minimum completion rules:

1. The criterion must be observable. "Communication" is acceptable if there is a definition; "nice person" or "cultural fit" are not.

2. Type helps avoid mixing different kinds of signal.

Professional skill shows the ability to do the work; Soft skill shows work behaviour; Motivation shows alignment with role expectations; Signal reliability shows how much the evidence can be trusted.

3. Weight reflects the cost of getting the criterion wrong.

The more the criterion affects the first 3-6 months of results, the higher the weight.

4. The definition is written before the interview, not after meeting the candidate.

Candidate scorecard
Candidate scorecardA scorecard layout that keeps the team anchored to criteria instead of personal impressions.

5. Mandatory assessment facts record what the interviewer must bring to the debrief.

Without assessment facts, the score remains an unverified impression.

6. The method must fit the criterion.

For example, discovery is best tested with a role-play scenario or case study, and reliability with behavioural interviews and references.

7. Rating 1 / 3 / 5 describes the scale anchors.

The interviewer must not invent the scale during the meeting.

Regulated / high-risk Responsibility, ethical behaviour, A mistake can create legal, finan-

reliability cial or safety risk

6. Scorecard, scales, weights and confidence

The scorecard is not about turning recruitment into arithmetic. Its purpose is to make candidate assessment

reproducible: to pre-determine what matters for the role, how each criterion is tested, what assessment facts are

considered sufficient and how confident the team is in the signal. A good scorecard protects the team from three

common errors: each interviewer assesses different things, a strong impression replaces facts and the final

decision is made before the assessment facts are discussed.

Universal scorecard template

Use one base format for all roles and adapt the criteria content to the role. In the template below, each row is

a separate criterion that must be linked to a role task and tested by a specific method. If a criterion cannot be

described through assessment facts, it cannot be used for decision-making.

Criterion Type Weight Definition Mandatory Method Rating 1 / 3 Confidence Notes

assessment / 5

facts

Criterion Professional 5-30% What the candidate Specific Structured 1 = insuffi- High / Questions, risks,

name skill / Soft must be able to facts, ex- interview, cient for Medium / contradictions,

skill / Motiva- do or demonstrate amples, work trial, role; 3 = work- Low and missing assess-

tion / Signal in the context of artefacts, case study, ing minimum; reason ment facts

reliability this role answers to portfolio 5 = strong

follow-up review, level for

questions, reference this role

work output, check,

candidate's role screening

Minimum completion rules:

1. The criterion must be observable. "Communication" is acceptable if there is a definition; "nice person"

or "cultural fit" are not.

2. Type helps avoid mixing different kinds of signal. Professional skill shows the ability to do the

work; Soft skill shows work behaviour; Motivation shows alignment with role expectations; Signal

reliability shows how much the evidence can be trusted.

3. Weight reflects the cost of getting the criterion wrong. The more the criterion affects the first 3-6

months of results, the higher the weight.

4. The definition is written before the interview, not after meeting the candidate.

5. Mandatory assessment facts record what the interviewer must bring to the debrief. Without assessment

facts, the score remains an unverified impression.

6. The method must fit the criterion. For example, discovery is best tested with a role-play scenario or

case study, and reliability with behavioural interviews and references.

7. Rating 1 / 3 / 5 describes the scale anchors. The interviewer must not invent the scale during the meeting.

8. Confidence is filled in separately from the rating. A high rating with low confidence means "it looks

like a strong signal, but there is little evidence."

9. Notes must not turn into a free-form essay. Write facts, short quotes, next-step questions and gaps.

Scale 1 / 3 / 5

Use practical scale anchors, not school grades:

Rating Meaning When to assign What should be in the assessment facts

1 Below working minimum or a risk Candidate does not demonstrate the right A concrete example of failure, absence

for the role behaviour, confuses the basic logic of the of the required action, a contradiction,

criterion, gives only generalities or shows a weak result, inability to explain

a red flag their contribution

3 Sufficient working level Candidate can perform the task in a A relevant work example, a clear

typical context but requires clear conditions, process, acceptable result, statement of

support or does not show consistency in constraints or level

challenging cases

5 Strong level for this role Candidate consistently demonstrates the Multiple relevant examples, complexity,

criterion under challenging conditions, independent role, measurable result,

explains trade-offs, sees risks and can mature explanation of decisions

improve the system around the task

If an intermediate rating of 2 or 4 is needed, use it only when the assessment facts fall between the scale anchors.

Do not turn the scale into "liked them a 4." An intermediate rating must be explainable through the neighbouring

scale anchors.

How to assign weights

Weights are not for pretty maths but for disciplined prioritisation. Without weights, the team often overvalues

bright but secondary signals: charisma, experience at a well-known company, a polished presentation or

resemblance to a successful employee.

Practical rules:

Situation Weighting Decision Example

Criterion is directly linked to the main business 20-30% For a CSM who must reduce churn, customer

result of the role risk diagnosis could carry 25%

Criterion is needed every day but is not the 10-20% Communication for most client-facing

primary differentiator roles

Criterion matters as a minimum threshold Pass / fail or 5-10% Ethical behaviour, legal eligibility,

basic literacy

Criterion can be developed after entry 5-10% Knowledge of a specific internal tool

without high risk

Criterion sounds desirable but is not linked Remove "Start-up mindset" if the role is

to role tasks actually process-driven and predictable

Recommended starting range for a role scorecard:

Type Typical total weight Comment

Professional skills 40-60% Tests the ability to produce a work outcome

Behavioural skills 25-40% Tests behaviour in an environment where

results are created with people

Motivation 10-20% Tests the durability of interest in the real role

Signal reliability Not always a weight, more often a confidence layer Does not replace criteria but shows how

much to trust the assessment facts

Weights must sum to 100% if the team uses a weighted view. But the final decision need not be an automatic

calculation: some criteria are threshold-based and may stop the process regardless of the average score.

Confidence rating

Confidence does not measure the candidate's quality but the quality of the signal. It answers the question:

"How confident are we that the assessment reflects real ability or risk?"

Confidence Meaning When to use

High Assessment facts are relevant to the role, Candidate broke down a similar case,

the next-step questions are specific, there are no explained their role, showed an artefact

major contradictions or completed a work trial

Medium Assessment facts are partially relevant but There is a good example but from a

there are constraints on context, depth or different context; details unverified; the

repeatability method was only an interview

Low Assessment facts are weak, indirect, Candidate spoke in generalities, interviewer

contradictory or obtained through an did not ask follow-up questions, no data on

inappropriate method a complex criterion

Record confidence as High - two relevant examples and a case, Medium - a similar example but

without metrics, Low - self-report only. A single word is not enough.

The difference between the score and confidence

The rating answers the question: "What level did the candidate demonstrate on this criterion?"

Confidence answers the question: "How reliably do we know that?"

Examples:

Situation Rating Confidence What to do

Candidate answers discovery questions 5 Medium Do not lower the rating automatically;

brilliantly but there is no role-play and add targeted verification or a

examples are only from small clients reference check

Candidate performed poorly on a work 1 High Treat as a strong negative assessment

trial for a key task; the method was close fact

to real work

SituationRatingConfidenceWhat to do
The interviewer feels the candidate is unreliable but the notes only contain "didn't like their style"Cannot be usedLowMove to questions / missing assessment facts, not to the decision
There is an average behavioural answer but the reference confirms the same pattern3HighConsider a consistent working level

When not to average scores. The mean is useful for review but dangerous as an automatic decision. Do not average ratings if:

1. There is a threshold criterion.

For example, ethical behaviour, legal eligibility, mandatory certification, safety, right to work with data.

2. The criteria tested different risk levels.

Weak customer risk diagnosis must not be offset by strong presentation if the role is responsible for churn.

3. Confidence differs significantly.

A single high rating with low confidence must not outweigh several medium high-confidence signals.

4. The methods are not comparable. A work trial on a real task is usually more reliable than a general conversation.

5. There is a red flag.

A red flag must be examined separately: confirm it, dismiss it or halt the process.

6. The role has mandatory constraint criteria.

For example, working hours, travel, language requirements, regulated access, conflict of interest.

7. The ratings reflect different stages.

A screening rating and a final case rating cannot be mechanically added together if they tested different criteria. Use the average only as a summary view after the debrief, not as a substitute for the decision. A proper debrief starts with the assessment facts per criterion, not with an overall candidate ranking. End-to-end example: minimal scorecards for BDM and Backend. Use these tables as a compact version of the scorecard. The full criteria library may be broader but the first working version is best kept short: 5-6 criteria, clear scale anchors and separate confidence.

Criterion BDMWeight135
B2B SaaS qualificationHighSpeaks in generalities about salesRuns a clear qualification logicBuilds a repeatable process and explains trade-offs
Discovery qualityHighPresents before understanding the clientAsks baseline questions and records the painIdentifies business / process / technical blockers and the next step
Payment domain learningMediumDoes not understand the payment context and does not know how to learnSees baseline payment constraintsConnects payment flow, risk, integration and buyer concerns
Funnel ownershipHighCannot explain the stages / reasonsRuns the funnel and next actionsAnalyses win / loss causes and improves the process
Stakeholder communicationMediumSpeaks the same way toAdapts language to the roleHolds complex

1. There is a threshold criterion. For example, ethical behaviour, legal eligibility, mandatory

certification, safety, right to work with data.

2. The criteria tested different risk levels. Weak customer risk diagnosis must not be offset by

strong presentation if the role is responsible for churn.

3. Confidence differs significantly. A single high rating with low confidence must not outweigh

several medium high-confidence signals.

4. The methods are not comparable. A work trial on a real task is usually more reliable than a general conversation.

5. There is a red flag. A red flag must be examined separately: confirm it, dismiss it or halt the process.

6. The role has mandatory constraint criteria. For example, working hours, travel, language requirements,

regulated access, conflict of interest.

7. The ratings reflect different stages. A screening rating and a final case rating cannot be mechanically

added together if they tested different criteria.

Use the average only as a summary view after the debrief, not as a substitute for the decision. A proper debrief

starts with the assessment facts per criterion, not with an overall candidate ranking.

End-to-end example: minimal scorecards for BDM and Backend

Use these tables as a compact version of the scorecard. The full criteria library may be broader but the first

working version is best kept short: 5-6 criteria, clear scale anchors and separate confidence.

Criterion BDM Weight 1 3 5

B2B SaaS qualification High Speaks in Runs a clear Builds a

generality qualification repeatable

about sales logic process and

explains

trade-offs

Discovery quality High Presents before Asks baseline Identifies business /

understanding questions process / technical

the client and records blockers and

the pain the next step

Payment domain learning Medium Does not See baseline Connects

understand payment payment flow,

the payment constraints risk, integration

context and and buyer

does not know concerns

how to learn

Funnel ownership High Cannot explain Runs the Analyses win /

the stages / funnel and loss causes and

reasons next actions improves the

process

Stakeholder communication Medium Speaks the same Adapts language Holds complex

way to everyone to the role multi-stakeholder

Criterion BDMWeight135
everyonestakeholdersmulti-stakeholder conversation
Learning agilityMediumBecomes defensive about gapsAcknowledges gaps and learnsBuilds their own learning cycle and validates hypotheses
Criterion BackendWeight135
Backend fundamentalsHighKnows the framework but poorly explains design decisionsConfidently designs typical backend tasksHolds API, data, async, alignment and evolution
Work ownershipHighDid not own the delivery after mergeParticipated in release / monitoringOwned changes end-to-end and learned from incidents
Reliability thinkingHighDoes not see failure modesNames typical risksProactively designs retries, idempotency, observability, rollback
Code qualityMediumWorks but fragile codeReadable code with baseline testsMaintainable code, good tests, review-friendly decisions
CollaborationMediumStruggles to explain trade-offsWorks normally with the product / teamHelps the team take balanced decisions
Domain learningMediumNot interested in the payment contextWilling to learnConnects product risk and engineering design

Example scorecard: Customer Success Manager for B2B SaaS

Role context: The CSM manages B2B SaaS clients after the sale, owns adoption, renewal readiness, risk

visibility, escalation management and qualitative handover of context between sales, implementation, support and

product. The role is not a pure support position: the candidate must understand the client's business objectives,

spot risk before churn signals and communicate without making empty promises.

Criterion Type Weight Definition Mandatory Method Rating 1 / 3 Confidence Notes

assessment / 5

facts

Discovery Professional 20% Can surface Can identify Role-play, 1 = asks superficial Fill in post- Check that

skill business objectives, business goals, customer case, questions and interview discovery is

solution criteria, structured interview; immediately sells not replaced

stakeholders, interview the product

success metrics and with follow-up

implementation questions the candidate asks;

constraints how goal, pain, sets; how they record the

barriers and the next step

translates them into a plan

Client risk Professional 25% Early at detecting Real escalation Case interview, 1 = reacts only Fill in post- Key criterion

diagnosis skill churn risk, low or incident review behavioural after complaints interview for the role;

adoption, stakeholder breakdown or case; interview, or renewal risk; weak high-

loss, value gaps or list of risk signals; reference confidence

implementation drift actions in the first checks cannot be

CriterionTypeWeightDefinitionMandatory assessment factsMethodRating 1 / 3 / 5ConfidenceNotes
and separating symptoms from root cause48 hours; who they involve3 = spots obvious signals and escalates; 5 = systematically diagnoses root cause, prioritises client portfolios, proposes a risk plan and communication rhythmoffset by charm
CommunicationSoft skill15%Structurally communicates status, risks, decisions and expectations to the client and internal teamsExample of a difficult letter or conversation; how they adapt the message for the client sponsor, user, product, salesStructured interview, writing sample, role-play scenario1 = speaks generally, over-promises, hides uncertainty; 3 = clearly communicates status and next steps; 5 = makes a complex situation clear, records the responsible person, trade-offs, risks and next actionFill in post-interviewAssess clarity and ownership, not extroversion
ReliabilitySoft skill15%Keeps commitments, next actions, account notes, internal context handover and escalations without constant oversightExample of a period with many clients; tracking system; reference feedback on follow-throughBehavioural interview, reference check1 = forgets commitments or blames external causes; 3 = reliable under normal load and clear processes; 5 = builds a next-action system, flags risks in advance, restores trust after a failureFill in post-interviewTest repeatability, not a single heroic case
Conflict maturitySoft skill10%Able to handle client or internal conflict without defensiveness, blame or empty concessionsSituation with an angry customer, disagreement with sales or product, escalation debriefBehavioural interview, role-play scenario1 = argues, blames or promises the impossible; 3 = stays calm and looks for a solution; 5 = validates the problem, separates facts from emotions, agrees on options, boundaries and the responsible personFill in post-interviewImportant for renewals and escalations
Motivation for client workMotivation15%Understands the reality of the CSM role: repetitive next actions, difficult clients, product implementation work, business reviews, pre-renewal pressureReasons for choosing the role; what energises and what drains; examples of sustained interest in client outcomesScreening, motivation interview, reference check1 = wants the role as a bridge into strategy or product without client routine; 3 = accepts client work and understands the main challenges; 5 = consciously chooses long-term client outcomes, can workFill in post-interviewDo not confuse motivation with an energetic interview style
CriterionTypeWeightDefinitionMandatory assessment factsMethodRating 1 / 3 / 5ConfidenceNotes
with routine and difficult conversations

The decision based on this scorecard must not be taken on a single average number. For this role, customer risk diagnosis is a near-threshold criterion: if the candidate scores 1 with high confidence, the process usually needs to be halted or a targeted check set only with very strong grounds. For motivation, a score of 3 is sufficient if professional skills are strong and role expectations have been discussed honestly. For communication and conflict maturity, it is important to look not only at the rating but also at the assessment facts: a calmer, less flashy candidate may be stronger than a candidate with polished but imprecise speech.

7. SOP: How to embed candidate assessment in the team

The SOPs in this section are designed as process documents that can be copied into a corporate knowledge base, ATS working guide or hiring operations manual and adapted to the role, stage, local legal requirements, internal policies and the specific ATS. They do not constitute legal advice. Before use in regulated hiring, mass recruitment, international processes or AI-assisted workflows, they must be reviewed with a legal or compliance lead. Each SOP below answers one operational question: who does what, when, on what inputs, which fields are mandatory, what counts as output and how to measure the quality of the process. If the company already has its own SOP format, transfer the content into it but do not remove the fields: Purpose, Scope, Roles, Inputs, Steps, SLA / timelines, Mandatory fields, Outputs, Exceptions, Metrics and Example. SOP: Professional Skills and Behavioural Skills

FieldProcess description
PurposeDefine work-related Professional Skills and Behavioural Skills for a specific role before interviews begin, so that the team assesses the ability to do the work rather than a CV, charisma or personal preferences.
ScopeUsed for new roles, backfills, level changes, launching a new hiring campaign and reviewing the scorecard after a calibration meeting. Not used for adjusting criteria to fit a favoured candidate.
RolesThe hiring manager owns the business outcomes and must-have task criteria. The recruiter runs the process and documents the criteria, checking consistency across stages. Interviewers verify whether the criteria are observable. Legal / HR compliance reviews criteria where mandatory. The role owner or hiring lead approves.
InputsDraft job description, business outcome of the role, success expectations for the first 3-6 months, key tasks, team context, failure modes from past hires, legal or compliance constraints, compensation / level range, existing competency framework if available.
Steps1. Start with the business outcome, not a skills list. Record what the role must change or protect in the business. 2. List 5-8 key tasks where poor performance would be costly. 3. For each task, define Professional Skills as practical abilities, not tool names. 4. Define Behavioural Skills only where they affect performance in this environment. 5. Add motivation risks that could lead to failure or departure of a strong person. 6. Remove criteria that are not linked to work, not

a score of 3 is sufficient if professional skills are strong and role expectations have been discussed honestly. For

communication and conflict maturity, it is important to look not only at the rating but also at the assessment

facts: a calmer, less flashy candidate may be stronger than a candidate with polished but imprecise speech.

7. SOP: How to embed candidate assessment in the team

The SOPs in this section are designed as process documents that can be copied into a corporate

knowledge base, ATS working guide or hiring operations manual and adapted to the role, stage, local

legal requirements, internal policies and the specific ATS. They do not constitute legal advice. Before

use in regulated hiring, mass recruitment, international processes or AI-assisted workflows, they must

be reviewed with a legal or compliance lead.

Each SOP below answers one operational question: who does what, when, on what inputs, which fields

are mandatory, what counts as output and how to measure the quality of the process. If the company

already has its own SOP format, transfer the content into it but do not remove the fields: Purpose,

Scope, Roles, Inputs, Steps, SLA / timelines, Mandatory fields, Outputs, Exceptions, Metrics and Example.

SOP: Professional Skills and Behavioural Skills

Field Process description

Purpose Define work-related Professional Skills and

Behavioural Skills for a specific role before

interviews begin, so that the team assesses the

ability to do the work rather than a CV, charisma

or personal preferences.

Scope Used for new roles, backfills, level changes,

launching a new hiring campaign and reviewing

the scorecard after a calibration meeting. Not

used for adjusting criteria to fit a favoured

candidate.

Roles The hiring manager owns the business outcomes

and must-have task criteria. The recruiter runs

the process and documents the criteria, checking

consistency across stages. Interviewers verify

whether the criteria are observable. Legal / HR

compliance reviews criteria where mandatory.

The role owner or hiring lead approves.

Inputs Draft job description, business outcome of the

role, success expectations for the first 3-6

months, key tasks, team context, failure modes

from past hires, legal or compliance constraints,

compensation / level range, existing competency

framework if available.

Steps 1. Start with the business outcome, not a skills

list. Record what the role must change or

protect in the business. 2. List 5-8 key tasks

where poor performance would be costly. 3.

For each task, define Professional Skills as

practical abilities, not tool names. 4. Define

Behavioural Skills only where they affect

performance in this environment. 5. Add

motivation risks that could lead to failure or

departure of a strong person.

6. Remove criteria that are not linked to work, not

observable or not testable.

7. Translate each criterion into scorecard fields: definition, assessment facts, method and scale

anchors 1 / 3 / 5.

8. Mark threshold criteria separately from weighted ones.

9. Check language for bias risk — for example cultural fit, executive presence, aggressive, native speaker, youthful

energy or similar labels.

10. Fix the criteria before the first interview with the candidate.

SLA / timelinesComplete before sourcing starts where possible; at minimum before the first recruiter screening. For urgent backfill roles, prepare a lightweight version within 1 working day and fix it before the panel interview. Review after every 5-8 interviewed candidates or after a calibration meeting if assessment facts show the criteria are unclear.
Mandatory fieldsRole name, level, business outcome, key tasks, criterion name, criterion type, definition, mandatory assessment facts, method, scale anchor rating 1 / 3 / 5, weight or threshold status, responsible interviewer, prohibited assessment facts, debrief date.
OutputsApproved role scorecard, interview plan, interviewer assignment by criterion, question / case bank, documented thresholds, notes on excluded criteria and why they were excluded.
ExceptionsIf the legal team or compliance require a criterion, mark it as mandatory and note the policy owner. If the role is experimental and tasks are unclear, create provisional criteria and schedule a review after the first hiring cycle. If the hiring manager asks to add a vague criterion, require observable behaviour and a link to role tasks before including it.
MetricsProportion of open roles with approved criteria before the first interview; number of criteria changed after candidates entered the process; interviewer agreement on definitions; percentage of debrief comments linked to criteria; rejected criteria due to weak link to role tasks.
ExampleFor a Customer Success Manager in B2B SaaS, "knows SaaS" is too broad. Convert it into Discovery, Client Risk Diagnosis, Communication, Reliability, Conflict Maturity and Motivation for Client Work. Each criterion receives a method and assessment facts requirement before interviews begin.
SOP: Completing scorecards
FieldProcess description
PurposeEach interviewer records the rating by criterion, the facts and confidence before the debrief, so that the hiring decision rests on observable data rather than memory, status or group influence.
ScopeApplies to recruiter screens, hiring manager interviews, panel interviews, technical and functional interviews, work trials, case studies and reference checks, if they influence the hiring decision.
RolesThe interviewer completes the scorecard. The recruiter checks completeness before the debrief. The hiring manager reads the assessment facts, not just the recommendation on the CV. The interview coordinator or ATS owner ensures the form is available and linked to the right stage.

Example. For a Customer Success Manager in B2B SaaS, "knows SaaS" is too

broad. Convert it into Discovery, Client Risk Diagnosis,

Communication, Reliability, Conflict Maturity and Motivation for

Client Work. Each criterion receives a method and assessment

facts requirement before interviews begin.

SOP: Completing scorecards

Field Process description

Purpose Each interviewer records the rating by criterion, the facts and

confidence before the debrief, so that the hiring decision

rests on observable data rather than memory, status or group

influence.

Scope Applies to recruiter screens, hiring manager interviews, panel

interviews, technical and functional interviews, work trials,

case studies and reference checks, if they influence the

hiring decision.

Roles The interviewer completes the scorecard. The recruiter checks

completeness before the debrief. The hiring manager reads the

assessment facts, not just the recommendation on the CV.

The interview coordinator or ATS owner ensures the form is

available and linked to the right stage.

FieldProcess description
InputsApproved role scorecard, interview stage brief, candidate CV or profile, interview notes, work trial output, case rubric, reference notes where applicable, ATS scorecard form.
Steps1. Open the scorecard before the interview and confirm which criteria you are responsible for. 2. During the interview, record facts, examples, follow-up questions and missing assessment facts. 3. After the interview and before the debrief, complete each assigned criterion. 4. Select a rating of 1, 3 or 5, or an approved intermediate rating, only when it is supported by assessment facts. 5. Assessment facts are mandatory for ratings 1, 3 and 5; strong positive and strong negative ratings require particularly specific notes. 6. Add confidence separately from the rating and explain why the confidence is High, Medium or Low. 7. Move unverified impressions to "questions / missing assessment facts" rather than into the decision recommendation. 8. Flag contradictions and next-step questions explicitly. 9. Submit the scorecard before reading other interviewers' opinions when the ATS allows it. 10. Do not change a submitted rating after the debrief unless the change and reason are documented.
SLA / timelinesComplete within 2 business hours after the interview and always before the debrief. For same-day hiring processes, complete within 30 minutes. Reference checks should be documented before the final decision meeting.
Mandatory fieldsCandidate name or ID, role, stage, interviewer, date, criterion, rating, notes with assessment facts, confidence, missing assessment facts, risks, recommendation if requested by the process, next-step questions, submission time.
OutputsCompleted scorecard in the ATS or shared hiring document, clear assessment facts by criterion, list of missing assessment facts, interviewer recommendation with confidence level, next-step questions for the next stage.
ExceptionsIf a criterion was not assessed, mark "not assessed" and explain why; do not draw conclusions from an irrelevant conversation. If the interview was interrupted, mark confidence as Low and request a next step if the criterion matters. If notes contain sensitive personal data unrelated to the role, remove or escalate per company policy.
MetricsScorecard completion rate before the debrief; percentage of ratings with assessment facts; percentage of unverified impressions moved to missing assessment facts; average time from interview end to scorecard submission; number of debrief delays caused by incomplete scorecards.
ExampleAn interviewer writes "works brilliantly with clients" after a CSM interview. This cannot be used as an assessment fact. A useful note reads: "In a risk case, identified sponsor change, low adoption and missing executive goal; proposed a 48-hour sponsor call, usage analysis and success plan. Rating 5, confidence Medium because no reference check yet."
SOP: Debrief
FieldProcess description
PurposeMake the hiring stage decision by reviewing criteria and assessment facts before opinions, resolving missing assessment facts and assigning a candidate communication owner.

8. Flag contradictions and next-step questions explicitly. 9. Submit the scorecard before reading

other interviewers' opinions when the ATS allows it. 10. Do not change a submitted rating

after the debrief unless the change and reason are documented.

SLA / timelines Complete within 2 business hours after the interview and

always before the debrief. For same-day hiring processes,

complete within 30 minutes. Reference checks should be

documented before the final decision meeting.

Mandatory fields Candidate name or ID, role, stage, interviewer, date,

criterion, rating, notes with assessment facts, confidence,

missing assessment facts, risks, recommendation if

requested by the process, next-step questions,

submission time.

Outputs Completed scorecard in the ATS or shared hiring

document, clear assessment facts by criterion, list of

missing assessment facts, interviewer recommendation

with confidence level, next-step questions for the next

stage.

Exceptions If a criterion was not assessed, mark "not assessed" and

explain why; do not draw conclusions from an irrelevant

conversation. If the interview was interrupted, mark

confidence as Low and request a next step if the criterion

atters. If notes contain sensitive personal data unrelated to

the role, remove or escalate per company policy.

Metrics Scorecard completion rate before the debrief; percentage

of ratings with assessment facts; percentage of unverified

impressions moved to missing assessment facts; average

time from interview end to scorecard submission; number

of debrief delays caused by incomplete scorecards.

Example. An interviewer writes "works brilliantly with clients" after a

CSM interview. This cannot be used as an assessment fact. A

useful note reads: "In a risk case, identified sponsor change,

low adoption and missing executive goal; proposed a 48-hour

sponsor call, usage analysis and success plan. Rating 5,

confidence Medium because no reference check yet."

SOP: Debrief

Field Process description

Purpose Make the hiring stage decision by reviewing criteria and

assessment facts before opinions, resolving missing assessment

facts and assigning a candidate communication owner.

FieldProcess description
ScopeUsed after any stage where multiple people or strong assessment facts influence the decision: after a panel interview review, post-case review, final interview, offer decision or rejection decision.
RolesThe recruiter runs the debrief and protects the order of "assessment facts first." The hiring manager owns the final stage decision within company policy. Interviewers present assessment facts for their assigned criteria. The bar raiser, HRBP, legal or compliance participate when mandatory. The candidate communication owner is named before the meeting ends.
InputsCompleted scorecards submitted before the debrief, role scorecard, interview plan, work trial rubric, candidate stage history, compensation and availability constraints if relevant, open questions from previous stages.
Steps1. Confirm that all mandatory scorecards are submitted. 2. Reiterate the role outcome and threshold criteria. 3. Review criteria before opinions: for each criterion, read the rating, assessment facts and confidence. 4. Discuss the assessment facts per criterion, not general impressions of the candidate. 5. Identify missing assessment facts and decide whether this is material. 6. Separate red flags from weak signals and unverified impressions. 7. Compare contradictory ratings by examining the method and assessment facts quality. 8. Choose one decision option: advance, reject, pause, schedule a next step, make an offer. 9. Assign the owner and deadline for candidate communication. 10. Record the decision rationale in the ATS or hiring document.
SLA / timelinesThe debrief is held within 1 working day after the last mandatory interview for active candidates. Candidate communication must be sent within the company's stated SLA for candidates, typically 1-2 working days after the decision.
Mandatory fieldsCandidate, role, stage, participants, criteria assessed, rating and confidence summary, key assessment facts, missing assessment facts, decision option, rationale, next action, candidate communication owner, communication deadline.
OutputsDocumented stage decision, candidate communication owner, plan for the next stage or rejection rationale, next-action assessment facts request if needed, updates to the ATS status.
ExceptionsIf scorecards are missing, postpone the debrief or explicitly note the absent assessment facts; do not replace them with verbal recollections. If a senior stakeholder states an opinion without assessment facts, the facilitator asks which criterion and which facts support it. If the legal team, compensation or access constraints arise, pause the decision and route the question to the appropriate owner.
MetricsProportion of debriefs with completed scorecards before the meeting; time from final interview to decision; percentage of decisions with documented rationale; number of decisions changed due to late assessment facts; candidate communication SLA adherence.
ExampleAt the final CSM debrief, the team reviews Client Risk Diagnosis before discussing "overall fit." One interviewer gave a 5 from a case study, another gave a 3 from behavioural responses. The team compares assessment facts, notes that the case study is closer to real work and confidence is higher, then advances with a next-step recommendation question on reliability.

3. Review criteria before opinions: for each criterion, read the

rating, assessment facts and confidence. 4. Discuss the assessment

facts per criterion, not general impressions of the candidate. 5.

Identify missing assessment facts and decide whether this is

material. 6. Separate red flags from weak signals and unverified

impressions. 7. Compare contradictory ratings by examining the

method and assessment facts quality. 8. Choose one decision

option: advance, reject, pause, schedule a next step, make an

offer. 9. Assign the owner and deadline for candidate

communication. 10. Record the decision rationale in the ATS or

hiring document.

SLA / timelines The debrief is held within 1 working day after the last mandatory

interview for active candidates. Candidate communication must be

sent within the company's stated SLA for candidates, typically 1-2

working days after the decision.

Mandatory fields Candidate, role, stage, participants, criteria assessed, rating

and confidence summary, key assessment facts, missing

assessment facts, decision option, rationale, next action,

candidate communication owner, communication deadline.

Outputs Documented stage decision, candidate communication owner,

plan for the next stage or rejection rationale, next-action

assessment facts request if needed, updates to the ATS status.

Exceptions If scorecards are missing, postpone the debrief or explicitly note

the absent assessment facts; do not replace them with verbal

recollections. If a senior stakeholder states an opinion without

assessment facts, the facilitator asks which criterion and which

facts support it. If the legal team, compensation or access

constraints arise, pause the decision and route the question to

the appropriate owner.

Metrics Proportion of debriefs with completed scorecards before the

meeting; time from final interview to decision; percentage of

decisions with documented rationale; number of decisions

changed due to late assessment facts; candidate communication

SLA adherence.

Example. At the final CSM debrief, the team reviews Client Risk

Diagnosis before discussing "overall fit." One interviewer gave

a 5 from a case study, another gave a 3 from behavioural

responses. The team compares assessment facts, notes that the

case study is closer to real work and confidence is higher, then

advances with a next-step recommendation question on

reliability.

SOP: Calibration meeting
FieldProcess description
PurposeKeep interviewers aligned on criteria, rating scale anchors and assessment facts quality so the same candidate behaviour is evaluated consistently across interviewers and time.
ScopeUsed by active hiring teams that conduct repeated interviews for similar or identical roles. Mandatory when a new role is launched, interviewers are added, rating discrepancies are high, pass-through looks unusual or candidate feedback suggests an inconsistent process.
RolesThe recruitment lead or recruiter facilitates the meeting. The hiring manager owns the role criteria. Interviewers bring examples and scorecards. The HRBP, talent operations, legal or DEI lead join when fairness, compliance or adverse impact risk needs to be examined.
InputsTwo real candidate cases from recent loops, anonymised where appropriate; completed scorecards; interview notes; pass-through data; offer / rejection results; candidate feedback; current rating scale anchors; interviewer training notes.
Steps1. Schedule a monthly or fortnightly cadence for active hiring teams; use the fortnightly format when hiring volume or rating discrepancies are high. 2. Select two real candidate cases before the meeting: one with a clear decision and one with a borderline or contested decision. 3. Review the scorecard criteria and scale anchors. 4. Compare ratings and assessment facts by criterion. 5. Find where interviewers used different definitions, different assessment fact thresholds or unverified impressions. 6. Update the scale anchors if the original wording is ambiguous. 7. Add training notes for interviewers: common mistakes and more precise follow-up questions. 8. Decide whether the stage or method needs adjustment. 9. Record the changes and the effective date. 10. Communicate the changes to all interviewers before the next cycle.
SLA / timelinesMonthly for stable hiring teams; fortnightly for active or high-volume hiring teams; ad hoc within 5 working days after a serious disagreement, process complaint or significant quality miss.
Mandatory fieldsMeeting date, roles covered, attendees, candidate cases reviewed, criteria discussed, rating differences, assessment facts differences, scale anchor updates, interviewer training notes, process changes, owner, deadline.
OutputsUpdated rating scale anchors, clarified assessment facts standards, interviewer training notes, action items for ATS or interview plan changes, documented calibration history.
ExceptionsDo not use calibration to retroactively adjust criteria to fit a favoured candidate. Do not discuss protected or irrelevant personal characteristics. If cases cannot be shared broadly, use anonymised summaries or a smaller authorised group.
MetricsInterviewer rating variance by criterion; percentage of ratings with sufficient assessment facts; pass-through rate by stage; offer acceptance and pre-hire performance quality where available; number of scale anchor updates; completion rate for interviewer training notes.

3. Review the scorecard criteria and scale anchors. 4. Compare

ratings and assessment facts by criterion. 5. Find where

interviewers used different definitions, different assessment fact

thresholds or unverified impressions. 6. Update the scale anchors

if the original wording is ambiguous. 7. Add training notes for

interviewers: common mistakes and more precise follow-up

questions. 8. Decide whether the stage or method needs

adjustment. 9. Record the changes and the effective date. 10.

Communicate the changes to all interviewers before the next

cycle.

SLA / timelines Monthly for stable hiring teams; fortnightly for active or

high-volume hiring teams; ad hoc within 5 working days after

a serious disagreement, process complaint or significant

quality miss.

Mandatory fields Meeting date, roles covered, attendees, candidate cases

reviewed, criteria discussed, rating differences, assessment

facts differences, scale anchor updates, interviewer training

notes, process changes, owner, deadline.

Outputs Updated rating scale anchors, clarified assessment facts

standards, interviewer training notes, action items for ATS or

interview plan changes, documented calibration history.

Exceptions Do not use calibration to retroactively adjust criteria to fit a

favoured candidate. Do not discuss protected or irrelevant

personal characteristics. If cases cannot be shared broadly,

use anonymised summaries or a smaller authorised group.

Metrics Interviewer rating variance by criterion; percentage of ratings

with sufficient assessment facts; pass-through rate by stage;

offer acceptance and pre-hire performance quality where

available; number of scale anchor updates; completion rate for

interviewer training notes.

FieldProcess description
ExampleThe CSM team reviews two candidates. In both cases, interviewers rated "Communication" differently: one rewarded a polished presentation, the other valued clear risk framing. The team updates the rating-5 anchor: clear status, responsible person, risk and next step are required. Then they add a training note: "Do not rate extroversion as communication."
SOP: Reviewing AI-assisted assessment
FieldProcess description
PurposeUse AI output as a draft preparation and review assistant without transferring hiring, rejection, level, compensation or final assessment decisions to AI.
ScopeApplies when AI summarises interview notes, extracts assessment facts, compares them to criteria, drafts next-step questions, checks scorecard completeness or highlights missing data. It does not permit automatic ranking, rejection, promotion, level or compensation decisions.
RolesThe interviewer owns the raw notes and rating. The recruiter or hiring manager owns use within the process. HR, legal or compliance approves use where mandatory. The AI tool owner manages access, data storage, privacy and model risk controls. A human remains responsible for all hiring decisions.
InputsApproved scorecard, human interview notes, candidate-provided materials permitted by policy, work trial rubric, AI usage policy, data handling rules, criteria list with weights, known missing assessment facts.
Steps1. Treat AI output as a draft or hypothesis, not a fact. 2. Verify each claim against the raw notes. 3. Check alignment with approved criteria: AI must not add new criteria after the interview. 4. Cross-check weights and thresholds against the scorecard. 5. Mark missing data without letting AI infer it. 6. Check for bias risk and sensitive indicators. 7. Do not use AI for hiring, rejection, level, compensation or final assessment decisions. 8. If AI proposes a rating, treat it as a prompt for human verification. 9. Log AI usage if company policy requires it. 10. Escalate bias, privacy or automated-decision risk to the process owner.
SLA / timelinesAI review, if used, runs after human notes are recorded and before the debrief. Human checks must be completed before AI summaries are used in the decision discussion. Policy and legal review must be completed before a new AI process is launched.
Mandatory fieldsTool name if mandatory, use case, source materials used, human reviewer, criteria checked, unconfirmed AI statements removed, missing assessment facts, bias risk check, final human rating, decision owner, audit note if mandatory.
OutputsHuman-validated assessment facts summary, cleaned missing-assessment-facts list, next-step questions, AI usage audit note where mandatory, rejected AI statements or bias-risk notes.

2. Verify each claim against the raw notes. 3. Check alignment

with approved criteria: AI must not add new criteria after the

interview. 4. Cross-check weights and thresholds against the

scorecard. 5. Mark missing data without letting AI infer it. 6.

Check for bias risk and sensitive indicators. 7. Do not use AI

for hiring, rejection, level, compensation or final assessment

decisions. 8. If AI proposes a rating, treat it as a prompt for

human verification. 9. Log AI usage if company policy requires

it. 10. Escalate bias, privacy or automated-decision risk to the

process owner.

SLA / timelines AI review, if used, runs after human notes are recorded and

before the debrief. Human checks must be completed before

AI summaries are used in the decision discussion. Policy and

legal review must be completed before a new AI process is

launched.

Mandatory fields Tool name if mandatory, use case, source materials used,

human reviewer, criteria checked, unconfirmed AI statements

removed, missing assessment facts, bias risk check, final

human rating, decision owner, audit note if mandatory.

Outputs Human-validated assessment facts summary, cleaned

missing-assessment-facts list, next-step questions, AI usage

audit note where mandatory, rejected AI statements or

bias-risk notes.

FieldProcess description
ExceptionsDo not use AI on candidate data if policy, consent, jurisdiction or vendor terms do not permit it. Do not transfer sensitive personal data unless explicitly permitted and necessary. If AI output contradicts human notes, the raw notes take priority. If a stakeholder asks AI to automatically rank candidates, halt the process and route the question to the hiring lead.
MetricsPercentage of AI summaries human-validated before the debrief; number of unconfirmed AI statements removed; number of missing-assessment-facts items identified; bias-risk escalations; AI process policy compliance; decision records showing a human owner.
ExampleAI summarises a CSM interview and writes: "low motivation because the candidate asked about work-life balance." The reviewer removes the unconfirmed inference, keeps the factual quote and adds a next-step question about expectations for client escalations and renewals. The hiring decision is not made on AI output.

8. Cases: How to handle complex situations

Case

1. Strong professional skills, weak behavioural skills

BlockAnalysis
SituationThe candidate confidently handles the technical / functional part: solves the case, demonstrates deep expertise and relevant results. However, in behavioural interviews, weak signals appear in communication, responsibility or handling conflict.
RiskThe team may overvalue professional skills and hire someone who creates rework, hides risks, damages trust or fails to hold stakeholders. Another risk is rejecting a strong expert due to general irritation, without separating trainable gaps from critical work risks.
What to checkWhich behavioural skills are truly needed for this role; how closely the weak signal links to job performance; whether the gap can be offset by management, onboarding or allocation of responsibilities; whether the behaviour is repeated across multiple sources.
Questions / assessment facts"Tell me about a conflict with a colleague or client: what was your contribution, what did you do, what did you change afterwards?" "When was the last time you proactively flagged bad news to a stakeholder?" Look for specifics, timelines, ownership, consequences and the ability to learn.
Recommended actionDo not decide on overall impression alone. Separate the criteria: professional skills rating, behavioural skills rating, confidence and impact on the role. If a soft skill is a threshold criterion, add a next step or reference check. If the gap is trainable, document the onboarding condition and the risk owner.
How to document"Professional skill X is confirmed by assessment facts A/B with high confidence. Soft skill Y has a weak signal: the candidate did not mention early risk escalation in two examples; impact on the role is high because the role requires client escalation. A reference check on reliability is needed."

8. Cases: How to handle complex situations

Case 1. Strong professional skills, weak behavioural skills

Block Analysis

Situation The candidate confidently handles the technical / functional

part: solves the case, demonstrates deep expertise and

relevant results. However, in behavioural interviews, weak

signals appear in communication, responsibility or handling

conflict.

Risk The team may overvalue professional skills and hire someone

who creates rework, hides risks, damages trust or fails to

hold stakeholders. Another risk is rejecting a strong expert

due to general irritation, without separating trainable gaps

from critical work risks.

What to check Which behavioural skills are truly needed for this role; how

closely the weak signal links to job performance; whether the

gap can be offset by management, onboarding or allocation

of responsibilities; whether the behaviour is repeated across

multiple sources.

Questions / assessment facts "Tell me about a conflict with a colleague or client: what was

your contribution, what did you do, what did you change

afterwards?" "When was the last time you proactively flagged

bad news to a stakeholder?" Look for specifics, timelines,

ownership, consequences and the ability to learn.

Recommended action Do not decide on overall impression alone. Separate the

criteria: professional skills rating, behavioural skills rating,

confidence and impact on the role. If a soft skill is a threshold

criterion, add a next step or reference check. If the gap is

trainable, document the onboarding condition and the risk

owner.

How to document "Professional skill X is confirmed by assessment facts A/B

with high confidence. Soft skill Y has a weak signal: the

candidate did not mention early risk escalation in two

examples; impact on the role is high because the role requires

client escalation. A reference check on reliability is needed."

Page 98 Page 99 Page 100 Page 101 Page 102 Page 103 Page 104 Page 105 Page 106 Page 107 Page 108 Page 109 Page 110 Page 111 Page 112 Page 113 Page 114 Page 115 Page 116