Chapter 6. What We Assess


| Autonomy | |
|---|---|
| Block | Content |
| Plain definition | The ability to move work forward without constant instructions, whilst maintaining alignment with goals and constraints. |
| Why it matters | Autonomy reduces management overhead and is especially important where tasks are ambiguous or the manager cannot oversee every step. |
| Where it shows | In ambiguous tasks, independent prioritisation, requesting context, making interim decisions and escalating. |
| What to ask | "Tell me about a task where you were given a goal but no detailed plan. How did you clarify expectations, choose the first steps and know when to escalate?" |
| Strong signals | Clarifies the goal and constraints, proposes a plan, iterates, reports progress and knows the limits of their autonomy. |
| Weak signals | Waits for detailed instructions, does lots of work without validating the approach, confuses autonomy with isolation. |
| Red flags | Ignores alignment, makes irreversible decisions without the necessary approval, hides lack of progress. |
| Scale 1 / 3 / 5 | 1 = stops without instructions or acts randomly; 3 = handles clear tasks independently but gets lost in uncertainty; 5 = turns an ambiguous goal into a plan, validates assumptions, pushes work forward and escalates in good time. |
| Assessment errors | Rewarding people who "don't ask questions" when they may simply be working without alignment and creating risk. |
| Mini case | Situation: manager asks to "improve client onboarding"; risk: the candidate starts with the wrong problem; how to test: ask for the first 48 hours of actions; resolution: a strong answer starts with the goal, metrics, stakeholders and quick fact-checking. |
| Judgement | |
| Block | Content |
| Plain definition | The ability to make sound decisions with incomplete information, considering impact, risk, trade-offs and deadlines. |
| Why it matters | In real work, complete data is rarely available; the quality of decision-making determines which risks the team takes consciously. |
| Where it shows | In prioritisation, escalation, hiring decisions, product choices, exceptions, incident response and resource allocation. |
| What to ask | "Tell me about a decision where you lacked sufficient data. What options did you consider, what risks did you accept and what would you revisit now?" |
| Strong signals | Names the decision criteria, alternatives, assumptions, reversibility, worst-case risk and the trigger for review. |
| Weak signals | Decides by habit or the loudest opinion, conflates urgent with important, cannot see second-order consequences. |
| Block | Content |
|---|---|
| Red flags | Confidently takes high-risk decisions without assessment facts, ignores contrary data, rationalises failure after the fact. |
| Scale 1 / 3 / 5 | 1 = chooses on whim or pressure; 3 = compares obvious options and risks; 5 = clearly states criteria, trade-offs, unknowns, worst-case protection and the review plan. |
| Assessment errors | Confusing decision maturity with confidence, seniority, a loud opinion or the answer matching the interviewer's views. |
| Mini case | Situation: client asks for a non-standard discount to close quickly; risk: margin and precedent; how to test: request a decision memo; resolution: assess criteria, alternatives, the approval route and the consequences. |
| Collaboration | |
| Block | Content |
| Plain definition | The ability to achieve results with other people through clear roles, trust, information sharing and shared responsibility. |
| Why it matters | Most work outcomes are created across functions; poor collaboration turns even strong specialists into a bottleneck. |
| Where it shows | In cross-functional projects, context handover, joint planning, feedback, stakeholder management and shared metrics. |
| What to ask | "Tell me about a project where success depended on another team. How did you agree on roles, expectations and resolve disagreements?" |
| Strong signals | Clarifies the shared goal, roles, dependencies and the right to decide, keeps others in the loop, acknowledges the team's contribution. |
| Weak signals | Works through personal back-channels without transparency, complains about other functions, does not manage dependencies. |
| Red flags | Sabotages shared decisions, takes credit for others' work, creates coalitions instead of solving the problem. |
| Scale 1 / 3 / 5 | 1 = works in isolation and blames other teams; 3 = collaborates normally with clear roles; 5 = builds alignment themselves, clarifies responsibility, reduces friction and helps the group reach the result. |
| Assessment errors | Confusing collaboration with friendliness, conflict-free behaviour or a willingness to help everyone at the expense of their own commitments. |
| Mini case | Situation: sales and product disagree about an urgent feature for a client; risk: conflicting goals; how to test: ask how the candidate would organise the discussion; resolution: a strong answer connects the facts, impact, the decision owner and the next step. |
| Conflict maturity | |
|---|---|
| Block | Content |
| Plain definition | The ability to deal with disagreement directly, respectfully and productively, without resorting to avoidance or personal battle. |
| Why it matters | Conflict is inevitable in complex work; conflict maturity determines whether the decision improves or trust breaks down. |
| Where it shows | In disagreement, feedback, performance conversations, stakeholder conflicts, prioritisation and customer escalations. |
| What to ask | "Think of a strong workplace disagreement. What was the other party's position, how did you verify the facts and how did you reach a resolution?" |
| Strong signals | Can honestly describe the opponent's position, separates people from the problem, seeks criteria, records the decision and next action. |
| Weak signals | Avoids difficult conversations, concedes without discussion, argues through status, only recalls the conflict as someone else's fault. |
| Red flags | Makes it personal, takes revenge after a decision, publicly undermines agreements, uses pressure instead of arguments. |
| Scale 1 / 3 / 5 | 1 = avoids conflict or turns it into a personal battle; 3 = discusses disagreement but needs moderation; 5 = creates a safe structure for the conversation, holds to criteria and drives through to a resolution. |
| Assessment errors | Assuming someone is mature because they never argue; the absence of conflict can mean avoidance, not maturity. |
| Mini case | Situation: the candidate considers a manager's decision wrong; risk: silent disagreement or a public argument; how to test: ask them to describe the conversation; resolution: assess the facts, the timeline, respect and commitment after the decision. |
| Systems thinking | |
| Block | Content |
| Plain definition | The ability to see connections, root causes, consequences and the feedback loop, not just the isolated symptom. |
| Why it matters | Without systems thinking, a person treats symptoms, creates local optimisations and pushes the problem to another part of the process. |
| Where it shows | In process improvement, incident analysis, org design, funnel diagnostics, product decisions and policy changes. |
| What to ask | "Tell me about a recurring problem you addressed. How did you find the systemic root cause, not just fix one case?" |
| Strong signals | Analyses inputs, context handover, incentives, constraints, metrics, unintended consequences and the feedback loop. |
| Weak signals | Sees only the immediate cause, proposes more control without changing the process, does not check side effects. |
| Block | Content |
|---|---|
| Red flags | Creates elaborate diagrams without practical action, ignores people and incentives, protects the system even when assessment facts show failure. |
| Scale 1 / 3 / 5 | 1 = reacts to a symptom with a one-off action; 3 = sees root causes and proposes process improvement; 5 = analyses connections, incentives, risks, metrics and changes the system so the problem does not recur. |
| Assessment errors | Confusing systems thinking with complex jargon, elaborate diagrams or abstract reasoning without a verifiable result. |
| Mini case | Situation: candidates regularly disappear after interviews; risk: the team blames the market; how to test: ask for a funnel diagnosis; resolution: a strong answer examines the timeline, communication, offer fit, interviewer behaviour and data quality. |
| Reliability | |
| Block | Content |
| Plain definition | The ability to consistently deliver on commitments with predictable quality and to flag deviations in advance. |
| Why it matters | Reliability builds trust in the person and the process; without it, the team must constantly re-check the work. |
| Where it shows | In deadlines, recurring tasks, documentation, data hygiene, operational processes, commitments to clients and compliance checks. |
| What to ask | "What recurring work commitments do you typically have each week or month? How do you ensure they are completed on time and without loss of quality?" |
| Strong signals | Uses a tracking system, keeps commitments, flags risk in advance, knows their capacity limits, documents what matters. |
| Weak signals | Relies on memory, frequently explains delays with urgent tasks, does not record agreements in writing. |
| Red flags | Regularly breaks promises without warning, hides status, fakes progress or quality. |
| Scale 1 / 3 / 5 | 1 = commitments are unpredictable and need constant monitoring; 3 = delivers core commitments under normal load; 5 = consistently manages commitments, capacity and risks even through change. |
| Assessment errors | Confusing reliability with the absence of errors; a reliable person can make mistakes but makes them visible and manageable. |
| Mini case | Situation: the role has many routine compliance tasks; risk: one missed check is costly; how to test: ask about their self-monitoring system; resolution: assess routines, reminders, backup, escalation and assessment facts. |
| Adaptability | |
|---|---|
| Block | Content |
| Plain definition | The ability to change approach when new data, conditions or constraints emerge, without losing sight of the goal or quality of work. |
| Why it matters | Markets, clients, priorities and processes change; a strong employee does not stick to an old plan when the context has shifted. |
| Where it shows | In changing priorities, a new manager, customer escalations, market shifts, reorgs, new tools and strategy changes. |
| What to ask | "Tell me about a situation when your original plan stopped working because conditions changed. How did you recognise it and what did you adjust?" |
| Strong signals | Notices new data, revisits assumptions, preserves the goal, involves stakeholders and explains the plan change. |
| Weak signals | Cling to the old approach, changes everything without criteria, sees change only as an external disruption. |
| Red flags | Sabotages new rules, creates chaos under the guise of flexibility, ignores mandatory constraints. |
| Scale 1 / 3 / 5 | 1 = resists change or loses the outcome; 3 = adapts after being told explicitly; 5 = notices the shift themselves, rebuilds the plan, explains trade-offs and stabilises the work of the team or client. |
| Assessment errors | Assuming someone is adaptable because they always agree; adaptability requires criteria, not endless acquiescence. |
| Mini case | Situation: a key client changes requirements a day before launch; risk: the team rushes to redo everything; how to test: ask for a response plan; resolution: a strong answer clarifies the impact, must-have criteria, risk, the decision owner and communication. |
| Stress tolerance | |
| Block | Content |
| Plain definition | The ability to maintain work quality, communication and decision maturity under pressure, without normalising chronic overload. |
| Why it matters | In stressful situations, mistakes, blunt messages and tunnel vision become costlier; the role requires resilience without glorifying burnout. |
| Where it shows | In incidents, peak workloads, client escalations, tight deadlines, uncertainty, public criticism and high-stakes decisions. |
| What to ask | "Describe a period of intense work pressure. How did you prioritise, what did you delegate or escalate and how did you control the quality of decisions?" |
| Strong signals | Narrows focus to critical tasks, states constraints, asks for help in good time, maintains baseline quality and recovery. |
| Weak signals | Works only through overtime, becomes sharp, loses priorities, does not report capacity limits. |
| Block | Content |
|---|---|
| Red flags | Boasts about constant burnout, snaps at people, hides mistakes under pressure, takes risky shortcuts. |
| Scale 1 / 3 / 5 | 1 = under pressure loses communication, quality or ethics; 3 = handles clear stress with support; 5 = preserves decision maturity, prioritises, escalates constraints and helps the system exit overload. |
| Assessment errors | Confusing stress tolerance with willingness to put up with poor management, chronic overtime or a toxic environment. |
| Mini case | Situation: a work incident and an angry client at the same time; risk: chaotic promises; how to test: ask for the first 30 minutes of actions; resolution: assess triage, the owner, communication rhythm, facts and recovery. |
| Ethical behaviour | |
| Block | Content |
| Plain definition | The ability to act honestly, follow the rules and protect trust, even when a quick result can be achieved by cutting corners. |
| Why it matters | Ethical breaches create legal, financial, reputational and team risks that often cost more than the short-term gain. |
| Where it shows | In handling data, client promises, hiring decisions, vendor selection, reporting, expenses, compliance and conflicts of interest. |
| What to ask | "Tell me about a situation where you were expected to deliver a result in a way that felt wrong or risky to you. What did you do?" |
| Strong signals | States the principle or rule, verifies the facts, escalates appropriately, proposes a safe alternative, prepared to lose short-term upside. |
| Weak signals | Reasons only about whether they will be caught, does not know the current rules, justifies grey areas with business pressure. |
| Red flags | Willing to distort data, promise the client the impossible, bypass compliance, use confidential information for unauthorised purposes. |
| Scale 1 / 3 / 5 | 1 = chooses the result even if it requires breaking the rules; 3 = follows explicit rules but gets lost in grey areas; 5 = foresees ethical risk, asks questions, escalates and finds a workable, legitimate option. |
| Assessment errors | Asking abstract questions like "are you an honest person?" instead of concrete situations with pressure, incentives and consequences. |
| Mini case | Situation: a manager asks to "improve" a report before a board meeting; risk: data distortion; how to test: ask for the response and escalation route; resolution: a strong answer clarifies the facts, preserves data integrity and proposes the correct wording. |
| Role Context | Which Behavioural Skills Become Critical | Why |
|---|---|---|
| Junior / trainee | Ability to learn, responsibility, reliability | Mistakes are expected but growth and discipline matter |
| Senior IC | Judgement, systems thinking, communication | The person influences complex decisions without direct authority |
| Manager | Communication, conflict maturity, ethical behaviour, reliability | The manager creates the environment and makes decisions for others |
| Client-facing | Communication, adaptability, stress tolerance | The client environment changes and demands trust |
| Regulated / high-risk | Responsibility, ethical behaviour, reliability | A mistake can create legal, financial or safety risk |
6. Scorecard, scales, weights and confidence
The scorecard is not about turning recruitment into arithmetic. Its purpose is to make candidate assessment reproducible: to pre-determine what matters for the role, how each criterion is tested, what assessment facts are considered sufficient and how confident the team is in the signal. A good scorecard protects the team from three common errors: each interviewer assesses different things, a strong impression replaces facts and the final decision is made before the assessment facts are discussed.
Universal scorecard template
Use one base format for all roles and adapt the criteria content to the role. In the template below, each row is a separate criterion that must be linked to a role task and tested by a specific method. If a criterion cannot be described through assessment facts, it cannot be used for decision-making.
| Criterion | Type | Weight | Definition | Mandatory assessment facts | Method | Rating 1 / 3 / 5 | Confidence | Notes |
| Criterion name | Professional skill / Soft skill / Motivation / Signal reliability | 5-30% | What the candidate must be able to do or demonstrate in the context of this role | Specific facts, examples, artefacts, answers to follow-up questions, work output, candidate's role | Structured interview, work trial, case study, portfolio review, reference check, screening | 1 = insufficient for the role; 3 = working minimum; 5 = strong level for this role | High / Medium / Low and the reason | Questions, risks, contradictions, missing assessment facts |
Minimum completion rules:
1. The criterion must be observable. "Communication" is acceptable if there is a definition; "nice person" or "cultural fit" are not.
2. Type helps avoid mixing different kinds of signal.
Professional skill shows the ability to do the work; Soft skill shows work behaviour; Motivation shows alignment with role expectations; Signal reliability shows how much the evidence can be trusted.
3. Weight reflects the cost of getting the criterion wrong.
The more the criterion affects the first 3-6 months of results, the higher the weight.
4. The definition is written before the interview, not after meeting the candidate.

5. Mandatory assessment facts record what the interviewer must bring to the debrief.
Without assessment facts, the score remains an unverified impression.
6. The method must fit the criterion.
For example, discovery is best tested with a role-play scenario or case study, and reliability with behavioural interviews and references.
7. Rating 1 / 3 / 5 describes the scale anchors.
The interviewer must not invent the scale during the meeting.
Regulated / high-risk Responsibility, ethical behaviour, A mistake can create legal, finan-
reliability cial or safety risk
6. Scorecard, scales, weights and confidence
The scorecard is not about turning recruitment into arithmetic. Its purpose is to make candidate assessment
reproducible: to pre-determine what matters for the role, how each criterion is tested, what assessment facts are
considered sufficient and how confident the team is in the signal. A good scorecard protects the team from three
common errors: each interviewer assesses different things, a strong impression replaces facts and the final
decision is made before the assessment facts are discussed.
Universal scorecard template
Use one base format for all roles and adapt the criteria content to the role. In the template below, each row is
a separate criterion that must be linked to a role task and tested by a specific method. If a criterion cannot be
described through assessment facts, it cannot be used for decision-making.
Criterion Type Weight Definition Mandatory Method Rating 1 / 3 Confidence Notes
assessment / 5
facts
Criterion Professional 5-30% What the candidate Specific Structured 1 = insuffi- High / Questions, risks,
name skill / Soft must be able to facts, ex- interview, cient for Medium / contradictions,
skill / Motiva- do or demonstrate amples, work trial, role; 3 = work- Low and missing assess-
tion / Signal in the context of artefacts, case study, ing minimum; reason ment facts
reliability this role answers to portfolio 5 = strong
follow-up review, level for
questions, reference this role
work output, check,
candidate's role screening
Minimum completion rules:
1. The criterion must be observable. "Communication" is acceptable if there is a definition; "nice person"
or "cultural fit" are not.
2. Type helps avoid mixing different kinds of signal. Professional skill shows the ability to do the
work; Soft skill shows work behaviour; Motivation shows alignment with role expectations; Signal
reliability shows how much the evidence can be trusted.
3. Weight reflects the cost of getting the criterion wrong. The more the criterion affects the first 3-6
months of results, the higher the weight.
4. The definition is written before the interview, not after meeting the candidate.
5. Mandatory assessment facts record what the interviewer must bring to the debrief. Without assessment
facts, the score remains an unverified impression.
6. The method must fit the criterion. For example, discovery is best tested with a role-play scenario or
case study, and reliability with behavioural interviews and references.
7. Rating 1 / 3 / 5 describes the scale anchors. The interviewer must not invent the scale during the meeting.
8. Confidence is filled in separately from the rating. A high rating with low confidence means "it looks
like a strong signal, but there is little evidence."
9. Notes must not turn into a free-form essay. Write facts, short quotes, next-step questions and gaps.
Scale 1 / 3 / 5
Use practical scale anchors, not school grades:
Rating Meaning When to assign What should be in the assessment facts
1 Below working minimum or a risk Candidate does not demonstrate the right A concrete example of failure, absence
for the role behaviour, confuses the basic logic of the of the required action, a contradiction,
criterion, gives only generalities or shows a weak result, inability to explain
a red flag their contribution
3 Sufficient working level Candidate can perform the task in a A relevant work example, a clear
typical context but requires clear conditions, process, acceptable result, statement of
support or does not show consistency in constraints or level
challenging cases
5 Strong level for this role Candidate consistently demonstrates the Multiple relevant examples, complexity,
criterion under challenging conditions, independent role, measurable result,
explains trade-offs, sees risks and can mature explanation of decisions
improve the system around the task
If an intermediate rating of 2 or 4 is needed, use it only when the assessment facts fall between the scale anchors.
Do not turn the scale into "liked them a 4." An intermediate rating must be explainable through the neighbouring
scale anchors.
How to assign weights
Weights are not for pretty maths but for disciplined prioritisation. Without weights, the team often overvalues
bright but secondary signals: charisma, experience at a well-known company, a polished presentation or
resemblance to a successful employee.
Practical rules:
Situation Weighting Decision Example
Criterion is directly linked to the main business 20-30% For a CSM who must reduce churn, customer
result of the role risk diagnosis could carry 25%
Criterion is needed every day but is not the 10-20% Communication for most client-facing
primary differentiator roles
Criterion matters as a minimum threshold Pass / fail or 5-10% Ethical behaviour, legal eligibility,
basic literacy
Criterion can be developed after entry 5-10% Knowledge of a specific internal tool
without high risk
Criterion sounds desirable but is not linked Remove "Start-up mindset" if the role is
to role tasks actually process-driven and predictable
Recommended starting range for a role scorecard:
Type Typical total weight Comment
Professional skills 40-60% Tests the ability to produce a work outcome
Behavioural skills 25-40% Tests behaviour in an environment where
results are created with people
Motivation 10-20% Tests the durability of interest in the real role
Signal reliability Not always a weight, more often a confidence layer Does not replace criteria but shows how
much to trust the assessment facts
Weights must sum to 100% if the team uses a weighted view. But the final decision need not be an automatic
calculation: some criteria are threshold-based and may stop the process regardless of the average score.
Confidence rating
Confidence does not measure the candidate's quality but the quality of the signal. It answers the question:
"How confident are we that the assessment reflects real ability or risk?"
Confidence Meaning When to use
High Assessment facts are relevant to the role, Candidate broke down a similar case,
the next-step questions are specific, there are no explained their role, showed an artefact
major contradictions or completed a work trial
Medium Assessment facts are partially relevant but There is a good example but from a
there are constraints on context, depth or different context; details unverified; the
repeatability method was only an interview
Low Assessment facts are weak, indirect, Candidate spoke in generalities, interviewer
contradictory or obtained through an did not ask follow-up questions, no data on
inappropriate method a complex criterion
Record confidence as High - two relevant examples and a case, Medium - a similar example but
without metrics, Low - self-report only. A single word is not enough.
The difference between the score and confidence
The rating answers the question: "What level did the candidate demonstrate on this criterion?"
Confidence answers the question: "How reliably do we know that?"
Examples:
Situation Rating Confidence What to do
Candidate answers discovery questions 5 Medium Do not lower the rating automatically;
brilliantly but there is no role-play and add targeted verification or a
examples are only from small clients reference check
Candidate performed poorly on a work 1 High Treat as a strong negative assessment
trial for a key task; the method was close fact
to real work
| Situation | Rating | Confidence | What to do |
|---|---|---|---|
| The interviewer feels the candidate is unreliable but the notes only contain "didn't like their style" | Cannot be used | Low | Move to questions / missing assessment facts, not to the decision |
| There is an average behavioural answer but the reference confirms the same pattern | 3 | High | Consider a consistent working level |
When not to average scores. The mean is useful for review but dangerous as an automatic decision. Do not average ratings if:
1. There is a threshold criterion.
For example, ethical behaviour, legal eligibility, mandatory certification, safety, right to work with data.
2. The criteria tested different risk levels.
Weak customer risk diagnosis must not be offset by strong presentation if the role is responsible for churn.
3. Confidence differs significantly.
A single high rating with low confidence must not outweigh several medium high-confidence signals.
4. The methods are not comparable. A work trial on a real task is usually more reliable than a general conversation.
5. There is a red flag.
A red flag must be examined separately: confirm it, dismiss it or halt the process.
6. The role has mandatory constraint criteria.
For example, working hours, travel, language requirements, regulated access, conflict of interest.
7. The ratings reflect different stages.
A screening rating and a final case rating cannot be mechanically added together if they tested different criteria. Use the average only as a summary view after the debrief, not as a substitute for the decision. A proper debrief starts with the assessment facts per criterion, not with an overall candidate ranking. End-to-end example: minimal scorecards for BDM and Backend. Use these tables as a compact version of the scorecard. The full criteria library may be broader but the first working version is best kept short: 5-6 criteria, clear scale anchors and separate confidence.
| Criterion BDM | Weight | 1 | 3 | 5 |
| B2B SaaS qualification | High | Speaks in generalities about sales | Runs a clear qualification logic | Builds a repeatable process and explains trade-offs |
| Discovery quality | High | Presents before understanding the client | Asks baseline questions and records the pain | Identifies business / process / technical blockers and the next step |
| Payment domain learning | Medium | Does not understand the payment context and does not know how to learn | Sees baseline payment constraints | Connects payment flow, risk, integration and buyer concerns |
| Funnel ownership | High | Cannot explain the stages / reasons | Runs the funnel and next actions | Analyses win / loss causes and improves the process |
| Stakeholder communication | Medium | Speaks the same way to | Adapts language to the role | Holds complex |
1. There is a threshold criterion. For example, ethical behaviour, legal eligibility, mandatory
certification, safety, right to work with data.
2. The criteria tested different risk levels. Weak customer risk diagnosis must not be offset by
strong presentation if the role is responsible for churn.
3. Confidence differs significantly. A single high rating with low confidence must not outweigh
several medium high-confidence signals.
4. The methods are not comparable. A work trial on a real task is usually more reliable than a general conversation.
5. There is a red flag. A red flag must be examined separately: confirm it, dismiss it or halt the process.
6. The role has mandatory constraint criteria. For example, working hours, travel, language requirements,
regulated access, conflict of interest.
7. The ratings reflect different stages. A screening rating and a final case rating cannot be mechanically
added together if they tested different criteria.
Use the average only as a summary view after the debrief, not as a substitute for the decision. A proper debrief
starts with the assessment facts per criterion, not with an overall candidate ranking.
End-to-end example: minimal scorecards for BDM and Backend
Use these tables as a compact version of the scorecard. The full criteria library may be broader but the first
working version is best kept short: 5-6 criteria, clear scale anchors and separate confidence.
Criterion BDM Weight 1 3 5
B2B SaaS qualification High Speaks in Runs a clear Builds a
generality qualification repeatable
about sales logic process and
explains
trade-offs
Discovery quality High Presents before Asks baseline Identifies business /
understanding questions process / technical
the client and records blockers and
the pain the next step
Payment domain learning Medium Does not See baseline Connects
understand payment payment flow,
the payment constraints risk, integration
context and and buyer
does not know concerns
how to learn
Funnel ownership High Cannot explain Runs the Analyses win /
the stages / funnel and loss causes and
reasons next actions improves the
process
Stakeholder communication Medium Speaks the same Adapts language Holds complex
way to everyone to the role multi-stakeholder
| Criterion BDM | Weight | 1 | 3 | 5 |
|---|---|---|---|---|
| everyone | stakeholders | multi-stakeholder conversation | ||
| Learning agility | Medium | Becomes defensive about gaps | Acknowledges gaps and learns | Builds their own learning cycle and validates hypotheses |
| Criterion Backend | Weight | 1 | 3 | 5 |
| Backend fundamentals | High | Knows the framework but poorly explains design decisions | Confidently designs typical backend tasks | Holds API, data, async, alignment and evolution |
| Work ownership | High | Did not own the delivery after merge | Participated in release / monitoring | Owned changes end-to-end and learned from incidents |
| Reliability thinking | High | Does not see failure modes | Names typical risks | Proactively designs retries, idempotency, observability, rollback |
| Code quality | Medium | Works but fragile code | Readable code with baseline tests | Maintainable code, good tests, review-friendly decisions |
| Collaboration | Medium | Struggles to explain trade-offs | Works normally with the product / team | Helps the team take balanced decisions |
| Domain learning | Medium | Not interested in the payment context | Willing to learn | Connects product risk and engineering design |
Example scorecard: Customer Success Manager for B2B SaaS
Role context: The CSM manages B2B SaaS clients after the sale, owns adoption, renewal readiness, risk
visibility, escalation management and qualitative handover of context between sales, implementation, support and
product. The role is not a pure support position: the candidate must understand the client's business objectives,
spot risk before churn signals and communicate without making empty promises.
Criterion Type Weight Definition Mandatory Method Rating 1 / 3 Confidence Notes
assessment / 5
facts
Discovery Professional 20% Can surface Can identify Role-play, 1 = asks superficial Fill in post- Check that
skill business objectives, business goals, customer case, questions and interview discovery is
solution criteria, structured interview; immediately sells not replaced
stakeholders, interview the product
success metrics and with follow-up
implementation questions the candidate asks;
constraints how goal, pain, sets; how they record the
barriers and the next step
translates them into a plan
Client risk Professional 25% Early at detecting Real escalation Case interview, 1 = reacts only Fill in post- Key criterion
diagnosis skill churn risk, low or incident review behavioural after complaints interview for the role;
adoption, stakeholder breakdown or case; interview, or renewal risk; weak high-
loss, value gaps or list of risk signals; reference confidence
implementation drift actions in the first checks cannot be
| Criterion | Type | Weight | Definition | Mandatory assessment facts | Method | Rating 1 / 3 / 5 | Confidence | Notes |
|---|---|---|---|---|---|---|---|---|
| and separating symptoms from root cause | 48 hours; who they involve | 3 = spots obvious signals and escalates; 5 = systematically diagnoses root cause, prioritises client portfolios, proposes a risk plan and communication rhythm | offset by charm | |||||
| Communication | Soft skill | 15% | Structurally communicates status, risks, decisions and expectations to the client and internal teams | Example of a difficult letter or conversation; how they adapt the message for the client sponsor, user, product, sales | Structured interview, writing sample, role-play scenario | 1 = speaks generally, over-promises, hides uncertainty; 3 = clearly communicates status and next steps; 5 = makes a complex situation clear, records the responsible person, trade-offs, risks and next action | Fill in post-interview | Assess clarity and ownership, not extroversion |
| Reliability | Soft skill | 15% | Keeps commitments, next actions, account notes, internal context handover and escalations without constant oversight | Example of a period with many clients; tracking system; reference feedback on follow-through | Behavioural interview, reference check | 1 = forgets commitments or blames external causes; 3 = reliable under normal load and clear processes; 5 = builds a next-action system, flags risks in advance, restores trust after a failure | Fill in post-interview | Test repeatability, not a single heroic case |
| Conflict maturity | Soft skill | 10% | Able to handle client or internal conflict without defensiveness, blame or empty concessions | Situation with an angry customer, disagreement with sales or product, escalation debrief | Behavioural interview, role-play scenario | 1 = argues, blames or promises the impossible; 3 = stays calm and looks for a solution; 5 = validates the problem, separates facts from emotions, agrees on options, boundaries and the responsible person | Fill in post-interview | Important for renewals and escalations |
| Motivation for client work | Motivation | 15% | Understands the reality of the CSM role: repetitive next actions, difficult clients, product implementation work, business reviews, pre-renewal pressure | Reasons for choosing the role; what energises and what drains; examples of sustained interest in client outcomes | Screening, motivation interview, reference check | 1 = wants the role as a bridge into strategy or product without client routine; 3 = accepts client work and understands the main challenges; 5 = consciously chooses long-term client outcomes, can work | Fill in post-interview | Do not confuse motivation with an energetic interview style |
| Criterion | Type | Weight | Definition | Mandatory assessment facts | Method | Rating 1 / 3 / 5 | Confidence | Notes |
|---|---|---|---|---|---|---|---|---|
| with routine and difficult conversations |
The decision based on this scorecard must not be taken on a single average number. For this role, customer risk diagnosis is a near-threshold criterion: if the candidate scores 1 with high confidence, the process usually needs to be halted or a targeted check set only with very strong grounds. For motivation, a score of 3 is sufficient if professional skills are strong and role expectations have been discussed honestly. For communication and conflict maturity, it is important to look not only at the rating but also at the assessment facts: a calmer, less flashy candidate may be stronger than a candidate with polished but imprecise speech.
7. SOP: How to embed candidate assessment in the team
The SOPs in this section are designed as process documents that can be copied into a corporate knowledge base, ATS working guide or hiring operations manual and adapted to the role, stage, local legal requirements, internal policies and the specific ATS. They do not constitute legal advice. Before use in regulated hiring, mass recruitment, international processes or AI-assisted workflows, they must be reviewed with a legal or compliance lead. Each SOP below answers one operational question: who does what, when, on what inputs, which fields are mandatory, what counts as output and how to measure the quality of the process. If the company already has its own SOP format, transfer the content into it but do not remove the fields: Purpose, Scope, Roles, Inputs, Steps, SLA / timelines, Mandatory fields, Outputs, Exceptions, Metrics and Example. SOP: Professional Skills and Behavioural Skills
| Field | Process description |
| Purpose | Define work-related Professional Skills and Behavioural Skills for a specific role before interviews begin, so that the team assesses the ability to do the work rather than a CV, charisma or personal preferences. |
| Scope | Used for new roles, backfills, level changes, launching a new hiring campaign and reviewing the scorecard after a calibration meeting. Not used for adjusting criteria to fit a favoured candidate. |
| Roles | The hiring manager owns the business outcomes and must-have task criteria. The recruiter runs the process and documents the criteria, checking consistency across stages. Interviewers verify whether the criteria are observable. Legal / HR compliance reviews criteria where mandatory. The role owner or hiring lead approves. |
| Inputs | Draft job description, business outcome of the role, success expectations for the first 3-6 months, key tasks, team context, failure modes from past hires, legal or compliance constraints, compensation / level range, existing competency framework if available. |
| Steps | 1. Start with the business outcome, not a skills list. Record what the role must change or protect in the business. 2. List 5-8 key tasks where poor performance would be costly. 3. For each task, define Professional Skills as practical abilities, not tool names. 4. Define Behavioural Skills only where they affect performance in this environment. 5. Add motivation risks that could lead to failure or departure of a strong person. 6. Remove criteria that are not linked to work, not |
a score of 3 is sufficient if professional skills are strong and role expectations have been discussed honestly. For
communication and conflict maturity, it is important to look not only at the rating but also at the assessment
facts: a calmer, less flashy candidate may be stronger than a candidate with polished but imprecise speech.
7. SOP: How to embed candidate assessment in the team
The SOPs in this section are designed as process documents that can be copied into a corporate
knowledge base, ATS working guide or hiring operations manual and adapted to the role, stage, local
legal requirements, internal policies and the specific ATS. They do not constitute legal advice. Before
use in regulated hiring, mass recruitment, international processes or AI-assisted workflows, they must
be reviewed with a legal or compliance lead.
Each SOP below answers one operational question: who does what, when, on what inputs, which fields
are mandatory, what counts as output and how to measure the quality of the process. If the company
already has its own SOP format, transfer the content into it but do not remove the fields: Purpose,
Scope, Roles, Inputs, Steps, SLA / timelines, Mandatory fields, Outputs, Exceptions, Metrics and Example.
SOP: Professional Skills and Behavioural Skills
Field Process description
Purpose Define work-related Professional Skills and
Behavioural Skills for a specific role before
interviews begin, so that the team assesses the
ability to do the work rather than a CV, charisma
or personal preferences.
Scope Used for new roles, backfills, level changes,
launching a new hiring campaign and reviewing
the scorecard after a calibration meeting. Not
used for adjusting criteria to fit a favoured
candidate.
Roles The hiring manager owns the business outcomes
and must-have task criteria. The recruiter runs
the process and documents the criteria, checking
consistency across stages. Interviewers verify
whether the criteria are observable. Legal / HR
compliance reviews criteria where mandatory.
The role owner or hiring lead approves.
Inputs Draft job description, business outcome of the
role, success expectations for the first 3-6
months, key tasks, team context, failure modes
from past hires, legal or compliance constraints,
compensation / level range, existing competency
framework if available.
Steps 1. Start with the business outcome, not a skills
list. Record what the role must change or
protect in the business. 2. List 5-8 key tasks
where poor performance would be costly. 3.
For each task, define Professional Skills as
practical abilities, not tool names. 4. Define
Behavioural Skills only where they affect
performance in this environment. 5. Add
motivation risks that could lead to failure or
departure of a strong person.
6. Remove criteria that are not linked to work, not
observable or not testable.
7. Translate each criterion into scorecard fields: definition, assessment facts, method and scale
anchors 1 / 3 / 5.
8. Mark threshold criteria separately from weighted ones.
9. Check language for bias risk — for example cultural fit, executive presence, aggressive, native speaker, youthful
energy or similar labels.
10. Fix the criteria before the first interview with the candidate.
| SLA / timelines | Complete before sourcing starts where possible; at minimum before the first recruiter screening. For urgent backfill roles, prepare a lightweight version within 1 working day and fix it before the panel interview. Review after every 5-8 interviewed candidates or after a calibration meeting if assessment facts show the criteria are unclear. |
| Mandatory fields | Role name, level, business outcome, key tasks, criterion name, criterion type, definition, mandatory assessment facts, method, scale anchor rating 1 / 3 / 5, weight or threshold status, responsible interviewer, prohibited assessment facts, debrief date. |
| Outputs | Approved role scorecard, interview plan, interviewer assignment by criterion, question / case bank, documented thresholds, notes on excluded criteria and why they were excluded. |
| Exceptions | If the legal team or compliance require a criterion, mark it as mandatory and note the policy owner. If the role is experimental and tasks are unclear, create provisional criteria and schedule a review after the first hiring cycle. If the hiring manager asks to add a vague criterion, require observable behaviour and a link to role tasks before including it. |
| Metrics | Proportion of open roles with approved criteria before the first interview; number of criteria changed after candidates entered the process; interviewer agreement on definitions; percentage of debrief comments linked to criteria; rejected criteria due to weak link to role tasks. |
| Example | For a Customer Success Manager in B2B SaaS, "knows SaaS" is too broad. Convert it into Discovery, Client Risk Diagnosis, Communication, Reliability, Conflict Maturity and Motivation for Client Work. Each criterion receives a method and assessment facts requirement before interviews begin. |
| SOP: Completing scorecards | |
| Field | Process description |
| Purpose | Each interviewer records the rating by criterion, the facts and confidence before the debrief, so that the hiring decision rests on observable data rather than memory, status or group influence. |
| Scope | Applies to recruiter screens, hiring manager interviews, panel interviews, technical and functional interviews, work trials, case studies and reference checks, if they influence the hiring decision. |
| Roles | The interviewer completes the scorecard. The recruiter checks completeness before the debrief. The hiring manager reads the assessment facts, not just the recommendation on the CV. The interview coordinator or ATS owner ensures the form is available and linked to the right stage. |
Example. For a Customer Success Manager in B2B SaaS, "knows SaaS" is too
broad. Convert it into Discovery, Client Risk Diagnosis,
Communication, Reliability, Conflict Maturity and Motivation for
Client Work. Each criterion receives a method and assessment
facts requirement before interviews begin.
SOP: Completing scorecards
Field Process description
Purpose Each interviewer records the rating by criterion, the facts and
confidence before the debrief, so that the hiring decision
rests on observable data rather than memory, status or group
influence.
Scope Applies to recruiter screens, hiring manager interviews, panel
interviews, technical and functional interviews, work trials,
case studies and reference checks, if they influence the
hiring decision.
Roles The interviewer completes the scorecard. The recruiter checks
completeness before the debrief. The hiring manager reads the
assessment facts, not just the recommendation on the CV.
The interview coordinator or ATS owner ensures the form is
available and linked to the right stage.
| Field | Process description |
|---|---|
| Inputs | Approved role scorecard, interview stage brief, candidate CV or profile, interview notes, work trial output, case rubric, reference notes where applicable, ATS scorecard form. |
| Steps | 1. Open the scorecard before the interview and confirm which criteria you are responsible for. 2. During the interview, record facts, examples, follow-up questions and missing assessment facts. 3. After the interview and before the debrief, complete each assigned criterion. 4. Select a rating of 1, 3 or 5, or an approved intermediate rating, only when it is supported by assessment facts. 5. Assessment facts are mandatory for ratings 1, 3 and 5; strong positive and strong negative ratings require particularly specific notes. 6. Add confidence separately from the rating and explain why the confidence is High, Medium or Low. 7. Move unverified impressions to "questions / missing assessment facts" rather than into the decision recommendation. 8. Flag contradictions and next-step questions explicitly. 9. Submit the scorecard before reading other interviewers' opinions when the ATS allows it. 10. Do not change a submitted rating after the debrief unless the change and reason are documented. |
| SLA / timelines | Complete within 2 business hours after the interview and always before the debrief. For same-day hiring processes, complete within 30 minutes. Reference checks should be documented before the final decision meeting. |
| Mandatory fields | Candidate name or ID, role, stage, interviewer, date, criterion, rating, notes with assessment facts, confidence, missing assessment facts, risks, recommendation if requested by the process, next-step questions, submission time. |
| Outputs | Completed scorecard in the ATS or shared hiring document, clear assessment facts by criterion, list of missing assessment facts, interviewer recommendation with confidence level, next-step questions for the next stage. |
| Exceptions | If a criterion was not assessed, mark "not assessed" and explain why; do not draw conclusions from an irrelevant conversation. If the interview was interrupted, mark confidence as Low and request a next step if the criterion matters. If notes contain sensitive personal data unrelated to the role, remove or escalate per company policy. |
| Metrics | Scorecard completion rate before the debrief; percentage of ratings with assessment facts; percentage of unverified impressions moved to missing assessment facts; average time from interview end to scorecard submission; number of debrief delays caused by incomplete scorecards. |
| Example | An interviewer writes "works brilliantly with clients" after a CSM interview. This cannot be used as an assessment fact. A useful note reads: "In a risk case, identified sponsor change, low adoption and missing executive goal; proposed a 48-hour sponsor call, usage analysis and success plan. Rating 5, confidence Medium because no reference check yet." |
| SOP: Debrief | |
| Field | Process description |
| Purpose | Make the hiring stage decision by reviewing criteria and assessment facts before opinions, resolving missing assessment facts and assigning a candidate communication owner. |
8. Flag contradictions and next-step questions explicitly. 9. Submit the scorecard before reading
other interviewers' opinions when the ATS allows it. 10. Do not change a submitted rating
after the debrief unless the change and reason are documented.
SLA / timelines Complete within 2 business hours after the interview and
always before the debrief. For same-day hiring processes,
complete within 30 minutes. Reference checks should be
documented before the final decision meeting.
Mandatory fields Candidate name or ID, role, stage, interviewer, date,
criterion, rating, notes with assessment facts, confidence,
missing assessment facts, risks, recommendation if
requested by the process, next-step questions,
submission time.
Outputs Completed scorecard in the ATS or shared hiring
document, clear assessment facts by criterion, list of
missing assessment facts, interviewer recommendation
with confidence level, next-step questions for the next
stage.
Exceptions If a criterion was not assessed, mark "not assessed" and
explain why; do not draw conclusions from an irrelevant
conversation. If the interview was interrupted, mark
confidence as Low and request a next step if the criterion
atters. If notes contain sensitive personal data unrelated to
the role, remove or escalate per company policy.
Metrics Scorecard completion rate before the debrief; percentage
of ratings with assessment facts; percentage of unverified
impressions moved to missing assessment facts; average
time from interview end to scorecard submission; number
of debrief delays caused by incomplete scorecards.
Example. An interviewer writes "works brilliantly with clients" after a
CSM interview. This cannot be used as an assessment fact. A
useful note reads: "In a risk case, identified sponsor change,
low adoption and missing executive goal; proposed a 48-hour
sponsor call, usage analysis and success plan. Rating 5,
confidence Medium because no reference check yet."
SOP: Debrief
Field Process description
Purpose Make the hiring stage decision by reviewing criteria and
assessment facts before opinions, resolving missing assessment
facts and assigning a candidate communication owner.
| Field | Process description |
|---|---|
| Scope | Used after any stage where multiple people or strong assessment facts influence the decision: after a panel interview review, post-case review, final interview, offer decision or rejection decision. |
| Roles | The recruiter runs the debrief and protects the order of "assessment facts first." The hiring manager owns the final stage decision within company policy. Interviewers present assessment facts for their assigned criteria. The bar raiser, HRBP, legal or compliance participate when mandatory. The candidate communication owner is named before the meeting ends. |
| Inputs | Completed scorecards submitted before the debrief, role scorecard, interview plan, work trial rubric, candidate stage history, compensation and availability constraints if relevant, open questions from previous stages. |
| Steps | 1. Confirm that all mandatory scorecards are submitted. 2. Reiterate the role outcome and threshold criteria. 3. Review criteria before opinions: for each criterion, read the rating, assessment facts and confidence. 4. Discuss the assessment facts per criterion, not general impressions of the candidate. 5. Identify missing assessment facts and decide whether this is material. 6. Separate red flags from weak signals and unverified impressions. 7. Compare contradictory ratings by examining the method and assessment facts quality. 8. Choose one decision option: advance, reject, pause, schedule a next step, make an offer. 9. Assign the owner and deadline for candidate communication. 10. Record the decision rationale in the ATS or hiring document. |
| SLA / timelines | The debrief is held within 1 working day after the last mandatory interview for active candidates. Candidate communication must be sent within the company's stated SLA for candidates, typically 1-2 working days after the decision. |
| Mandatory fields | Candidate, role, stage, participants, criteria assessed, rating and confidence summary, key assessment facts, missing assessment facts, decision option, rationale, next action, candidate communication owner, communication deadline. |
| Outputs | Documented stage decision, candidate communication owner, plan for the next stage or rejection rationale, next-action assessment facts request if needed, updates to the ATS status. |
| Exceptions | If scorecards are missing, postpone the debrief or explicitly note the absent assessment facts; do not replace them with verbal recollections. If a senior stakeholder states an opinion without assessment facts, the facilitator asks which criterion and which facts support it. If the legal team, compensation or access constraints arise, pause the decision and route the question to the appropriate owner. |
| Metrics | Proportion of debriefs with completed scorecards before the meeting; time from final interview to decision; percentage of decisions with documented rationale; number of decisions changed due to late assessment facts; candidate communication SLA adherence. |
| Example | At the final CSM debrief, the team reviews Client Risk Diagnosis before discussing "overall fit." One interviewer gave a 5 from a case study, another gave a 3 from behavioural responses. The team compares assessment facts, notes that the case study is closer to real work and confidence is higher, then advances with a next-step recommendation question on reliability. |
3. Review criteria before opinions: for each criterion, read the
rating, assessment facts and confidence. 4. Discuss the assessment
facts per criterion, not general impressions of the candidate. 5.
Identify missing assessment facts and decide whether this is
material. 6. Separate red flags from weak signals and unverified
impressions. 7. Compare contradictory ratings by examining the
method and assessment facts quality. 8. Choose one decision
option: advance, reject, pause, schedule a next step, make an
offer. 9. Assign the owner and deadline for candidate
communication. 10. Record the decision rationale in the ATS or
hiring document.
SLA / timelines The debrief is held within 1 working day after the last mandatory
interview for active candidates. Candidate communication must be
sent within the company's stated SLA for candidates, typically 1-2
working days after the decision.
Mandatory fields Candidate, role, stage, participants, criteria assessed, rating
and confidence summary, key assessment facts, missing
assessment facts, decision option, rationale, next action,
candidate communication owner, communication deadline.
Outputs Documented stage decision, candidate communication owner,
plan for the next stage or rejection rationale, next-action
assessment facts request if needed, updates to the ATS status.
Exceptions If scorecards are missing, postpone the debrief or explicitly note
the absent assessment facts; do not replace them with verbal
recollections. If a senior stakeholder states an opinion without
assessment facts, the facilitator asks which criterion and which
facts support it. If the legal team, compensation or access
constraints arise, pause the decision and route the question to
the appropriate owner.
Metrics Proportion of debriefs with completed scorecards before the
meeting; time from final interview to decision; percentage of
decisions with documented rationale; number of decisions
changed due to late assessment facts; candidate communication
SLA adherence.
Example. At the final CSM debrief, the team reviews Client Risk
Diagnosis before discussing "overall fit." One interviewer gave
a 5 from a case study, another gave a 3 from behavioural
responses. The team compares assessment facts, notes that the
case study is closer to real work and confidence is higher, then
advances with a next-step recommendation question on
reliability.
| SOP: Calibration meeting | |
|---|---|
| Field | Process description |
| Purpose | Keep interviewers aligned on criteria, rating scale anchors and assessment facts quality so the same candidate behaviour is evaluated consistently across interviewers and time. |
| Scope | Used by active hiring teams that conduct repeated interviews for similar or identical roles. Mandatory when a new role is launched, interviewers are added, rating discrepancies are high, pass-through looks unusual or candidate feedback suggests an inconsistent process. |
| Roles | The recruitment lead or recruiter facilitates the meeting. The hiring manager owns the role criteria. Interviewers bring examples and scorecards. The HRBP, talent operations, legal or DEI lead join when fairness, compliance or adverse impact risk needs to be examined. |
| Inputs | Two real candidate cases from recent loops, anonymised where appropriate; completed scorecards; interview notes; pass-through data; offer / rejection results; candidate feedback; current rating scale anchors; interviewer training notes. |
| Steps | 1. Schedule a monthly or fortnightly cadence for active hiring teams; use the fortnightly format when hiring volume or rating discrepancies are high. 2. Select two real candidate cases before the meeting: one with a clear decision and one with a borderline or contested decision. 3. Review the scorecard criteria and scale anchors. 4. Compare ratings and assessment facts by criterion. 5. Find where interviewers used different definitions, different assessment fact thresholds or unverified impressions. 6. Update the scale anchors if the original wording is ambiguous. 7. Add training notes for interviewers: common mistakes and more precise follow-up questions. 8. Decide whether the stage or method needs adjustment. 9. Record the changes and the effective date. 10. Communicate the changes to all interviewers before the next cycle. |
| SLA / timelines | Monthly for stable hiring teams; fortnightly for active or high-volume hiring teams; ad hoc within 5 working days after a serious disagreement, process complaint or significant quality miss. |
| Mandatory fields | Meeting date, roles covered, attendees, candidate cases reviewed, criteria discussed, rating differences, assessment facts differences, scale anchor updates, interviewer training notes, process changes, owner, deadline. |
| Outputs | Updated rating scale anchors, clarified assessment facts standards, interviewer training notes, action items for ATS or interview plan changes, documented calibration history. |
| Exceptions | Do not use calibration to retroactively adjust criteria to fit a favoured candidate. Do not discuss protected or irrelevant personal characteristics. If cases cannot be shared broadly, use anonymised summaries or a smaller authorised group. |
| Metrics | Interviewer rating variance by criterion; percentage of ratings with sufficient assessment facts; pass-through rate by stage; offer acceptance and pre-hire performance quality where available; number of scale anchor updates; completion rate for interviewer training notes. |
3. Review the scorecard criteria and scale anchors. 4. Compare
ratings and assessment facts by criterion. 5. Find where
interviewers used different definitions, different assessment fact
thresholds or unverified impressions. 6. Update the scale anchors
if the original wording is ambiguous. 7. Add training notes for
interviewers: common mistakes and more precise follow-up
questions. 8. Decide whether the stage or method needs
adjustment. 9. Record the changes and the effective date. 10.
Communicate the changes to all interviewers before the next
cycle.
SLA / timelines Monthly for stable hiring teams; fortnightly for active or
high-volume hiring teams; ad hoc within 5 working days after
a serious disagreement, process complaint or significant
quality miss.
Mandatory fields Meeting date, roles covered, attendees, candidate cases
reviewed, criteria discussed, rating differences, assessment
facts differences, scale anchor updates, interviewer training
notes, process changes, owner, deadline.
Outputs Updated rating scale anchors, clarified assessment facts
standards, interviewer training notes, action items for ATS or
interview plan changes, documented calibration history.
Exceptions Do not use calibration to retroactively adjust criteria to fit a
favoured candidate. Do not discuss protected or irrelevant
personal characteristics. If cases cannot be shared broadly,
use anonymised summaries or a smaller authorised group.
Metrics Interviewer rating variance by criterion; percentage of ratings
with sufficient assessment facts; pass-through rate by stage;
offer acceptance and pre-hire performance quality where
available; number of scale anchor updates; completion rate for
interviewer training notes.
| Field | Process description |
|---|---|
| Example | The CSM team reviews two candidates. In both cases, interviewers rated "Communication" differently: one rewarded a polished presentation, the other valued clear risk framing. The team updates the rating-5 anchor: clear status, responsible person, risk and next step are required. Then they add a training note: "Do not rate extroversion as communication." |
| SOP: Reviewing AI-assisted assessment | |
| Field | Process description |
| Purpose | Use AI output as a draft preparation and review assistant without transferring hiring, rejection, level, compensation or final assessment decisions to AI. |
| Scope | Applies when AI summarises interview notes, extracts assessment facts, compares them to criteria, drafts next-step questions, checks scorecard completeness or highlights missing data. It does not permit automatic ranking, rejection, promotion, level or compensation decisions. |
| Roles | The interviewer owns the raw notes and rating. The recruiter or hiring manager owns use within the process. HR, legal or compliance approves use where mandatory. The AI tool owner manages access, data storage, privacy and model risk controls. A human remains responsible for all hiring decisions. |
| Inputs | Approved scorecard, human interview notes, candidate-provided materials permitted by policy, work trial rubric, AI usage policy, data handling rules, criteria list with weights, known missing assessment facts. |
| Steps | 1. Treat AI output as a draft or hypothesis, not a fact. 2. Verify each claim against the raw notes. 3. Check alignment with approved criteria: AI must not add new criteria after the interview. 4. Cross-check weights and thresholds against the scorecard. 5. Mark missing data without letting AI infer it. 6. Check for bias risk and sensitive indicators. 7. Do not use AI for hiring, rejection, level, compensation or final assessment decisions. 8. If AI proposes a rating, treat it as a prompt for human verification. 9. Log AI usage if company policy requires it. 10. Escalate bias, privacy or automated-decision risk to the process owner. |
| SLA / timelines | AI review, if used, runs after human notes are recorded and before the debrief. Human checks must be completed before AI summaries are used in the decision discussion. Policy and legal review must be completed before a new AI process is launched. |
| Mandatory fields | Tool name if mandatory, use case, source materials used, human reviewer, criteria checked, unconfirmed AI statements removed, missing assessment facts, bias risk check, final human rating, decision owner, audit note if mandatory. |
| Outputs | Human-validated assessment facts summary, cleaned missing-assessment-facts list, next-step questions, AI usage audit note where mandatory, rejected AI statements or bias-risk notes. |
2. Verify each claim against the raw notes. 3. Check alignment
with approved criteria: AI must not add new criteria after the
interview. 4. Cross-check weights and thresholds against the
scorecard. 5. Mark missing data without letting AI infer it. 6.
Check for bias risk and sensitive indicators. 7. Do not use AI
for hiring, rejection, level, compensation or final assessment
decisions. 8. If AI proposes a rating, treat it as a prompt for
human verification. 9. Log AI usage if company policy requires
it. 10. Escalate bias, privacy or automated-decision risk to the
process owner.
SLA / timelines AI review, if used, runs after human notes are recorded and
before the debrief. Human checks must be completed before
AI summaries are used in the decision discussion. Policy and
legal review must be completed before a new AI process is
launched.
Mandatory fields Tool name if mandatory, use case, source materials used,
human reviewer, criteria checked, unconfirmed AI statements
removed, missing assessment facts, bias risk check, final
human rating, decision owner, audit note if mandatory.
Outputs Human-validated assessment facts summary, cleaned
missing-assessment-facts list, next-step questions, AI usage
audit note where mandatory, rejected AI statements or
bias-risk notes.
| Field | Process description |
|---|---|
| Exceptions | Do not use AI on candidate data if policy, consent, jurisdiction or vendor terms do not permit it. Do not transfer sensitive personal data unless explicitly permitted and necessary. If AI output contradicts human notes, the raw notes take priority. If a stakeholder asks AI to automatically rank candidates, halt the process and route the question to the hiring lead. |
| Metrics | Percentage of AI summaries human-validated before the debrief; number of unconfirmed AI statements removed; number of missing-assessment-facts items identified; bias-risk escalations; AI process policy compliance; decision records showing a human owner. |
| Example | AI summarises a CSM interview and writes: "low motivation because the candidate asked about work-life balance." The reviewer removes the unconfirmed inference, keeps the factual quote and adds a next-step question about expectations for client escalations and renewals. The hiring decision is not made on AI output. |
8. Cases: How to handle complex situations
Case
1. Strong professional skills, weak behavioural skills
| Block | Analysis |
| Situation | The candidate confidently handles the technical / functional part: solves the case, demonstrates deep expertise and relevant results. However, in behavioural interviews, weak signals appear in communication, responsibility or handling conflict. |
| Risk | The team may overvalue professional skills and hire someone who creates rework, hides risks, damages trust or fails to hold stakeholders. Another risk is rejecting a strong expert due to general irritation, without separating trainable gaps from critical work risks. |
| What to check | Which behavioural skills are truly needed for this role; how closely the weak signal links to job performance; whether the gap can be offset by management, onboarding or allocation of responsibilities; whether the behaviour is repeated across multiple sources. |
| Questions / assessment facts | "Tell me about a conflict with a colleague or client: what was your contribution, what did you do, what did you change afterwards?" "When was the last time you proactively flagged bad news to a stakeholder?" Look for specifics, timelines, ownership, consequences and the ability to learn. |
| Recommended action | Do not decide on overall impression alone. Separate the criteria: professional skills rating, behavioural skills rating, confidence and impact on the role. If a soft skill is a threshold criterion, add a next step or reference check. If the gap is trainable, document the onboarding condition and the risk owner. |
| How to document | "Professional skill X is confirmed by assessment facts A/B with high confidence. Soft skill Y has a weak signal: the candidate did not mention early risk escalation in two examples; impact on the role is high because the role requires client escalation. A reference check on reliability is needed." |
8. Cases: How to handle complex situations
Case 1. Strong professional skills, weak behavioural skills
Block Analysis
Situation The candidate confidently handles the technical / functional
part: solves the case, demonstrates deep expertise and
relevant results. However, in behavioural interviews, weak
signals appear in communication, responsibility or handling
conflict.
Risk The team may overvalue professional skills and hire someone
who creates rework, hides risks, damages trust or fails to
hold stakeholders. Another risk is rejecting a strong expert
due to general irritation, without separating trainable gaps
from critical work risks.
What to check Which behavioural skills are truly needed for this role; how
closely the weak signal links to job performance; whether the
gap can be offset by management, onboarding or allocation
of responsibilities; whether the behaviour is repeated across
multiple sources.
Questions / assessment facts "Tell me about a conflict with a colleague or client: what was
your contribution, what did you do, what did you change
afterwards?" "When was the last time you proactively flagged
bad news to a stakeholder?" Look for specifics, timelines,
ownership, consequences and the ability to learn.
Recommended action Do not decide on overall impression alone. Separate the
criteria: professional skills rating, behavioural skills rating,
confidence and impact on the role. If a soft skill is a threshold
criterion, add a next step or reference check. If the gap is
trainable, document the onboarding condition and the risk
owner.
How to document "Professional skill X is confirmed by assessment facts A/B
with high confidence. Soft skill Y has a weak signal: the
candidate did not mention early risk escalation in two
examples; impact on the role is high because the role requires
client escalation. A reference check on reliability is needed."