Performance management

Designing a performance management system people trust

How to design a performance management system that employees and managers trust, covering purpose, goals, ratings, calibration, rewards and software choices.

Key points

  • Decide what the system is for before designing forms or buying software, because development and pay decisions pull the design in different directions.
  • Trust comes from clear goals, consistent manager judgement and a visible, fair link to consequences, not from the sophistication of the tool.
  • Calibration is the single most useful control for fairness, and forced distribution is a blunt substitute for it.
  • Software should automate a design you have already agreed, not supply the design for you.

Ask employees in most organisations what they think of performance management and you will hear the same things: goals set late, reviews written in a rush, ratings that depend on which manager you have, and a bonus outcome that seems unrelated to all of it. Ask managers and they will tell you it takes too long and changes nothing. Both are usually right.

The problem is rarely the form or the software. It is that the system was never designed as a system. This article sets out the design choices that matter, in the order I would make them, with a design checklist you can use to review your current approach or build a new one.

Start with what the system is for

A performance management system can serve several purposes: aligning effort with strategy, developing people, differentiating pay, informing promotion and succession, and managing underperformance. These purposes pull in different directions.

Development needs honest conversations about weaknesses. Pay differentiation needs defensible distinctions between people. When the same conversation must do both, employees understandably spend it defending their rating rather than discussing how to improve. That is not a character flaw. It is a design outcome.

So the first design decision is to rank the purposes. If pay differentiation is essential, and in most private sector organisations it is, then design for it deliberately and protect development by holding some conversations separately from the rating discussion. If your entity is a government body where pay is largely fixed by grade, you have more room to design for development and alignment, and the rating matters less.

A performance system that tries to do everything with one conversation ends up doing none of it well.

Goals that people can actually be judged on

Most of the unfairness people feel in performance management starts at goal setting, long before any rating.

Good goals have three properties. They connect visibly to the unit's plan, so employees can see why the goal matters. They are specific enough that the employee and manager would agree, independently, on whether they were met. And there are few of them. Five to seven objectives is usually plenty; twelve is a list of tasks.

A few practical design points:

  • Cascade with judgement. Mechanical cascades, where every level copies and slices the level above, produce goals that are technically aligned and practically meaningless. Ask each unit what it will contribute to the level above, then let managers translate that into individual goals.
  • Mix results and behaviours. Most systems assess both what was achieved and how. Keep the behavioural element tied to a short set of competencies or values with clear indicators, otherwise it becomes a place for vague impressions.
  • Allow change. Plans change mid-year, especially in organisations with shifting mandates. Build a simple mechanism to revise goals at the mid-year point, with manager approval, rather than assessing people against objectives that no longer exist.
  • Separate KPIs from objectives. A KPI tracks the ongoing health of a role (turnaround time, error rate). An objective describes a change or result for this year. Both are useful, but mixing them makes goal sheets long and unclear. OKRs can help teams set stretch outcomes, but I would not use OKR scores directly for individual ratings; they are designed to be ambitious, and scoring people on them punishes ambition.

Ratings, calibration and distributions

Ratings are where trust is won or lost.

Keep the scale simple

Choose a scale that managers can apply consistently. Each level needs a plain description of what it means in behaviour and results, not just labels such as "exceeds". Three to five levels is typical. The more you rely on ratings for pay, the more you need clear distinctions at the top.

Calibrate before you communicate

Calibration is a meeting where managers review proposed ratings together, compare evidence, and adjust so that the same standard applies across teams. It is the most effective fairness control I know. It exposes lenient and harsh raters, surfaces people who are invisible to senior leaders, and forces managers to justify ratings with evidence.

For calibration to work, managers must bring evidence rather than adjectives, HR must facilitate firmly, and the session must happen before any rating is shared with employees. Once a rating is communicated, changing it costs far more trust than calibrating it properly in the first place.

Be careful with forced distribution

Rigid forced curves tell managers that a fixed share of their team must be rated low regardless of actual performance. In small teams this is statistically meaningless, and it teaches employees that ratings are rationed rather than earned. A guided distribution, applied across a large population and used as a check during calibration rather than a quota, is usually a better balance.

If the rating has no visible consequence, people stop taking the process seriously. If the consequence is poorly explained, they stop trusting it.

Design the link explicitly. For pay, that usually means a merit or bonus matrix that shows how rating and position in the salary range translate into outcomes, with budgets agreed before calibration. For promotion, performance should be a necessary condition but not a sufficient one; readiness for the next role is a separate judgement. For development, the review should feed into an individual development plan that someone checks later in the year.

Make the rules public inside the organisation. Employees do not need to see everyone's outcome, but they should be able to understand how their own rating led to their own result.

A design checklist

Use this to review an existing system or to brief a design team. Each "no" is a likely source of mistrust.

Design area Question to answer Yes / No
Purpose Have we ranked the system's purposes, and does the design reflect that ranking?
Goals Are individual goals set within the first six weeks of the cycle, and are most of them measurable?
Goals Is there a formal mid-year mechanism to revise goals when plans change?
Behaviours Are competencies or values assessed against written indicators, not impressions?
Check-ins Is there a minimum cadence of documented manager conversations during the year?
Ratings Does each rating level have a plain definition that managers can apply?
Calibration Are ratings calibrated across managers before anyone is told their rating?
Consequences Is the link from rating to pay, promotion and development written down and explained?
Underperformance Is there a clear, fair process for performance improvement plans with defined support?
Managers Are managers trained and assessed on how well they manage performance?
Appeals Can employees raise a concern about their review through a defined route?
Review Do we analyse rating patterns by unit, grade, gender and nationality each year?

The last row matters in the GCC. With large expatriate populations and national workforce programmes such as Emiratisation and Saudisation, you should check that ratings are not systematically different by nationality or other groups without a performance reason. If they are, find out why before anyone else does.

Choosing software after the design, not before

Performance management software is useful. It removes spreadsheets, reminds managers of deadlines, stores evidence and makes analysis possible. But I regularly see organisations buy a platform and adopt its default workflow as their performance philosophy.

Agree the design first. Then evaluate tools on whether they support your goal structure, your rating and calibration process, your check-in cadence, Arabic and English use where relevant, integration with your HR system of record, and reporting that HR can actually use. Ask vendors to demonstrate your process, not their standard demo.

Also consider what should be automated beyond workflow. AI agents are increasingly able to draft goal suggestions from unit plans, summarise check-in notes before a review, and flag rating patterns for calibration. These can save managers real time, but a manager must still own the judgement and the conversation.

Where to start

  1. Write down, in ranked order, what your performance system is for, and get the executive team to agree to it.
  2. Run the checklist above against your current process and list every "no" with an owner.
  3. Pull last year's ratings and look at the distribution by unit and manager; the variation will tell you where calibration is most needed.
  4. Fix goal setting for the next cycle before changing anything else, because it is the foundation for every later step.
  5. Only then decide whether your current tool fits the design, or whether you need a different one.

If your system also needs to shift towards more frequent conversations, my article on continuous performance management without losing rigour covers how to do that without weakening accountability. At Humanyx, we design performance systems end to end, from purpose and policy to calibration and the tools that support them.

FAQ

Questions HR leaders ask

What are the main components of a performance management system?

A complete system covers goal setting, ongoing check-ins and feedback, a formal review with some form of assessment, calibration across managers, and a clear link to pay, promotion and development. Policies, manager capability and a supporting tool hold these together.

Should we use a 5-point or 3-point rating scale?

Fewer points are easier to apply consistently, and many organisations find three or four levels enough to separate genuine differences in performance. Choose based on how you will use the ratings; if they drive a differentiated bonus, you need enough levels to make the differentiation meaningful and defensible.

Is forced distribution legal and advisable in the GCC?

Organisations in the region do use guided distributions, but a rigid forced curve tends to damage trust, particularly in small teams where the curve makes little statistical sense. A guided distribution applied at a larger population level, combined with proper calibration, is usually a better balance. Check any approach against the employment rules that apply to your entity.

Want to apply this in your organisation?

Thirty minutes with Karim on what you're facing. No preparation needed.