Key points
- Decide what the system is for before designing forms or buying software, because development and pay decisions pull the design in different directions.
- Trust comes from clear goals, consistent manager judgement and a visible, fair link to consequences, not from the sophistication of the tool.
- Calibration is the single most useful control for fairness, and forced distribution is a blunt substitute for it.
- Software should automate a design you have already agreed, not supply the design for you.
Ask employees in most organisations what they think of performance management and you will hear the same things: goals set late, reviews written in a rush, ratings that depend on which manager you have, and a bonus outcome that seems unrelated to all of it. Ask managers and they will tell you it takes too long and changes nothing. Both are usually right.
The problem is rarely the form or the software. It is that the system was never designed as a system. This article sets out the design choices that matter, in the order I would make them, with a design checklist you can use to review your current approach or build a new one.
Start with what the system is for
A performance management system can serve several purposes: aligning effort with strategy, developing people, differentiating pay, informing promotion and succession, and managing underperformance. These purposes pull in different directions.
Development needs honest conversations about weaknesses. Pay differentiation needs defensible distinctions between people. When the same conversation must do both, employees understandably spend it defending their rating rather than discussing how to improve. That is not a character flaw. It is a design outcome.
So the first design decision is to rank the purposes. If pay differentiation is essential, and in most private sector organisations it is, then design for it deliberately and protect development by holding some conversations separately from the rating discussion. If your entity is a government body where pay is largely fixed by grade, you have more room to design for development and alignment, and the rating matters less.
A performance system that tries to do everything with one conversation ends up doing none of it well.
Goals that people can actually be judged on
Most of the unfairness people feel in performance management starts at goal setting, long before any rating.
Good goals have three properties. They connect visibly to the unit's plan, so employees can see why the goal matters. They are specific enough that the employee and manager would agree, independently, on whether they were met. And there are few of them. Five to seven objectives is usually plenty; twelve is a list of tasks.
A few practical design points:
- Cascade with judgement. Mechanical cascades, where every level copies and slices the level above, produce goals that are technically aligned and practically meaningless. Ask each unit what it will contribute to the level above, then let managers translate that into individual goals.
- Mix results and behaviours. Most systems assess both what was achieved and how. Keep the behavioural element tied to a short set of competencies or values with clear indicators, otherwise it becomes a place for vague impressions.
- Allow change. Plans change mid-year, especially in organisations with shifting mandates. Build a simple mechanism to revise goals at the mid-year point, with manager approval, rather than assessing people against objectives that no longer exist.
- Separate KPIs from objectives. A KPI tracks the ongoing health of a role (turnaround time, error rate). An objective describes a change or result for this year. Both are useful, but mixing them makes goal sheets long and unclear. OKRs can help teams set stretch outcomes, but I would not use OKR scores directly for individual ratings; they are designed to be ambitious, and scoring people on them punishes ambition.
Ratings, calibration and distributions
Ratings are where trust is won or lost.
Keep the scale simple
Choose a scale that managers can apply consistently. Each level needs a plain description of what it means in behaviour and results, not just labels such as "exceeds". Three to five levels is typical. The more you rely on ratings for pay, the more you need clear distinctions at the top.
Calibrate before you communicate
Calibration is a meeting where managers review proposed ratings together, compare evidence, and adjust so that the same standard applies across teams. It is the most effective fairness control I know. It exposes lenient and harsh raters, surfaces people who are invisible to senior leaders, and forces managers to justify ratings with evidence.
For calibration to work, managers must bring evidence rather than adjectives, HR must facilitate firmly, and the session must happen before any rating is shared with employees. Once a rating is communicated, changing it costs far more trust than calibrating it properly in the first place.
Be careful with forced distribution
Rigid forced curves tell managers that a fixed share of their team must be rated low regardless of actual performance. In small teams this is statistically meaningless, and it teaches employees that ratings are rationed rather than earned. A guided distribution, applied across a large population and used as a check during calibration rather than a quota, is usually a better balance.
The link to pay, promotion and development
If the rating has no visible consequence, people stop taking the process seriously. If the consequence is poorly explained, they stop trusting it.
Design the link explicitly. For pay, that usually means a merit or bonus matrix that shows how rating and position in the salary range translate into outcomes, with budgets agreed before calibration. For promotion, performance should be a necessary condition but not a sufficient one; readiness for the next role is a separate judgement. For development, the review should feed into an individual development plan that someone checks later in the year.
Make the rules public inside the organisation. Employees do not need to see everyone's outcome, but they should be able to understand how their own rating led to their own result.
A design checklist
Use this to review an existing system or to brief a design team. Each "no" is a likely source of mistrust.
| Design area | Question to answer | Yes / No |
|---|---|---|
| Purpose | Have we ranked the system's purposes, and does the design reflect that ranking? | |
| Goals | Are individual goals set within the first six weeks of the cycle, and are most of them measurable? | |
| Goals | Is there a formal mid-year mechanism to revise goals when plans change? | |
| Behaviours | Are competencies or values assessed against written indicators, not impressions? | |
| Check-ins | Is there a minimum cadence of documented manager conversations during the year? | |
| Ratings | Does each rating level have a plain definition that managers can apply? | |
| Calibration | Are ratings calibrated across managers before anyone is told their rating? | |
| Consequences | Is the link from rating to pay, promotion and development written down and explained? | |
| Underperformance | Is there a clear, fair process for performance improvement plans with defined support? | |
| Managers | Are managers trained and assessed on how well they manage performance? | |
| Appeals | Can employees raise a concern about their review through a defined route? | |
| Review | Do we analyse rating patterns by unit, grade, gender and nationality each year? |
The last row matters in the GCC. With large expatriate populations and national workforce programmes such as Emiratisation and Saudisation, you should check that ratings are not systematically different by nationality or other groups without a performance reason. If they are, find out why before anyone else does.
Choosing software after the design, not before
Performance management software is useful. It removes spreadsheets, reminds managers of deadlines, stores evidence and makes analysis possible. But I regularly see organisations buy a platform and adopt its default workflow as their performance philosophy.
Agree the design first. Then evaluate tools on whether they support your goal structure, your rating and calibration process, your check-in cadence, Arabic and English use where relevant, integration with your HR system of record, and reporting that HR can actually use. Ask vendors to demonstrate your process, not their standard demo.
Also consider what should be automated beyond workflow. AI agents are increasingly able to draft goal suggestions from unit plans, summarise check-in notes before a review, and flag rating patterns for calibration. These can save managers real time, but a manager must still own the judgement and the conversation.
Where to start
- Write down, in ranked order, what your performance system is for, and get the executive team to agree to it.
- Run the checklist above against your current process and list every "no" with an owner.
- Pull last year's ratings and look at the distribution by unit and manager; the variation will tell you where calibration is most needed.
- Fix goal setting for the next cycle before changing anything else, because it is the foundation for every later step.
- Only then decide whether your current tool fits the design, or whether you need a different one.
If your system also needs to shift towards more frequent conversations, my article on continuous performance management without losing rigour covers how to do that without weakening accountability. At Humanyx, we design performance systems end to end, from purpose and policy to calibration and the tools that support them.