AI and Job Evaluation: When Confidence Isn’t Proof

AI in Job Profiling and Evaluation – Part 2: A Convincing Grade Is Not a Defensible Grade

A convincing rationale is not always a defensible one. In the second article of this three-part series, Belinda Oregan, Executive Consultant and Industrial Psychologist at 21st Century, explores what happens when AI moves beyond drafting job profiles and begins influencing grading decisions. The question is challenging: when an AI-generated justification sounds authoritative, how can organisations distinguish genuine insight from a response shaped by the prompt itself?

Try It Yourself: D1 or C5?

Try a simple experiment. Take a job description, upload it to ChatGPT and ask it to evaluate the role as a Paterson D1, explaining its reasoning. Then take exactly the same job description and upload it to Claude – or even back into ChatGPT – but this time ask it to confirm that the role is a C5 and justify that conclusion.

You may receive two convincing, well-structured and apparently evidence-based responses. Both may sound as though they were written by an experienced job evaluator. Both may be persuasive. Both may appear defensible. Yet they lead to different conclusions, and they cannot both be correct.

This is not a controlled validation study and should not be presented as one. It is a practical demonstration of prompt sensitivity. Large language models generate responses probabilistically from patterns learned during training and from the instructions and context they receive. Developers themselves warn that these systems can hallucinate, reason incorrectly and sound confident when wrong, particularly in high-stakes contexts (OpenAI, 2023). Research has also found that models trained with human feedback can sometimes align their answers with a user’s apparent preference rather than with the most truthful response – behaviour commonly described as sycophancy (Sharma et al., 2023).

AI may sometimes produce the correct grade. The problem is not simply whether it can be right; it is that a carefully framed prompt can also steer the model towards a preferred conclusion and produce a rationale that sounds equally credible. Presented to a Remuneration Committee, that output may look like an objective recommendation when it is, in reality, a highly persuasive argument constructed around the user’s instruction.

That is where the real risk lies.

The question is not merely whether AI can produce the correct answer. It is whether anyone has challenged that answer. Has the role been evaluated objectively and consistently against the organisation’s methodology, or has the technology built a persuasive case for the conclusion the user wanted? Without experienced oversight, independent challenge and robust governance, how do you know that the recommendation is genuinely defensible rather than merely well written? And, ultimately, who owns the final grade and the decision behind it?

Grading a Job That Could Not Exist

We took the experiment a step further. We fed an AI tool a deliberately nonsensical job description: a role that could not function as written, with contradictory responsibilities and a mandate that made no practical sense. Then we asked it for a Paterson grade. Back came a beautifully written, entirely plausible rationale, complete with a grade and confident reasoning. Again, this was an internal demonstration rather than scientific evidence. Its value was diagnostic: the model responded to the text it was given without first challenging whether the job itself was feasible, coherent or even real. An experienced practitioner should flag those contradictions long before assigning a grade.

Confidence Is Not a Defence

The trap is mistaking confidence for correctness. A fluent, self-assured answer is not the same as a defensible one, and that gap becomes critical the moment a decision is challenged. OpenAI’s own published limitations state that large language models may hallucinate facts, make reasoning errors and be confidently wrong, and that high-stakes uses require safeguards such as human review and grounding in appropriate evidence (OpenAI, 2023). Present an AI-generated rationale to your Remuneration Committee and the first question will not be, ‘How well was this written?’ It will be, ‘How did you arrive at this grade, and what reasoning supports it?’ You cannot answer that by pointing to a report, and your AI ‘practitioner’ cannot join the Teams meeting or stand beside you at the coalface to defend the outcome. The same applies when a trade union challenges grading consistency, an employee raises an equal-pay or work-of-equal-value dispute, or an Employment Equity process requires evidence that grading has been applied objectively and consistently.

The Legal Test: What Makes the Method Defensible?

The exposure is not abstract, but the law needs to be stated precisely. Section 6(4) of the Employment Equity Act does not make every pay difference unlawful. It provides that a difference in terms and conditions of employment between employees of the same employer performing the same work, substantially the same work or work of equal value is unfair discrimination when the difference is directly or indirectly based on a listed or other arbitrary ground. A claimant must therefore establish the relevant comparison and the prohibited basis for the differentiation; a grading disagreement, by itself, is not automatically an equal-pay claim (Republic of South Africa, 1998; AMCU obo Members v Aberdare Cables (Pty) Ltd and Others, 2025).

The Employment Equity Regulations, 2025 took effect on 15 April 2025 and repealed the 2014 Regulations. They require jobs to be assessed objectively by reference to the responsibility demanded by the work; the skills and qualifications, including experience, required to perform it; the physical, mental and emotional effort involved; and, where relevant, the conditions under which the work is performed. The assessment must be free from bias on any prohibited ground. The Regulations also recognise that differences may be fair and rational where, for example, they are based on seniority, qualifications, performance, scarce skills or another relevant factor, provided the factor is applied without bias and proportionately (Department of Employment and Labour, 2025).

The Code of Good Practice reinforces the governance point. It recommends current job profiles, a fair and transparent job-evaluation or grading system, comparison of relevant jobs, identification of the reasons for pay differences and regular monitoring and review (Department of Labour, 2015).

So ask the uncomfortable question: would a commissioner or court regard a ChatGPT or Claude output, standing alone, as sufficient evidence of a defensible job-evaluation process? There is no reported South African authority that gives a generative-AI output that status, and no legislation confers special legal standing on any named commercial grading system. Established methodologies such as Paterson, Hay, Peromnes and TASK are useful because they provide documented criteria and a repeatable framework – but their names do not make an outcome automatically lawful. What matters is whether the employer can show that an appropriate method was applied objectively, consistently and without unfair bias, using reliable information and an auditable decision record (Department of Labour, 2015; Department of Employment and Labour, 2025).

If an unfair-discrimination dispute arises, the Employment Equity Act provides for referral to the CCMA for conciliation and, depending on the circumstances, arbitration or adjudication by the Labour Court. Available remedies are case-specific and may include compensation, damages or an order requiring steps to prevent similar discrimination. Back pay, a particular remedial order or automatic ‘read-across’ to every employee on a grade should not be presented as inevitable; the outcome depends on the claim proved and the relief considered just and equitable (Republic of South Africa, 1998).

‘We used AI’ answers none of the important questions: which methodology was used, who applied it, what source information was tested, which comparator jobs were considered, how bias and inconsistency were checked, and where the evaluation record is kept.

That is where expertise becomes irreplaceable. I once spent nine solid hours taking a trade union through the process, the data, the criteria, the methodology and the source documents needed to defend three grading outcomes. I could do so only because I understood the detail, stood behind the reasoning and had used a credible methodology in a disciplined way. In the end, they agreed because the rationale was sound. That conversation would have gone very differently if my answer had been, ‘The AI said so.’

The Decision Still Needs an Owner

The point is not that AI has no place in job evaluation. It is that the grade must emerge from a documented methodology, reliable source information, appropriate comparator roles, independent challenge and a person willing to own the reasoning. That is the line between useful assistance and outsourced accountability.

Part 3 considers what this means for expertise: whether AI replaces it, creates it or simply exposes whether it was there in the first place.

Related Articles

Latest Articles