SECTION
If AI Writes 80% of Your Performance Reviews, Whose Reviews Are They?
Listen to this article:
- In the first half of 2026, the major performance platforms converged on the same feature: AI that drafts a review from data the system already holds.
- The claim attached to the most confident version is that AI compiles roughly 80% of the first draft, with managers keeping full editorial control.
- Drafting was never the expensive part. Noticing, deciding, calibrating, and holding a genuine conversation are, and none of them got easier.
- Editorial control over a finished draft is a veto. Vetoing a document is a different task from forming a judgment.
- The useful distinction is when the AI runs. Before the manager concludes, it sharpens their thinking. After, it polishes a conclusion nobody reached.
- Reviews are the evidence base for pay decisions, and European pay transparency rules require objective criteria. A rationale nobody formed is not a rationale.
Performance platforms now offer to write the review for you. The claim attached to the most confident version of the feature is that AI can compile roughly 80% of the first draft, assembled from goals, feedback, 1:1 notes, and prior reviews. Managers keep full editorial control.
Read that as a product claim, and it sounds like a relief, right? Read it as a question about authorship, and it gets more interesting. If software gathered the evidence, chose the structure, and wrote the sentences, how did the manager contribute?
The answer might be a great deal. It might also be a signature. The difference matters to the person receiving the performance review, and it does not show up in a demo.
The feature the category converged on
Over the first half of 2026, the best-known performance platforms arrived at versions of the same idea within months of each other. It is worth describing plainly, because the pattern matters more than whose logo is on it.
At its core is review drafting. The system already holds objectives, feedback, check-in notes, and prior reviews, so it assembles them into a structured draft, sometimes reaching into chat and project tools for more evidence. The better implementations show their sources, so a manager can trace where a claim came from.
Two related patterns sit alongside it. An assistant that joins the 1:1 itself, summarizing the conversation by topic, capturing action items with owners, and offering coaching notes to both people. And a layer that coaches the manager through the harder conversations, including how to explain a pay decision to someone who does not like the answer.
Writing reviews is one of the least popular tasks in the performance cycle, so the convergence is a response to a real complaint. The open question is whether writing was the part worth fixing.
Drafting was never the expensive part
Ask a manager what makes a review hard, and the writing is rarely the answer. The hard parts sit earlier.
Noticing what someone did across 11 months, when the details that mattered happened in March, and nobody wrote them down, is the hard part. Deciding what is actually true about their work, which usually means choosing between two defensible readings, is another great example. Calibrating so that "exceeds expectations" means roughly the same thing in your team as in the team next door. Then holding a conversation the person believes, including the part they will not enjoy.
Writing it up comes after all of that, and it is the cheapest step. It is also the only one that looks like work from the outside, which may be why it got automated first.
A Mirro customer put the calibration problem better than we would have: "High Performance" should mean the same thing in Sales as it does in IT. No drafting tool moves that. Two managers with different standards now produce two fluent documents instead of two rough ones, and the disagreement underneath is harder to see.
What AI drafting genuinely fixes
None of that makes auto-drafting worthless, and pretending otherwise would be dishonest.
It fixes the blank page. A manager who has been staring at an empty form for forty minutes is not thinking deeply; they are avoiding starting. Having something on the screen breaks that.
It also helps with the manager who writes nothing. Every HR lead knows the ones who submit three sentences of warm praise for everyone, every cycle, and have done it for years. A draft grounded in real evidence beats that by any standard you choose.
Retrieval is the strongest case of the three. Finding what someone did across months of scattered notes, meeting docs, and half-remembered wins is real work with no judgment in it. Automating that is a straightforward good.
That last one carries the whole distinction.
Editorial control is a veto, not a judgment
The vendors have thought about this, and the guardrails are more careful than the headline numbers suggest.
The commitments run along the same lines across the category. The AI does not assign the rating. It does not submit anything on the manager's behalf. Every generated sentence can be edited or deleted; nothing reaches the employee without the manager's review, and the manager stays accountable for tone, accuracy, and outcome.
Take those at face value. They are real commitments, not cover.
They also describe a veto. A veto is something you exercise against a thing that already exists, which changes the manager's job from forming an opinion to auditing one. "Is anything in here wrong?" is a much easier question than "what do I actually think about this person's year?" It is also the question a finished draft invites.
Editing pulls you toward what you are editing. The draft becomes the starting point, and moving away from it costs effort you are free not to spend, particularly at 6 pm with eight reviews left to file. Nothing stops a manager from doing the harder thinking. The default simply stops requiring it.
Then the meeting happens. A manager who didn't think hard arrives holding a document they didn't write, defending judgments they didn't form. The employee asks one question past the surface, and the answer runs out. People notice that immediately. It is how you end up with a review that is complete, accurate, well structured, and changes nothing.
Two kinds of AI in a review cycle
Before the manager has concluded anything, AI can do a lot of good. It can surface evidence they forgot. It can flag that a rating sits oddly against the rest of their own team's ratings. It can catch language that is vague, unevidenced, or loaded in a way the writer did not intend, and ask what a claim rests on. All of that sharpens a judgment while leaving it to the manager.
After the manager has been handed a conclusion, the same capabilities do something narrower. They check the text the AI wrote against evidence the AI selected. That is quality control on a draft. It is useful, and it is not the same as pressure on a human's thinking.
These platforms already do checks like this. Several will flag tone mismatches, bias patterns, and claims that lack evidence, and they say so plainly in their own documentation. The difference is not what the software can do. It is what exists on the screen before the manager has decided anything.
It matters more when pay is attached
This would be a philosophical quibble if reviews stayed inside the performance cycle. They do not. They are the evidence base for who gets promoted and who gets paid more.
European pay transparency rules now require employers to justify pay differences using objective, gender-neutral criteria such as skills, effort, responsibility, and working conditions. That obligation lands on documentation that already exists, and performance reviews are most of it.
So the question stops being abstract. When an employee asks why a colleague earns more, or a regulator asks the same thing in writing, the answer has to trace back to a judgment somebody made and can still explain a year later. A fluent paragraph assembled from available AI signals is not that.
Mirro customers describe the gap directly. Salary decisions lack objective justification. Or, from a sales conversation, in the plainest possible form: why does employee X get that salary?
Where Mirro stands
Mirro has not shipped AI review drafting, and this is not a promise that it never will. It is a statement about where the line sits.
Automate retrieval. Automate consistency checks. Leave the judgment and the rating with the person accountable for them.
In practice, that makes the dull infrastructure matter more than a generator. Mirro is self-serviced, so admins build their own cycles, questions, and audiences without filing a request. The Question Library holds a stable set of question IDs per company, which is what makes an answer in Sales comparable to the same answer in IT, and comparability is the precondition for calibration. Configurable 9-box and 4-box grids give calibration somewhere to happen. Perspective selection sets who answers, and visibility controls set who sees the result, as configuration rather than convention.
The blank page has a less glamorous fix too. Continuous performance triggers check-ins from hire dates and on recurring cycles, and the Audience Builder targets them by role, tenure, team, or location. Evidence accumulates through the year instead of being reconstructed in one sitting in December. A manager who has had six short conversations does not need 80% of a draft.
Because performance, objectives, recognition, and engagement sit alongside salary ranges and pay equity analytics in the same system, a pay decision keeps the reasoning that justifies it. For a small team, that is the practical version of this whole argument: reviews your managers actually thought about, without adding a process nobody has time to run. Setup takes weeks, and each cycle saves hours of manual configuration.
Frequently Asked Questions
-
Should AI write performance reviews?
-
Is it wrong to use AI to draft a performance review?
-
Can employees tell if their review was written by AI?
-
Does AI make performance reviews less biased?
It shouldn't write the judgment, but it can safely do most of the work around it. Gathering evidence, checking a rating against your other ratings, and flagging vague or loaded language all improve a review. Producing the assessment itself moves the manager from deciding to editing, which is a smaller job than it looks.
It’s not wrong, but the order matters more than most teams realize. Using AI to find what you forgot is different from using it to reach a conclusion you then approve. The first improves your judgment. The second replaces the moment where the judgment would have formed.
Often, though not from the prose, which is usually fine. They notice in the conversation when a question one layer past the document goes unanswered. A review is judged on whether the person across the table can defend it.
It can help, but it can also hide bias. Flagging loaded language or a rating that doesn’t match your own pattern is real bias reduction. Writing confident, well-evidenced prose around an unexamined judgment makes the same bias harder to spot, because the output no longer looks careless.
See what performance, objectives, and pay data look like in one system, built so the judgment stays with your managers: try Mirro.