What problem does it solve?
A formal review asks a manager to remember and evaluate six or twelve months of a person's work, usually while running the same exercise for every other person they manage at once. Windmill describes this playing out at Rho as the company grew: managers spent significant time gathering work artifacts and feedback across separate tools, reviews stretched longer than intended, and leaders lacked a unified view of performance across teams.
Windmill describes this at Case Status, where what Case Status called the "blank page problem", managers and employees staring at an empty review form and trying to reconstruct months of work from memory, played out against a prior process, a patchwork of Google Docs and forms, that was both time consuming and made it hard to capture the full scope of an employee's contributions. Windmill also argues that mid sized companies often see the largest gains from this kind of tool, because they typically lack the dedicated HR resources of large enterprises to run and chase a formal cycle by hand.
How does it work?
- Collect the year's context continuously. The assistant connects to the tools work already happens in, such as project trackers, chat and documents, and to goals set earlier in the cycle, so it has real material to draw from rather than starting from nothing at review time.
- Gather structured feedback. Peers, the manager and, where used, the employee's own self review answer a short set of questions; the assistant chases outstanding responses so the cycle does not stall on one missing input.
- Draft the review. The assistant synthesises the work history and feedback into a first draft organised against the organization's review template and competencies, citing the specific examples it drew from.
- The manager edits and owns it. The manager rewrites, adds judgement the data cannot capture, and is accountable for the final rating and text; the draft is a starting point, not the answer.
- Coordinate the cycle. The assistant tracks who still owes feedback, reminds them, and gives HR or the manager's own manager visibility into where the cycle stands, without a dedicated coordinator running it by hand.
- Audience
- Employee facing
- Autonomy
- Copilot
- Adoption
- Early adopters
- Channels
- Internal tools, Email
What is it worth?
Benchmarks are computed from the public deployments below: one data point per organization per KPI, with who made each claim.
| KPI | Median | Reported range | Data points | Claimed by |
|---|---|---|---|---|
| Cycle time reduction | Too few to pool | 84% | 1 | 1 vendor |
| Handling time reduction | Too few to pool | 83% | 1 | 1 vendor |
Value drivers: Employee productivity, Lower cost to serve, Speed and cycle time.
Indicative value
An organization with 2,000 employees who receive a formal performance review each cycle
USD 80,000 to USD 960,000
Manager and employee review time cost avoided per year
How this is calculated
Formula: employees * cyclesPerYear * hoursSavedPerReview * hourlyCost. The low scenario uses every low input, the high scenario every high input.
| Input | Low | High | Basis |
|---|---|---|---|
| Employees reviewed per cycle employees, employees | 2,000 | 2,000 | The reference organization. |
| Formal review cycles per year cyclesPerYear, cycles per year | 1 | 2 | Editorial assumption, replace with your own; many organizations run one annual and one mid year cycle. |
| Manager and employee hours saved per review hoursSavedPerReview, hours per review | 1 | 3 | Editorial assumption, replace with your own. Rho's CFO is quoted comparing the AI assisted draft to a review he estimated would otherwise have taken him 3 hours, which anchors the upper bound; the Case Status figure is treated here as a reduction in elapsed cycle time, not in hours worked, so it is not used to derive this range. |
| Blended fully loaded hourly cost of a manager or employee hourlyCost, USD per hour | 40 | 80 | Editorial assumption. Replace with your own blended rate. |
What it leaves out: Gross avoided time cost only. It leaves out the platform cost, the value of more consistent and timely reviews, and any effect on retention, development or promotion decisions, which this page's evidence does not measure.
Who already uses it?
2 public deployments, strongest evidence first. Grades: A regulator or audit, B the organization itself, C vendor case study, D anonymous or estimate.
Case Status
United States · Technology and software · 2025
Case Status, a client experience platform for law firms, used Windmill's AI review agent to solve what Windmill calls the "blank page problem": managers and employees starting a review with an empty form and having to reconstruct months of work from memory. Windmill surfaces work context from Slack and other connected tools and generates a summary that becomes the starting point for every review, replacing a prior process built on Google Docs and forms.
- Cycle time reduction: 84%, full review cycle, versus the prior Google Docs and forms process
"Case Status completed their entire performance review cycle in 84% less time compared to their previous process, while dramatically improving the experience."
Claimed by: vendor
Rho
United States · Technology and software · 2025
Rho, a fintech platform for startups, runs its full performance review cycle on Windmill's AI review agent. The People team replaced a process where managers spent significant time gathering work artifacts and feedback across separate tools with a cycle that drafts each review from collected feedback and work history: self reviews complete in about 2.5 days, 360 reviews in 5 days, and the full self to manager cycle in 8 days, against an industry average that Windmill describes as several weeks.
- Handling time reduction: 83%, total hours spent on reviews
"Results: 83% reduction in total hours spent on reviews and 93% of employees preferred Windmill to the prior process."
Claimed by: vendor
How do you implement it?
A model agnostic playbook: what to prepare, the order to build in, and what goes wrong.
Data you need
- The organization's review template, competencies and rating scale
- Goals set earlier in the cycle for each employee, where the organization uses them
- Access to the work context tools the assistant should draw from, scoped per employee
- A clear policy on what AI generated text may and may not be used for
Systems to integrate
- HR information system for employee, manager and cycle data
- Chat and project tools (for example Slack, Jira, GitHub) the assistant summarises from
- Document storage for prior reviews and goal documents
- Identity provider for single sign on and access scoping
Complexity: Medium
Drafting from feedback already collected in a form is straightforward. The value comes from connecting to where work actually happens (chat, project tools, documents) so the draft has real, specific material; without that, the assistant just reformats whatever a person typed into a box, which saves little time.
- 1
Decide what the draft may and may not touch
Write down, before launch, that the AI draft is a starting point only, the manager owns the final text and rating, and the system is not used directly in pay or promotion decisions.
- 2
Connect real work context, not just a form
Wire the assistant to the tools work happens in for a pilot group before rolling out broadly, so the draft cites specific, verifiable examples rather than generic language.
- 3
Keep feedback collection structured
Use a short, consistent set of questions for peer and self feedback so the assistant can synthesise reliably, and let it chase outstanding responses automatically.
- 4
Require a human edit before anything is final
Do not allow a draft to be submitted unedited; require the manager to open and modify the text, and log that the review was edited before submission.
- 5
Pilot with one team's cycle
Run a full cycle with one function first, measure cycle time, hours and employee preference, and fix template and integration gaps before the next team joins.
- 6
Decide how sensitive topics are handled
Performance issues that touch conduct, health or personal circumstances should route to a person and a private conversation, not into an AI drafted written record.
Guardrails
- The manager reviews and edits every draft; nothing is sent to the employee unedited
- The system is not used to set pay, bonus or promotion decisions directly from its output
- Feedback and drafts are visible only to the people the organization's existing review process would show them to
- Draft text is clearly marked as AI assisted until the manager has reviewed and approved it
- Conversation and draft logs have a defined retention period and restricted access
KPIs to instrument
- Total manager and employee hours spent on the cycle, before and after
- Cycle length from opening to closing the review period
- Employee preference for the AI assisted process versus the prior one
- Share of drafts materially rewritten by the manager, as a check that editing is real
- On time completion rate across the organization
Human in the loop
The manager owns every rating and every word of the final review. HR owns the template, the cycle policy and what counts as a sensitive topic that must go to a person instead of into a draft, and reviews a sample of cycles for consistency and fairness across teams.
Common failure modes
- Generic, uneditable prose
- A draft that reads fluently but says nothing specific gets rubber stamped rather than improved. Require citations to specific work items in the draft and sample reviews for genuine editing.
- The draft becomes the decision
- Under time pressure, managers submit the draft with only cosmetic changes, so the AI's synthesis quietly becomes the evaluation. Track the share of drafts materially edited and make manager training explicit about this risk.
- Feedback collected without consent context
- Peers do not realise their comments will be summarised and shown to the subject, which damages trust in the feedback process. Be explicit up front about what is collected and how it is used.
- Sensitive matters written into a permanent record
- A conduct or health related issue gets synthesised into formal review text instead of handled as a private conversation. Define these topics in advance and route them to a person, not the drafting flow.
What are the risks and rules?
EU AI Act
High risk
Annex III point 4(b) lists AI systems intended to monitor and evaluate the performance and behaviour of workers as high risk. Synthesising an employee's work history and feedback into a performance evaluation is very plausibly profiling of a natural person under GDPR Article 4(4), which expressly covers analysing or predicting a person's "performance at work". Article 6(3)'s last subparagraph makes an Annex III system high risk regardless of the derogations whenever it performs such profiling, so a tool built this way is high risk by default however much the manager edits the output. The derogations in Article 6(3), including a narrow procedural task or improving the result of a previously completed human activity, do not fit drafting an evaluation from scratch; the closest is point (d), a preparatory task ahead of a human assessment, which only has a chance of applying to a design that avoids profiling altogether, for example one that only surfaces raw facts without synthesising a judgement. Where that derogation is argued, the documentation duty under Article 6(4) falls on the provider of the system, and only on the deploying organization when it builds the tool itself. Because the tool is high risk by default, Article 26(7) requires informing affected workers and their representatives before it is put into use in the workplace, whatever the tool's output is used for; using the same system's output directly in pay, promotion or termination decisions removes any doubt and triggers the full high risk regime. Annex III's high risk obligations apply from 2 December 2027.
Guidance
- Annex III, high risk AI systems referred to in Article 6(2) (European Union, Europe). Point 4(b) lists monitoring and evaluating the performance and behaviour of workers as a high risk employment use.
- Article 6, classification rules for high risk AI systems (European Union, Europe). Sets out the narrow procedural task, prior human activity and preparatory task derogations in Article 6(3), the rule in its last subparagraph that profiling of natural persons always makes an Annex III system high risk regardless of those derogations, and the documentation duty in Article 6(4), which falls on the provider of the system.
- Article 26, obligations of deployers of high risk AI systems (European Union, Europe). Article 26(7) requires informing affected workers and their representatives before a high risk AI system is put into use in the workplace.
Controls to put in place
- Written policy that the AI draft is a starting point and the manager owns the final rating and text
- Documented Article 6(3) and 6(4) assessment from the system's provider, covering whether the design profiles workers
- Inventory entry for the assistant with an accountable HR owner
- Sampled review of edited versus unedited drafts, to check editing is real
- Defined list of sensitive topics that route to a person instead of the drafting flow
Frequently asked questions
- Does AI decide an employee's rating?
- The sources on this page do not describe the AI setting a rating, and neither Windmill customer story says who decides it. Design the system so the manager decides: require the manager to open, edit and approve every draft before it reaches the employee, and keep the rating a manager judgement that the tool never sets on its own.
- How much time does AI assisted drafting actually save?
- Windmill reports an 83% reduction in total hours spent on reviews at Rho and an 84% faster full cycle at Case Status, both alongside a reported 93% of employees preferring the new process to the prior one. Both figures cover the whole cycle (reminders, feedback collection, synthesis) rather than drafting on its own, and neither evidence record on this page is a drafting only deployment. Windmill's companies list page separately describes JPMorgan Chase using AI for drafting only, and cites a Boston Consulting Group figure of a 40% reduction in writing time when staff use AI to draft reviews; that is Windmill's own secondhand claim about a different deployment, not evidence recorded on this page.
- Is this high risk under the EU AI Act?
- By default, yes. Annex III point 4(b) lists monitoring and evaluating worker performance as high risk, and Article 6(3)'s last subparagraph makes an Annex III system high risk regardless of the derogations whenever it profiles natural persons; synthesising someone's work history and feedback into a performance evaluation is plausibly profiling under GDPR's definition. Only a design that avoids profiling and fits a narrow procedural or preparatory task can argue for the derogation, and even then the documentation duty under Article 6(4) falls on the provider of the system, not the deploying organization, unless the organization built it itself.
- What should stay out of an AI drafted review?
- Conduct issues, health matters and anything tied to a protected characteristic should be handled in a private conversation and only entered into the formal record by a person, not synthesised automatically from feedback text.
How to cite this page
Blits.ai AI Use Case Library, "AI agent for drafting employee performance reviews", last verified 28 September 2026, https://www.blits.ai/ai-use-cases/performance-review-drafting-agent. Licensed under CC BY 4.0. Method: how we verify use cases.
Changelog
- 28 September 2026: First published