When AI can 10x anyone's output volume, output volume stops being a performance signal. The durable signals are judgment quality, goal attainment, and the story behind the numbers. If your review process captures none of those monthly, it captures nothing.
This piece is the methodology companion to The Verification Penalty. If that article named the problem — traditional metrics systematically underscore employees who verify, challenge, and correct AI output — this one answers the practical question that follows: what do we actually measure instead?
The answer is not to abandon quantitative measurement. It is to stop treating output volume as the primary signal and build a model that combines three things output volume alone cannot tell you.
The Three Signals That Survive the AI Era
Each of these signals captures something the others miss. Each alone is gameable or incomplete. Together they produce a defensible picture of contribution.
Qualitative narrative — the story behind the output
Monthly structured responses from both the employee (wins, challenges, Blockers raised, support requested) and the manager (recognition of specific contributions, primary coaching focus). This is the signal that captures judgment calls, error prevention, cross-functional contribution, and all the work that does not appear in any dashboard. It is also the signal most likely to reveal when a high-output employee's numbers are AI-assisted rather than judgment-driven.
Quantitative results — calibrated to role expectations
Measurable outcomes mapped to the specific benchmarks defined for the employee's job level and function. Not raw volume — revenue generated by a senior account executive is not the same signal as revenue generated by a first-quarter sales hire. Calibrated against what the role expects, at the level the employee is performing at. Quantitative results without this calibration layer are comparisons between incompatible things.
Goal attainment — what the employee and organization agreed on
Whether the specific goals defined at the start of the review period were achieved, and at what level. Goals capture both organizational priority and individual accountability — what the employee was supposed to do, not just what they happened to do. Goal attainment without the other two signals produces a score entirely dependent on how well goals were written.
Why Combining Them Beats Any Single Measure
The case for a composite score is not just that it captures more. It is that each signal disciplines the others.
High quantitative output with thin qualitative narrative and missed goals is a signal worth interrogating — not a guaranteed high score. Strong qualitative narrative with weak quantitative results warrants a different conversation than weak narrative with strong results. Goal attainment that happened despite significant documented Blockers reads differently than the same attainment in a frictionless quarter.
A Calibrated Performance Score built from all three creates a record where the components are visible alongside the composite — reviewers, HR leaders, and employees can see what drove the number, not just what the number is. That auditability is what makes the score defensible when it matters: in a compensation discussion, a promotion decision, or a workforce reduction.
| Signal | What it captures | What it misses alone |
|---|---|---|
| Qualitative narrative | Judgment, challenges cleared, contributions recognized | Unverifiable without quantitative grounding |
| Quantitative results | Measurable output against benchmarks | Inflatable by AI; silent on how results were achieved |
| Goal attainment | Alignment between individual contribution and org priority | Entirely dependent on goal quality |
| Calibrated composite | A defensible, auditable record across all three dimensions | Requires consistent monthly collection — cannot be reconstructed annually from memory |
The Fairness Mechanic: Why Silence Should Exclude, Not Zero
One of the most consequential design decisions in any performance scoring model is what happens when manager input is absent.
The standard approach — treating a missing manager response as a zero contribution for that period — introduces a fairness problem immediately. The employee has no control over whether their manager submits a check-in response. A manager who goes silent for three consecutive months should not reduce an employee's score through their absence. That outcome penalizes the employee for the manager's failure.
The alternative is exclusion. When a manager does not respond within the defined window, that period is excluded from the employee's score calculation rather than counted as a zero. The employee's Calibrated Performance Score reflects only the periods where full two-sided signal exists.
The distinction matters at scale. In any organization with Ghost Managers — managers who are technically present in the org chart but consistently absent from the management process — the exclusion mechanic protects employees from carrying a performance penalty they did not earn.
Mapping Rating Scales: Bringing Different Systems Onto the Same Spine
Many organizations have existing rating scales — typically 1–5 or 1–4 — that need to translate into a consistent scoring framework. Employees cannot be fairly compared across departments if their results are scored on different scales.
The mapping approach that works is range-anchoring, not averaging. Each point on a manager rating scale corresponds to a range on a 0–100 scoring spine:
| Manager rating | Score range | What it reflects |
|---|---|---|
| 5 — Exceptional | 90–100 | Demonstrably above role expectations; specific behavioral evidence required |
| 4 — Strong | 75–89 | Consistently meeting and selectively exceeding expectations |
| 3 — Solid | 55–74 | Meeting expectations; development opportunities identified |
| 2 — Developing | 35–54 | Partially meeting expectations; improvement plan warranted |
| 1 — Below expectations | 0–34 | Not meeting core role requirements; formal review required |
Range-anchoring means a manager who rates "5" is not assigning 100 — they are assigning somewhere in the 90–100 band, calibrated against the qualitative evidence provided and goal attainment in that period. No single input carries the whole result.
What This Requires: Monthly, Not Annual
The blended model described here requires one thing that annual review cycles structurally cannot provide: consistent, current data.
Qualitative signal reconstructed from memory in December for a January-to-December performance year is not the same signal as qualitative narrative collected in January, February, March, and every month thereafter. The verification work done in February — the assumption challenged, the error caught — will not appear in an annual narrative written eleven months later. The December review is a recency-weighted summary, not a full-year record.
Monthly structured check-ins are not an additional administrative burden on top of the annual review. They are the data collection mechanism that makes the annual review — or any performance decision — grounded in something other than end-of-year impressions.