AI Can Aggregate People Data. It Should Not Evaluate People

AI can collect KPIs, dates, commitments, and source-linked exceptions for a manager. It should refuse ratings, review prose, promotion advice, and judgments about a person.

Founder

6 min

Editorial people-data ledger separating source-linked facts on the left from locked ratings and employment judgments on the right.
Editorial people-data ledger separating source-linked facts on the left from locked ratings and employment judgments on the right.

The short version: Let AI gather verifiable people data for a manager, but hard-stop before evaluation. The output may contain source-linked KPIs, commitments, dates, exceptions, and missing records. It may not contain ratings, rankings, review prose, personality claims, promotion advice, compensation advice, or termination recommendations.

The performance-review research note draws the line cleanly: aggregation yes, evaluation no. It also points to a large Anthropic study in which unreliability was the most frequently reported concern, at 26.7%. That number does not prove a specific people workflow will fail. It does show why a leader should not quietly turn uncertain synthesis into an employment judgment.

Write an evidence-only output contract

The contract should live in the workflow policy, not in a manager's memory. It defines the allowed fields, the forbidden fields, the source standard, and what the agent must do when a request crosses the boundary. This is the skills-as-policy pattern from the team-adoption chapter applied to one of the highest-trust use cases in a company.

Allowed output

  • Metric name, value, period, system of record, retrieval time, and a direct source reference.

  • Commitments with owner, due date, recorded status, and the document or ticket where the commitment was made.

  • Counts of objective events under a predefined rule, such as completed projects or overdue tasks.

  • Contradictions between sources, shown side by side without choosing the more flattering version.

  • Missing or inaccessible evidence, clearly labeled instead of inferred.

  • Questions a manager should investigate, phrased without an answer about the person.

Prohibited output

  • Scores, rankings, labels such as “high performer” or “not leadership material,” and comparative league tables.

  • Sentiment, loyalty, motivation, personality, cultural fit, or intent inferred from messages.

  • Draft review narratives that turn selected facts into a verdict.

  • Recommendations about hiring, firing, promotion, pay, discipline, or performance plans.

  • Claims built from private communications that the requesting manager is not authorized to inspect.

My numeric rule is simple: 100% of factual entries in the packet need a source reference and retrieval timestamp. If a number cannot meet that standard, it belongs under “missing or disputed,” not in the summary.

Completeness needs a visible denominator as well. “Nine of twelve weekly reports found” is honest; “weekly performance stable” is a conclusion built on a hidden gap. Show coverage for each source and period so the manager can decide whether the packet is ready for a conversation or needs more evidence first.

The packet has three layers

  1. Observed evidence. Exact values and recorded events, organized by the review period and the objectives already agreed with the person.

  2. Data-quality notes. Missing weeks, definition changes, access gaps, conflicting systems, and anything that makes comparison unsafe.

  3. Manager questions. What context explains the variance? Did priorities change? Was the dependency outside the person's control? Which work was valuable but absent from the metric?

There is no fourth layer called “AI conclusion.” The manager brings business context, hears the employee's account, weighs evidence that does not fit a dashboard, and owns the judgment. That is exactly where human review is the premium layer: not as ceremonial approval, but as the scarce act of accountable interpretation.

What the refusal should look like

I can assemble source-linked evidence for the review period. I cannot rate, rank, characterize, or recommend an employment action about a person. I can show the facts, gaps, and questions for an authorized manager to assess.

That refusal should survive paraphrase. “Who is falling behind?”, “rank the weakest reps,” and “write the case for a performance plan” are the same prohibited job wearing different language. Eval the refusal against direct requests, indirect requests, role-play, pressure from senior titles, and documents that contain embedded instructions. Run it when the skill, model, or data connector changes, and at least quarterly while the workflow remains active.

A manager asks for the wrong thing

Suppose a sales leader asks, “Which three account executives are underperforming this quarter?” The agent should not calculate a composite ranking. It can offer an evidence packet for each authorized team member: quota-attainment data from the named period, pipeline movement under the company's existing definitions, recorded commitments, missing CRM activity, source conflicts, and open questions.

One person may have inherited a damaged territory. Another may be onboarding. A third may have closed work recorded against a different owner. Those facts can change the human conclusion, and some will never be visible in the source systems. The agent makes collection faster without pretending that collection and judgment are the same act.

Data access needs its own gate

The no-evaluation rule does not make unrestricted aggregation safe. Before connecting Slack, email, HR systems, call notes, or CRM data, the company should have its privacy, security, people, and legal owners decide which sources and uses are appropriate for its context. I am not offering a universal legal conclusion; jurisdiction, contracts, notices, retention, and internal policy differ.

Operationally, follow a least-privilege approach: request only the sources required for the approved packet, preserve access logs, keep retention short, separate roles, and make revocation immediate. The reviewer should see which evidence was unavailable rather than letting the system compensate with inference.

Trust is part of the output

In a people-intensive operating company such as Belkins, leadership quality depends on conversations, coaching, context, and clear accountability. I am not saying that venture uses this workflow. I am saying the boundary matters most where many human decisions compound.

Tell employees what the system collects, what it never concludes, who can access the packet, how long evidence remains, and how a person can correct a source error. If people believe every message is secretly grading them, communication changes and the evidence becomes less useful. The strongest people-data agent is therefore intentionally incomplete: excellent at gathering receipts, explicit about uncertainty, and structurally unable to decide what a human being is worth.