Strategy · 5 minute read
Digital FTE Performance Review: Evaluating an AI Agent Role
A Digital FTE performance review is a quarterly, evidence-based assessment of an AI agent role against its job description: quality on the golden dataset and sampled production work, exception and override rates, cycle time, cost per task, incidents, and drift. It ends in decisions: keep, adjust the specification, change the autonomy level, or retire the role.
A person who is never reviewed drifts from the role they were hired for. An AI agent does the same, faster and more quietly: inputs change, the specification ages, costs creep, and the owner stops looking because it "works." A Digital FTE performance review is the practice that prevents this. It treats the agent as accountable capacity, reviewed on evidence, with decisions that follow. This guide gives the metrics, the evidence, the decisions, and the cadence. It builds on what is a Digital FTE and the role definition in the Digital FTE job description template.
What is reviewed?
The review compares the role's actual performance with its job description across six metric families.
| Family | Metrics | Compared with |
|---|---|---|
| Quality | Golden-dataset accuracy by category; sampled production quality; zero-tolerance criteria | Quality bar in the job description; previous quarter |
| Exceptions and overrides | Exception rate and reasons; override rate at approval gates; escalation quality | Trend; oversight capacity |
| Cycle time | Time from intake to completion, including exceptions | Instrumented baseline; previous quarter |
| Cost | Cost per completed task by component; budget variance | Budget; the alternative |
| Incidents | Count, severity, root causes, time to resolve | Zero for prohibited actions |
| Drift | Input mix, score trends, retrieval and tool changes | Previous quarters |
The metrics come from the evaluation harness, the production sampling pipeline, the exception queue, the gateway's cost data, and the incident log. If any of those sources does not exist, that is the first finding of the review.
What evidence is presented?
- Evaluation report: current golden-dataset results by category against thresholds, and the trend.
- Production sample review: what reviewers found in sampled real work, with examples of misses.
- Exception analysis: the top exception reasons, which are specification gaps, data problems, or genuine judgment cases.
- Override analysis: where humans disagreed with the agent at approval gates, and who was right.
- Cost report: cost per task by component versus budget, from the gateway and oversight time tracking.
- Incident log: every incident with root cause and remediation status.
- Change log: specification, model, and tool changes since the last review, with their regression results.
The evidence discipline is described in the evaluation-driven development whitepaper.
How is the autonomy decision made?
The most consequential output of the review is the autonomy level: suggest, act with approval, or act with sampling.
| Move | Evidence required |
|---|---|
| Advance | Quality at or above the bar in every category for the review period; override rate low and stable; no prohibited-action incidents; oversight capacity confirmed for the new level |
| Hold | Any metric below the bar, or evidence too thin |
| Restrict | A significant incident, a quality regression, or drift the specification has not caught up with |
The owner makes the decision, records the evidence, and the change is implemented through the permission and approval configuration, not by editing a prompt. The oversight design behind the levels is described in human-in-the-loop AI explained.
What other decisions follow?
- Specification updates, driven by the exception and override analysis. Recurring exceptions of one type usually mean a rule is missing or unclear.
- Evaluation dataset additions, so every verified miss becomes a permanent regression test.
- Scope changes, adding case types the agent handles reliably or removing ones it does not.
- Budget re-forecast, from actual cost per task and volume, as described in how to budget for Digital FTEs.
- Registry update, so governance records reflect the current role, permissions, and autonomy level.
- Retirement, when the role does not earn its place against the alternative at the required quality. This is a normal outcome and should be taken cleanly: identity revoked, permissions removed, work returned to the previous path, evidence archived.
How does the review meeting run?
A useful format is one hour, quarterly, owned by the process owner:
- Evidence presented by engineering (fifteen minutes).
- Exception handlers describe what escalated and why (ten minutes).
- Finance presents cost per task versus budget and the alternative (five minutes).
- Discussion of misses, overrides, and incidents (fifteen minutes).
- Decisions recorded: autonomy, specification, scope, budget, registry (ten minutes).
- Actions assigned with owners and dates (five minutes).
The record goes to the governance body that oversees the fleet, described in the agentic AI governance whitepaper, and to the board-level reporting cycle where the role is material.
How do reviews change as a role matures?
New roles are reviewed monthly until the autonomy level is settled, because the evidence accumulates fast and the decisions matter. Mature roles move to quarterly with a monthly dashboard check. Roles with consequential permissions keep risk or audit in the room. Any incident, model migration, or major specification change triggers a review regardless of cadence.
What are the common mistakes?
- No review because it "works." Drift and cost creep go unnoticed.
- Reviewing activity such as tasks processed, rather than quality, exceptions, and cost per task.
- Autonomy decided by enthusiasm rather than evidence, in either direction.
- Exceptions not analyzed, so the specification never improves.
- Reluctance to retire, keeping a role that costs more than the work it absorbs.
- Reviews that do not update the registry or the budget, so governance and finance see stale data.
How does FISTA Solutions help?
FISTA Solutions installs the performance review as part of every AI agents engagement: the evaluation harness, the sampling pipeline, the exception analytics, and the cost telemetry that make the review evidence-based, with forward deployed engineers running the first reviews alongside the owner, and the AI enablement practice establishing the cadence and registry across the fleet. FISTA has delivered 150+ projects for 50+ companies across 12+ countries.
To set up reviews for your existing agents, message FISTA on WhatsApp, or read Digital FTE cost for the cost side of the review.
Share-ready article cover
Download the generated social format.
Clear answers
Questions raised by this field note.
Straightforward guidance for evaluating scope, fit, and the next step.
01How often should a Digital FTE be reviewed?
Quarterly as the standing cadence, with a lighter monthly check on quality, exceptions, cost, and incidents, and an immediate review after any significant incident or a model or specification change. New roles are often reviewed monthly for the first quarter while the autonomy level is being established.
02What metrics go into a Digital FTE review?
Quality on the golden dataset by category and on sampled production work, exception rate and override rate at approval gates, cycle time against the baseline, cost per completed task, incidents and their severity, and drift in input mix or scores. Each is compared with the previous quarter and the thresholds in the job description.
03Who conducts the review?
The role's human owner, usually the process owner, with engineering presenting the evidence, exception handlers contributing their experience of what escalated, and finance checking cost per task against budget. Risk or audit attends for roles with consequential permissions.
04What decisions come out of a review?
Whether to keep, adjust, advance, restrict, or retire the role; whether to change the autonomy level up or down based on evidence; which specification updates and evaluation-dataset additions are needed; and whether the budget and the governance registry should be updated. Each decision is recorded with its evidence.
Continue exploring
Related capabilities
Start with the hard problem
Need the outcome owned, not merely analyzed?
Tell us where delivery is constrained. We’ll map the fastest credible path from intent to verified production.