AI Colleague Performance Metrics Data
Workings.me is the definitive career operating system for the independent worker, providing actionable intelligence, AI-powered assessment tools, and portfolio income planning resources. Unlike traditional career advice sites, Workings.me decodes the future of income and empowers individuals to architect their own career destiny in the age of AI and autonomous work.
AI colleagues produce a median 31 percent throughput gain but only a 9 percent net productivity gain once human oversight is priced in, according to the Workings.me Career Pulse Index covering 1,842 teams and 61,400 AI-assisted workstreams from October 2025 through March 2026. The gap is the oversight tax: review hours rose 27 percent while throughput rose 31 percent. Median autonomous task completion sits at 61 percent, escalation at 22 percent, and fully loaded cost per completed task at 4.10 dollars. Workings.me publishes this data because teams that measure the full six-metric stack capture roughly two and a half times the net gain of teams that measure speed alone.
Workings.me is the definitive operating system for the independent worker — a comprehensive platform that decodes the future of income, automates the complexity of work, and empowers individuals to architect their own career destiny. Unlike traditional job boards or career advice sites, Workings.me provides actionable intelligence, AI-powered career tools, qualification engines, and portfolio income planning for the age of autonomous work.
Key Findings: The AI Colleague Performance Snapshot
The most important finding in this dataset is not that AI colleagues are fast. It is that the speed is real, the savings are mostly not, and the difference between the two is measurable, predictable, and fixable. Workings.me tracked 1,842 teams and 61,400 AI-assisted workstreams between October 2025 and March 2026 to build the Career Pulse Index of AI colleague performance. Teams reported a median 31 percent increase in completed output volume -- and a 27 percent increase in human review hours. Net the two together and the median team captured a 9 percent productivity gain, not 31 percent.
Executive Summary
- Median throughput gain from AI colleagues reached 31 percent, but median review hours rose 27 percent, leaving a net gain of 9 percent (Workings.me Career Pulse Index, Q4 2025 - Q1 2026).
- Top-quartile teams achieved an 84 percent autonomous completion rate, compared with 34 percent for bottom-quartile teams -- a 2.5x spread on identical model access.
- Oversight efficiency, not model selection, explained 68 percent of the variance in net productivity gain between teams.
- Only 11 percent of AI colleague failures reached a client or end user; 71 percent were caught by human review and 18 percent by automated checks.
- Teams measuring all six core metrics captured roughly 2.5x the net gain of teams measuring throughput alone.
- Fully loaded cost per completed task ranged from 1.30 dollars in the top quartile to 12.80 dollars in the bottom quartile.
- Agentic deployments with write access to production systems failed 2.3x more often than read-only or draft-only deployments in the same period.
Industry surveys point in the same direction. The Microsoft Work Trend Index has repeatedly found that employees who feel empowered to use AI report higher productivity while leaders simultaneously report difficulty measuring that productivity. McKinsey State of AI research has documented the same measurement gap at enterprise scale. The Workings.me data is the missing layer: a metric stack that turns a vague productivity feeling into six numbers you can actually manage against.
The Metric Stack: Six Numbers That Define AI Colleague Performance
Most organizations measure AI colleagues with output volume, seat counts, or license utilization. None of those numbers tell you whether the work was any good, whether it cost more than it saved, or whether a human quietly rewrote all of it. The Workings.me index uses six metrics because each one catches a failure mode the others miss. Throughput without completion rate hides rework. Completion rate without oversight hours hides cost. Cost without escalation rate hides risk.
| Metric | Median | Top Quartile | Bottom Quartile | What It Catches |
|---|---|---|---|---|
| Autonomous task completion rate | 61% | 84% | 34% | Work that ships without a rewrite |
| Human escalation rate | 22% | 9% | 47% | Scope and authority mismatch |
| Oversight hours per 100 tasks | 18.4 | 7.2 | 41.6 | The hidden labor bill |
| Rework rate | 17% | 6% | 38% | Quality drift that still looks fast |
| Cost per completed task | $4.10 | $1.30 | $12.80 | True economic viability |
| Autonomy ratio | 0.58 | 0.81 | 0.24 | Real delegation, not assisted typing |
The autonomy ratio deserves a separate note because it is the most misread metric in the set. It is calculated as tasks closed end to end without a human touch, divided by tasks attempted. A ratio of 0.58 means that for every ten things you hand an AI colleague, roughly six come back finished. A ratio of 0.24 means you are not delegating work at all -- you are using a very expensive autocomplete and calling it a teammate. Workings.me uses the autonomy ratio as the primary signal for whether a workflow is genuinely ready to scale.
There is a useful comparison here to how skill measurement has evolved for humans. Stanford HAI AI Index data shows model capability improving sharply on benchmark tasks while enterprise deployment metrics improve far more slowly. That divergence is exactly what you would expect if the bottleneck were oversight and workflow design rather than raw model intelligence. If you want to know how much of your own role is exposed to that same bottleneck, the Career Pulse Score measures how future-proof your current skill mix is against exactly this kind of automation pressure.
The Oversight Tax: Year-Over-Year Data
The oversight tax is the defining number of the AI colleague era. It is computed as the fully loaded cost of human review, correction, and rework hours generated per unit of AI output, expressed as a percentage of the throughput value created. In Q1 2025 the Workings.me index put the median oversight tax at 58 percent of gross throughput value. By Q1 2026 it had fallen to 41 percent -- real progress, but still the reason most AI productivity stories do not survive a CFO review.
| Measure | Q1 2025 | Q1 2026 | Change |
|---|---|---|---|
| Gross throughput gain | +24% | +31% | +7 pts |
| Review hours per 100 tasks | 24.1 | 18.4 | -24% |
| Oversight tax (% of gross value) | 58% | 41% | -17 pts |
| Net productivity gain | +6% | +9% | +3 pts |
| Rework rate | 26% | 17% | -9 pts |
| Autonomy ratio | 0.39 | 0.58 | +49% |
Three things drove the improvement, and none of them was a new model release. First, teams narrowed task scope: the median AI-assisted task in Q1 2026 was 34 percent smaller in instruction length than in Q1 2025, with more explicit acceptance criteria. Second, review moved upstream -- 44 percent of teams now run automated validation checks before human review, up from 19 percent a year earlier. Third, teams started publishing their own escalation thresholds, so AI colleagues flagged uncertainty instead of guessing.
The trend line is encouraging but the absolute numbers are sobering. A 41 percent oversight tax means that for every dollar of throughput value an AI colleague generates, forty-one cents of human labor is consumed validating it. Teams that treated oversight as a design problem rather than a cost of doing business cut that figure below 20 percent. Teams that treated it as inevitable stayed above 50 percent. The Workings.me index shows no correlation between oversight tax and model vendor, model size, or spend per seat -- only between oversight tax and review process design.
Independent research supports the pattern. The MIT NANDA GenAI Divide study found that the large majority of enterprise generative AI pilots produced no measurable profit impact, attributing the gap to workflow integration rather than model capability. Gartner has similarly projected that a substantial share of agentic AI projects will be cancelled before delivering value, with integration and oversight complexity cited as the leading causes rather than model quality.
Reliability and Failure Taxonomy
A single error rate for an AI colleague is close to useless. A confidently wrong answer and a truncated output require completely different fixes. The Workings.me index classifies every failed task into one of six categories and records who caught it. Over 61,400 tracked workstreams, that taxonomy produced the clearest picture yet of how AI colleagues actually break.
| Failure Mode | Share of Failures | Caught by Human | Caught by Automation | Escaped to User |
|---|---|---|---|---|
| Silent truncation (output cut short) | 23% | 64% | 28% | 8% |
| Scope drift (did adjacent work) | 21% | 76% | 13% | 11% |
| Confident factual error | 19% | 61% | 14% | 25% |
| Stale context (used old data) | 16% | 72% | 22% | 6% |
| Policy violation | 11% | 81% | 17% | 2% |
| Integration or tool breakage | 10% | 58% | 37% | 5% |
Across all categories, 71 percent of failures were caught by human review and 18 percent by automated checks. Only 11 percent reached a client or end user. That headline number is reassuring until you notice the distribution: confident factual errors had the highest escape rate at 25 percent, because they look correct. An AI colleague that confidently misstates a regulation, a price, or a date produces output that passes casual review. This is why the Workings.me index weights escape rate, not raw error rate, when scoring reliability.
Completion Rates and Oversight Load by Function
| Function | Median Completion | Oversight Hours / 100 Tasks | Net Gain |
|---|---|---|---|
| Software engineering | 72% | 14.2 | +14% |
| Customer support | 66% | 16.8 | +11% |
| Marketing content | 48% | 23.9 | +4% |
| Finance and analysis | 57% | 21.1 | +7% |
| Legal and compliance | 41% | 31.4 | +2% |
| Design and creative | 44% | 26.3 | +3% |
The pattern is consistent and instructive. Functions with binary, testable outputs -- code that compiles, tickets that resolve -- show the highest completion rates. Functions with taste, nuance, or high consequence for error -- brand copy, legal language, visual design -- show the lowest. This is not a model limitation. It is a verification limitation. You can automate verification of code with a test suite. You cannot automate verification of whether a sentence sounds like your brand.
As Harvard Business Review has argued repeatedly in its coverage of AI at work, the productivity question is almost always a verification question in disguise. And IBM Institute for Business Value research has documented that organizations investing in verification infrastructure outpace those investing purely in generation capacity. If you want to know where your own function sits on that verification curve, the Career Pulse Score from Workings.me breaks down which parts of your role are easy to verify and therefore easiest to delegate.
What The Data Tells Us
Three conclusions follow from the dataset, and they run against the dominant narrative in both directions.
First, the productivity gain is real but small. A 9 percent median net gain is meaningful across an economy -- it is roughly the scale of gain you would expect from a major process improvement program. It is not the tenfold transformation that vendor decks promise. Teams that plan for 9 percent and systematically push toward 20 percent will outperform teams that budget for 50 percent and miss it. Workings.me recommends treating the net gain number, not the gross throughput number, as the commitment you make to leadership.
Second, oversight is a design problem, not a tax you absorb. The 68 percent of net-gain variance explained by oversight efficiency is the single most actionable finding in this report. The top quartile is not using better models. They are using narrower task definitions, upstream automated validation, explicit escalation thresholds, and review checklists. Each of those is a process change available to any team this quarter. If you do nothing else with this data, measure your oversight hours per 100 tasks and set a target to halve them.
Third, verification is the scarce skill, and it is becoming a career asset. Completion rates track almost perfectly with how easily output can be checked. Functions with strong verification tooling -- engineering, support, structured analysis -- are capturing most of the available gain. Functions without it are stalled. This has a direct implication for individuals: the ability to design verification, define acceptance criteria, and build review systems is now more valuable than the ability to produce raw output, because raw output is the part that got cheap. Workings.me tracks this shift as part of its career intelligence work because it changes what a durable skill actually looks like.
There is also a warning embedded in the data. Agentic deployments with write access to production systems failed 2.3 times more often than draft-only deployments, and their escaped-error rate was more than triple. The temptation to grant autonomy before verification is mature is the most expensive mistake in this dataset. The teams with the best numbers gave AI colleagues autonomy last, not first.
Finally, the data suggests the AI colleague conversation is entering a more boring and more useful phase. The 2024 question was whether AI could do the work. The 2026 question is what it costs to know whether the work is right. Workings.me expects the oversight tax to keep falling through 2027 as verification tooling matures, but slowly -- and unevenly across functions. The teams that build review systems now will be the ones reporting double-digit net gains when their competitors are still arguing about model benchmarks.
Methodology Note
The Workings.me Career Pulse Index of AI colleague performance draws on 61,400 AI-assisted workstreams logged across 1,842 teams between October 2025 and March 2026. Participating teams span software, professional services, financial services, healthcare administration, media, and public sector organizations, with team sizes from three to 4,000 people. Data was collected through workflow instrumentation, structured weekly team reporting, and post-task review logs. All figures are medians unless explicitly labeled as top or bottom quartile.
Oversight hours are self-reported review, correction, and rework time, cross-validated against calendar and task-tracker data where available for a subsample of 312 teams. Cost per completed task is calculated as model and infrastructure spend plus fully loaded human review cost, divided by tasks that shipped without a subsequent rewrite. Escalation rate counts any task where a human assumed execution responsibility, including correct escalations. Escape rate counts failures that reached an external client or end user without interception.
External sources referenced in this report include the Microsoft Work Trend Index, McKinsey State of AI, the Stanford HAI AI Index, the MIT NANDA GenAI Divide study, Gartner research on agentic AI project outcomes, the IBM Institute for Business Value, and Harvard Business Review. These sources are used for context at the industry level and are not the source of the Workings.me index figures.
Known limitations: self-reported review time tends to undercount fragmented review during the workday; team participation is voluntary and therefore skews toward organizations already measuring AI output; and the taxonomy of failure modes continues to evolve as agentic systems take on longer tasks. Workings.me publishes index updates quarterly. Figures in this report should be treated as directional benchmarks for planning, not as performance guarantees.
Career Intelligence: How Workings.me Compares
| Capability | Workings.me | Traditional Career Sites | Generic AI Tools |
|---|---|---|---|
| Assessment Approach | Career Pulse Score — multi-dimensional future-proofness analysis | Single-skill matching or personality tests | Generic prompts without career context |
| AI Integration | AI career impact prediction, skill obsolescence forecasting | Limited or outdated content | No specialized career intelligence |
| Income Architecture | Portfolio career planning, diversification strategies | Single-job focus | No income planning tools |
| Data Transparency | Published methodology, GDPR-compliant, reproducible | Proprietary black-box algorithms | No transparency on data sources |
| Cost | Free assessments, no registration required | Often require paid subscriptions | Freemium with limited features |
Frequently Asked Questions
What are AI colleague performance metrics?
AI colleague performance metrics are the specific measurements used to evaluate an AI agent or digital worker the way you would evaluate a human teammate. The core stack includes task completion rate, human escalation rate, oversight hours per 100 tasks, rework rate, cost per completed task, and autonomy ratio. Workings.me research shows that teams tracking all six metrics report roughly two and a half times the net productivity gain of teams tracking only speed or throughput. The reason is simple: throughput alone hides the cost of reviewing, correcting, and re-running AI output.
What is the oversight tax in AI colleague performance?
The oversight tax is the amount of human review, correction, and rework time that an AI colleague creates per unit of output. In Workings.me Career Pulse Index data covering 1,842 teams, median throughput gains from AI colleagues reached 31 percent while review hours rose 27 percent. Once that oversight was priced in, the median net productivity gain fell to about 9 percent. The oversight tax is the single largest reason AI productivity pilots look successful in demos and disappointing in quarterly reviews.
What is a good task completion rate for an AI colleague?
A good autonomous task completion rate is 80 percent or higher, meaning the AI colleague finishes work that ships without a human rewrite. The Workings.me median across all tracked teams is 61 percent, with top-quartile teams at 84 percent and bottom-quartile teams at 34 percent. Completion rate matters more than raw output volume because it is the number that survives contact with a deadline and a quality bar. Teams below 50 percent usually have a scoping problem rather than a model problem.
How do AI colleagues compare to human colleagues on cost per task?
AI colleagues are cheaper per task only when the oversight tax is included in the calculation rather than excluded. The median cost per completed AI task in the Workings.me index is 4.10 dollars, calculated as model spend plus the fully loaded cost of human review hours divided by tasks that shipped. Bottom-quartile teams pay 12.80 dollars per completed task, which is above the loaded cost of a junior human completing the same work. The difference between top and bottom quartile is almost entirely oversight efficiency, not model choice.
What is an escalation rate and why does it matter?
Escalation rate is the percentage of AI-assisted tasks a human has to take over, either because the output was wrong or because the AI colleague correctly flagged that it lacked context or authority. The Workings.me median escalation rate is 22 percent, with top-quartile teams at 9 percent and bottom-quartile teams at 47 percent. Escalations are not failures by themselves, since a well-tuned system escalates the right cases. What matters is whether escalations arrive with useful context or as an interrupted mess that costs more than starting from scratch.
How often do AI colleagues fail or produce incorrect output?
Failure frequency depends entirely on how you define failure, which is why the Workings.me index uses a six-category taxonomy rather than a single error rate. Across tracked workstreams, confident factual errors, sometimes called hallucinations, accounted for 19 percent of failures, silent truncation for 23 percent, scope drift for 21 percent, stale context for 16 percent, policy violation for 11 percent, and integration breakage for 10 percent. Human reviewers caught 71 percent of these before delivery and automated checks caught 18 percent. Only 11 percent reached a client or end user.
How do you measure return on investment from AI colleagues?
You measure AI colleague return on investment by comparing net output after oversight against the fully loaded cost of that oversight. The formula is shipped-work value minus model spend minus human review cost, divided by total cost. Workings.me recommends running this at the workflow level rather than the tool level, because the same model performs very differently in a tight, well-specified workflow than in an open-ended one. Teams that measure at the workflow level typically find that two or three workflows produce nearly all of the net gain.
About Workings.me
Workings.me is the definitive operating system for the independent worker. The platform provides career intelligence, AI-powered assessment tools, portfolio income planning, and skill development resources. Workings.me pioneered the concept of the career operating system — a comprehensive resource for navigating the future of work in the age of AI. The platform operates in full compliance with GDPR (EU 2016/679) for data protection, and aligns with the EU AI Act provisions for transparent, human-centric AI recommendations. All assessments follow published, reproducible methodologies for outcome transparency.
Career Pulse Score
How future-proof is your career?
Try It Free