Skip to content
Back Can AI Agents Replace Freelancers? What Three Benchmarks Actually Measured

Can AI Agents Replace Freelancers? What Three Benchmarks Actually Measured

Best agent on real freelance projects: 2.5% at launch, about 16% now. Which means it still fails five out of six paid jobs.

Blago Yanakiev
Blago Yanakiev

Sep 08, 2026

AI Future of Work studie
TL;DR

Scale AI and the Center for AI Safety tested agents on 240 real, already paid freelance projects. At launch in October 2025 the best agent finished 2.5% of them end to end. On the live leaderboard read on 7 September 2026, the top score is 15.80%: six times higher, and still failing five of every six paid jobs. Upwork found that pairing an agent with a human expert lifted completion rates by up to 70%. The job is not competing with an agent. It is selling supervised delivery.

Every few months a chart claims freelance work is about to be automated. The most useful test of that claim is a benchmark that took real freelance projects, ones a human had already been paid for, and handed them to agents.

The Remote Labor Index

Scale AI and the Center for AI Safety published the Remote Labor Index on 29 October 2025. The method matters more than the headline. They collected 240 completed freelance projects across 23 domains: video production, CAD and architectural drawing, graphic design, game development, writing, data retrieval, interactive data visualisation. Combined original earnings: $143,991. Median project value: $200. Median duration: about 11.5 hours. Ordinary paid jobs, not toy tasks.

At launch, the best-performing agent was Manus, which completed 2.5% of the projects end to end, earning $1,720 of the pool. Every tested model scored under 3%.

The failure pattern is the interesting part. Agents did comparatively well on audio and image generation, some data retrieval and some writing. They failed on work requiring complex editing, tool use and precise multi-step specifications. Of the failures, 45.6% were quality problems, 35.7% were incomplete or malformed deliverables, 17.6% were technical or file integrity issues, and 14.8% were internal inconsistencies. They produce something, and the something is not deliverable.

One caveat: Scale AI sells AI evaluation, so this is its own benchmark rather than neutral ground truth. The saving grace is that the method is checkable and the jobs were real.

The number on the writing day

The RLI has a live leaderboard, which makes it the rare AI statistic you can re-check yourself. Read on 7 September 2026, the top entry sits at 15.80%, with the next at 8.33%, then 6.25%, and a cluster of models between roughly 2.5% and 5%. Still the same 240 projects.

Treat that 15.80% as a photograph, not a fact. It will be wrong within weeks, in one direction. Two things are worth saying anyway. The trend is real: a sixfold rise in under a year. And the level is still low: even the leader fails roughly five of every six real projects, and the failures cluster in exactly the multi-step, judgment-heavy work experienced freelancers sell.

Where the human actually adds value

Upwork ran the complementary experiment and published the Human+Agent Productivity Index on 13 November 2025. Across more than 300 real client projects, deliberately picked as simple, well-defined and low-complexity, pairing an AI agent with a human expert increased completion rates by up to 70% compared with agents working alone.

The second figure is the one to quote at clients: those simple, well-defined projects, the ones where an agent has a real shot on its own, make up less than 6% of Upwork's total gross services volume. Upwork is a marketplace with a commercial interest in "freelancers are still needed", so weigh the methodology rather than the headline. But the direction survives the discount.

Microsoft's 2026 Work Trend Index, based on 20,000 workers across ten countries plus its own product telemetry, adds the skills side. Asked what human capabilities matter most in AI-heavy work, respondents named quality control of AI output (50%) and critical thinking (46%). The report also found organisational factors driving more than twice the impact on AI outcomes that individual factors do (67% against 32%), while 65% of AI users fear falling behind and only 13% say they are recognised for reinventing their work.

Put those three together and you get the actual job description: someone who runs the agent, catches what it got wrong, and stands behind the result.

What that looks like in practice

Lucas Chevillard does email marketing and CRM strategy, works with seven to ten clients a year and three at a time, and spent 2026 building himself a system that turns recorded client calls into to-dos, content ideas and draft deliverables. His numbers on stage: 57 calls processed, 159 content ideas generated, and the write-up after a discovery call down from one to two hours to about ten minutes.

The line that matters is what he does with the output:

"Now I know this is not the document I'll give to my client. That's for sure not. But it's a pretty good starting base for me to have all the important information to process and to read when I deliver such a project." (Lucas Chevillard)

He is equally firm on the review step: everything he hands over, he goes through thoroughly, because "if you create or just write with AI, whoever is delivering that, you really need to stand by everything that's written."

His most useful question for anyone building this is a positioning question, not a technical one. Working through where your time goes, he asks: what do you bring that is not in any tool? Why is the output yours? That answer is the part the agent cannot supply, and the part you are charging for.

He is also clear on the client conversation. He does not sell "I will build you an agent". He is hired to audit or advise, notices a bottleneck, and builds one small thing to show what is possible. On recording calls he flagged his own gap: he wants the permission written into his client contracts and has not finished doing it, which is the right instinct given GDPR.

Why the work does not run out

Julien Look is a software engineer and consultant who works on AI adoption inside companies. His talk was about why those projects fail, and he showed the audience two numbers he works with: about 88% of people inside organisations already use AI in some form, while only about 5% of AI projects capture measurable value. His figures, presented on his slides.

His diagnosis is not technical. Asked whether the biggest barrier is technology, budget or change management, the room mostly picked change management, and he agreed. The failure mode he described from a large industrial client is trust: people do not believe the output, so they stop using the system. His fix is making each step traceable and involving the people who do the work in designing the workflow, before any tool appears.

Then the part that concerns your invoice:

"Team sizes are shrinking. That is very true, because you don't need a ten-person team to build SaaS products anymore. Three people might be enough." (Julien Look)

And immediately after:

"Even if teams are getting smaller, we are moving at a way faster pace and organizations are still struggling to drive this whole adoption. I think there will be only more work created in the future for people in our position." (Julien Look)

That is the honest version of the outlook. Smaller teams per project, more projects, and a persistent shortage of people who can make an adoption actually land. His practical tip for freelancers pitching this work: transformation runs in years, not in a six-month project, so price and contract accordingly.

Pricing the middle, without inventing a number

Here is what nobody can tell you honestly: there is no primary research on pricing agent-assisted delivery. Everything on the subject is agency marketing with no disclosed method. Anyone quoting a specific premium for "human plus agent" made it up.

The Upwork index does support a direction. Value concentrates in supervised delivery rather than at either extreme, so the defensible moves are:

  • Price the outcome and the accountability, not the hours the agent saved you. The client is buying that you stand behind it.
  • Do not pass the full time saving through as a discount. If you cut two hours off a deliverable and cut the price by two hours, you have handed the entire gain to the client and kept the risk.
  • Charge separately for the review layer where it is heavy. Verification is now a service, and Microsoft's respondents named it the top human skill.
  • Say what is automated when the client asks, and be able to show your check. That is a differentiator while most competitors are quiet about it.

What to do on Monday

  1. List your last ten projects and mark each one as something an agent could plausibly finish alone. If more than one or two qualify, that is your exposure.
  2. Take one recurring deliverable and time the agent-assisted version end to end, including your review. That number is your real productivity gain, not the demo's.
  3. Write one paragraph for your proposals about what you check and what you take responsibility for. Verification is a sellable line item.
  4. Put recording and AI-processing permission into your contract template if client calls go anywhere near a model.
  5. Pick the one part of your output that is unmistakably your judgment and make it more visible in the deliverable, not less.

For what all this does to rates on both sides of the market, read AI premium or AI discount. For the build side with real costs, two freelancers show their AI assistants. And for keeping your own judgment sharp while using the tools daily, use AI without losing your edge has the study and the method. A free 9am profile puts you in front of the companies who have worked out they need the supervisor, not just the agent.

Freelance Unlocked is co-organized by 9am together with Uplink and freelancermap. This article draws on the sessions of Julien Look and Lucas Chevillard at Freelance Unlocked 2026. Watch the full talks above, and join us at the next edition: freelanceunlocked.com.

Blago Yanakiev

Co-founder & CPO

Blago is a product leader and SaaS founder. He runs product at 9am and directs events and growth for the Freelance Unlocked conference.

9am profileLinkedIn

Latest Articles

Germany's Income Tax Reform: What Actually Changes for the Self-Employed

Germany's Income Tax Reform: What Actually Changes for the Self-Employed

45% from €250,000, a new 47% step from €280,000. For a sole trader on €300,000 profit that is about €1,235 a year.

No Ambition: Why Germany Announces Reforms Nobody Feels, and Who Profits

No Ambition: Why Germany Announces Reforms Nobody Feels, and Who Profits

152 measures in the subjunctive, a new tax bracket €30,000 above the old one, 44% for the AfD. Three examples, one pattern.

Germany's Leaked Freelancer Status Reform: What 'New Self-Employment' Would Cost You

Germany's Leaked Freelancer Status Reform: What 'New Self-Employment' Would Cost You

Legal certainty in exchange for 16.74% of every invoice: what the leaked 'new self-employment' draft asks, and where it stands today.