Home Works Services AI Blog About Contact

AI Workflow Notes · AI Workflow BOFU

AI Agent ROI: 3 Variables That Decide in 90 Days

The math behind AI agent ROI: task volume, loaded hourly cost, and error cost. Three worked examples across solo, small team, and mid-size builds.

Keng · · ~ 7 min read · AI ROIAI AgentAutomation ROIBusiness AutomationClaude CodeSolopreneurSmall Business AI

A founder DM’d me on LinkedIn last month. He runs a 6-person marketing agency in Austin, wanted to know if a $12,000 agent build would pay back. I asked three questions. None about his stack or model or framework. Three business questions. Payback came out to 47 days.

Most agent conversations start in the wrong place. People argue Claude vs GPT, self-hosted vs API, LangChain vs raw SDK, before doing the math on whether the thing is worth building. Then they burn 6 weeks and $15K on something that saves 4 hours a week, and blame the tech.

After running my studio for a decade and auditing 20+ small business agent builds this year, I can tell you the ROI math is boring. Boring math prevents expensive mistakes. Three variables decide almost every case: task volume per month, fully-loaded hourly cost of whoever does it now, and error rate times cost per error. Multiply, divide the build cost by three months, compare. Answer before you write a prompt.

Here’s the formula, three worked examples, and where the math breaks.

Variable 1: Task Volume Per Month

Volume is the easiest number to get and the most under-measured. Count the exact task, not the category.

“Email responses” isn’t a task. “Reply to inbound sales inquiries with pricing plus a calendar link” is. First is too vague to measure. Second you can count from your inbox in 5 minutes.

Benchmarks from client audits this year:

  • Solo founder: 40-120 recurring tasks worth automating, mostly admin, content, inbound routing
  • Small team (3-8 people): 200-500 tasks across ops, sales enablement, reporting
  • Mid-size (15-40 people): 800-2,500 tasks, concentrated in support, compliance, data reconciliation

The trap: teams count the tasks they hate and ignore the boring ones they’ve stopped noticing. When I run an audit I ask people to log every discrete action for 3 days. Volume comes in 2-3x higher than the gut estimate.

Volume also compounds. A task running 20 times a month today, in a growing business, runs 60 times a month in 18 months. Build cost is fixed. If you’re on a growth trajectory, discount your payback calculation by 20-40%.

One caveat. Not every high-volume task is worth automating. If it takes 30 seconds and runs 50 times a month, that’s 25 minutes of human time, not enough to justify a $3K build. The formula below catches this automatically.

Variable 2: Fully-Loaded Hourly Cost

Salary isn’t the number. Fully-loaded hourly cost is.

Rule of thumb for North American markets:

  • Admin or VA: $30/hr loaded (base $20 plus benefits, software, management overhead)
  • Coordinator or associate: $50/hr loaded
  • Specialist or senior IC: $100/hr loaded
  • Founder time: $150-300/hr, and this is where people underprice worst

Loaded means base comp, employer taxes, benefits, PTO amortized, software seats, infrastructure, and management attention. Dividing salary by 2,080 hours misses most of it. Real number sits at 1.5-1.8x that.

For founders it’s different. Your hourly cost is the highest-value alternative use of that hour. If a client-work hour generates $250 in revenue, the hour you spend copying data between spreadsheets costs you $250, not the $0 you’re paying yourself in salary. Founders undervalue their own time by 3-5x, which is why solo automation ROI looks wrong on paper but feels right in practice.

Contractor rates are the cleanest input. Currently paying a VA $25/hr? Use $30 (add 20% for management overhead). Would need to hire an ops person at $65K base? Use $50/hr loaded. Alternative is your specialist at $120K? Use $100/hr.

Don’t use the cheap offshore rate unless that’s your realistic alternative. If you’d never actually hire an $8/hr VA to handle regulated customer data, don’t use $8/hr in your math. Use the rate of the person you’d actually hire.

Variable 3: Error Rate Times Cost Per Error

Everyone forgets this variable, and it’s often the largest.

Every manual process has an error rate. Repetitive data entry runs 1-4% depending on fatigue and attention. Customer-facing message routing gets misrouted 5-15% in teams I’ve audited. Manual invoice reconciliation misses 2-8% of exceptions.

Cost per error varies wildly:

  • Wrong pricing quoted to a lead: $0-2,000 (lost deal or margin hit)
  • Missed follow-up on a warm inbound: $500-10,000 depending on average contract value
  • Miscategorized expense in bookkeeping: $50-500 (accountant time plus tax risk)
  • Compliance violation from missed disclosure: $1,000-50,000+
  • Duplicate order fulfillment: goods plus shipping plus customer damage

Math: route 200 sales inquiries a month, misroute 8% at $1,200 average deal impact, that’s $19,200/month in error cost. A well-instrumented agent at 1% error saves you $16,800/month before you count labor hours.

I’ve seen agent builds with negative labor ROI justify themselves entirely on error cost reduction. A regulated fintech client had an $18K build that saved 6 hours of ops time per week (small) but reduced compliance misses from 3 per quarter to 0 (huge). Three avoided misses in year one saved them roughly $75K in penalties and remediation.

If you genuinely can’t estimate error cost, use a conservative $0. Formula still often works. Do the honest exercise first: write down the last three times a human error in this workflow cost you money. Add them up, divide by the period. The real number is usually meaningful.

The 90-Day Payback Formula

Whole thing on one line.

Monthly waste = (task volume × human hours per task × loaded hourly cost) + (error rate × cost per error × task volume)

Build it if: monthly waste × 3 > build cost

Three months is the payback window I use because it matches how fast tools change. If you can’t recoup in 90 days, the risk of the underlying model, pricing, or agent framework changing under you starts to matter more than the theoretical 12-month ROI. I watched two client builds get partially obsoleted by Claude Skills launching in late 2025. The ones that paid back in the first quarter didn’t care.

Worked examples at three scales.

Solo founder, $3K build: 60 tasks/month × 0.4 hours × $150/hr founder time = $3,600/month labor plus $500/month error cost = $4,100/month waste. Payback: 22 days. Build it. Variables here are usually founder time (undervalued) and content or admin volume (higher than you think). See the Solo Stack Method for how I structure small builds around Signal/Strategy/Skill/Ship.

Small team, $10K build: 300 tasks/month × 0.25 hours × $50/hr coordinator = $3,750/month labor plus $2,000/month error cost = $5,750/month waste. Payback: 52 days. Build it. Sweet spot for most 5-15 person businesses. Enough volume to justify custom work, few enough people that a single agent takes real headcount pressure off.

Mid-size, $25K build: 1,200 tasks/month × 0.15 hours × $80/hr blended = $14,400/month labor plus $5,000/month error cost = $19,400/month waste. Payback: 39 days. Build it. Mid-size wins come from volume, not per-task savings. Individual task is often small. Aggregate is enormous.

Where the math breaks: task volume below 20 per month, human time per task below 5 minutes, no measurable error cost, no growth tailwind. If three of those four apply, don’t build. Use a $20/month SaaS tool or leave it manual.

Doing the Math Yourself

If the three variables look promising, you don’t need me to build the case. The math is the case. What I help with, when clients come to us at enterprise consulting, is the audit that gets you to accurate numbers. Most teams underestimate volume by 2x and undervalue their time by 3x, and both errors hide ROI already sitting in the workflow. We publish the full framework and a calculator on that page, along with the phased rollout structure we use with clients. If you’re deciding on a Q4 build, start there before talking to any vendor. For the conceptual foundation the whole approach sits on, the Solo Stack Method walks through the four-layer framework we teach.

FAQ

What counts as a “task” for volume counting?

A discrete recurring action with a clear input and output. “Reply to inbound RFPs with our capabilities deck plus a 20-minute intro slot” is a task. “Manage our sales pipeline” is not, that’s a workflow made of many tasks. When in doubt, ask yourself: can I describe this in 3 sentences? If yes, it’s a task. If it needs a paragraph, break it down further before counting. Undercounting granularity is the most common measurement error I see in audits.

What goes into “fully-loaded hourly cost”?

Base comp, employer taxes, health and retirement benefits, PTO amortized, software seats, workspace, and management overhead (roughly 10-20% of a manager’s time per direct report). For contractors, use their bill rate plus 15-20%. For founder time, use the highest-value alternative use of that hour, usually client billing rate or revenue per hour if you’re product-led. Never use base salary divided by 2,080.

What if error cost is hard to quantify?

Qualitative exercise first. Write down the last three times a human error in this workflow cost you money, time, or reputation. If you can name three concrete incidents in the last 12 months, error cost is real and probably significant. Estimate conservatively: take the smallest of the three, treat it as the average. If you can’t name any incidents, either the rate is genuinely low or you’re not measuring failures, and defaulting to $0 is honest.

Why 90 days and not a 12-month ROI?

Two reasons. Model and tool landscape changes fast enough that 12-month projections carry real obsolescence risk. A build that pays back in month 10 might get eaten by a native platform feature in month 8. Second, 90-day payback is a strong signal that task volume and error cost are actually high enough to matter. If the honest math needs 8 months to break even, the candidate is marginal and probably not worth the ongoing maintenance and prompt-tuning either.