How we price an agentic retainer: the three numbers that set the floor
Key takeaways
- Hour-based pricing fails for agentic services because the client pays for outcomes, not sessions — once agents run overnight, there is no honest hourly count.
- The floor for an agentic retainer has three components: replacement cost (what a full-time hire would cost to do the same work), interrupt load (how much human attention the stack still demands per month), and knowledge transfer overhead (the cost of reversing the engagement).
- Scoping by outcomes, not tasks, prevents scope creep: define what the engine produces each month, not every step it takes to produce it.
- We raised effective hourly rate from $140 to $610 over four quarters without raising headline prices — the model changed, not the number on the invoice.
Most agency pricing is built on a simple lie: time in equals value out. Bill the hours, collect the check. It worked fine when every deliverable required a human sitting at a keyboard. It does not work when your stack runs overnight and processes 800 tasks while everyone sleeps.
Here's the concrete problem. An agent we run for a client executed 847 data enrichment and outreach tasks between 11 PM and 6 AM on a Tuesday. Under hourly billing, what do you charge? The compute cost was $4.20. The human oversight that night was zero. But the outcome — 847 qualified records updated, 200 sequences triggered — would have taken a two-person ops team roughly 40 hours to replicate manually. Billing $4.20 is absurd. Billing 40 hours of phantom labor is dishonest. The hourly model doesn't bend here — it breaks.
We spent about six months trying to patch the old model before accepting that it needed replacing. What we landed on is a floor-based retainer built from three components. Each component is independently defensible. Together they set a price that reflects what the client is actually receiving.
Component 1: Replacement Cost
The first question we ask is: what would it cost this client to produce the same outcomes without us?
For most engagements, the honest answer is a full-time senior hire — sometimes two. A senior growth ops manager in a major US market runs $110k–$130k in base salary, plus benefits, equity, and management overhead. Call it $120k/yr all-in, or $10,000/month. That's the floor anchor. We are not cheaper than that number — we are the alternative to it, and we deliver more throughput.
This isn't a negotiating tactic. It's a factual benchmark. If a client can hire a single person to do what our stack does, they should. We only make sense when the scope exceeds what one hire can cover, or when the speed and consistency of an agentic system materially outperforms a human team. In those cases, $10k/mo is not a ceiling — it's the starting point for the conversation.
Worked example: a client running a mid-market e-commerce brand needed continuous catalog optimization, paid search monitoring, and weekly performance reporting. Staffing that function fully would require one senior analyst ($90k) and one coordinator ($55k) — $145k/yr, or roughly $12,100/mo before overhead. Our retainer for that scope sits at $11,500/mo. The replacement cost floor makes that number self-evident.
Component 2: Interrupt Load
Agents are not autonomous. They surface edge cases, flag anomalies, and occasionally do something unexpected that requires a human to make a judgment call. That human attention has a cost, and it belongs in the floor.
We track interrupt load per engagement. For a typical mid-complexity retainer, it runs 10–15 hours of senior attention per month — reviews, escalation handling, monitoring checks, and the occasional 'why did it do that' investigation. At our internal rate of $250/hr for senior practitioner time, 12 hours/mo = $3,000/mo floor contribution.
This number matters for two reasons. First, it's real — if we don't price it, we absorb it as margin erosion. Second, it communicates something important to the client: this is not a set-it-and-forget-it system. There is a human in the loop, and that human is expensive and intentional. Clients who understand interrupt load stop asking why they can't just 'turn it on and walk away.'
Interrupt load also scales with client complexity. A client with clean data, clear decision rules, and low exception volume might run 6 hours/mo. A client with messy CRM data, frequent campaign pivots, and a high-touch approval process might run 20. We scope it explicitly during discovery and build it into the floor accordingly.
Component 3: Knowledge Transfer Overhead
Every engagement accumulates institutional knowledge — about the client's systems, their edge cases, their preferences, their history. That knowledge lives in our agents, our documentation, and our team's heads. If the client ever wants to leave, extracting and transferring that knowledge costs real time and money.
We price this as an ongoing floor component, not a one-time exit fee. Here's why: the switching cost the client implicitly holds is a form of value we're continuously creating. Ignoring it in the pricing model means we're delivering value we're not capturing.
In practice, we estimate knowledge transfer overhead at roughly $1,500–$2,500/mo for a standard engagement, based on the documentation, versioning, and handoff readiness work we maintain continuously. This isn't a penalty for leaving — it's the cost of keeping the engagement portable and auditable. Clients who ask about it usually appreciate the transparency. It signals that we're building something they could own, not a black box they're renting.
The three components together — replacement cost, interrupt load, and knowledge transfer overhead — give us a floor that's grounded in real costs and real value. For a typical mid-market engagement, that floor lands between $14,500 and $16,000/mo before any performance or outcome-based layer.
Presenting This to Clients Who Think in Hours
Every client asks the same question eventually: 'How many hours is that?' It's a reasonable question from someone who has only ever bought services by the hour. Our job is to redirect it without being dismissive.
The language we use: 'We don't sell hours — we sell a monthly outcome. The invoice line item is [specific deliverable], not a time block. If you want to benchmark the value, the right comparison is what it would cost you to staff this function internally, not what a freelancer charges per hour.' Then we walk them through the replacement cost math. Most clients, once they see the staffing comparison, stop asking about hours. The ones who don't are usually not the right fit for an agentic retainer.
The deeper reframe is this: hourly billing creates a perverse incentive where efficiency is penalized. If we build an agent that does in 2 hours what used to take 20, hourly billing would cut our revenue by 90%. Outcome-based pricing aligns our incentives with the client's — we win when they win, and we're rewarded for building systems that get faster over time, not slower.
The Onboarding Mistake
The most common pricing error we see — and made ourselves early on — is treating Month 1 the same as Month 6. Onboarding a new client is not steady-state work. It is discovery, integration, configuration, and knowledge acquisition compressed into 30 days. In our experience, onboarding costs roughly 3x a steady-state month in actual labor and attention.
If your retainer is $12,000/mo and you charge $12,000 for Month 1, you've just subsidized the client's onboarding to the tune of $24,000. That's margin you will never recover. The fix is straightforward: either charge a separate onboarding fee (we typically price this at 1.5–2x the monthly retainer) or structure Month 1 at a higher rate with a clear explanation of why. Clients who push back hard on onboarding pricing are signaling that they don't value the setup work — which is a useful signal to have before you've done it.
We now present onboarding as a distinct line item in every proposal: 'Month 1 — Onboarding & Integration: $X. Months 2+: $Y/mo.' It's cleaner, it's honest, and it protects margin on the work that is genuinely the most labor-intensive part of any engagement.
Price the Outcome, Not the Session
The Avakata pricing rule is one sentence: price the outcome, not the session. In practice, this means every invoice line item is a monthly outcome — 'Paid search management and optimization,' 'Agentic catalog enrichment and monitoring,' 'Weekly performance reporting and strategic review' — not a time block. The client is buying a result that recurs. The floor components ensure we're not delivering that result at a loss. The outcome framing ensures the client understands what they're paying for. When both sides of that equation are clear, pricing conversations get shorter and retention gets longer.
Frequently asked questions
How do you price AI agent services for clients?
Price agentic engagements on a monthly outcome retainer, not hourly sessions. The floor is built from three components: replacement cost (what the client would pay a full-time hire to produce the same output), interrupt load (the overhead of managing async agent operations and exception handling), and knowledge transfer overhead (the cost of embedding domain context into the system). Hourly pricing breaks down when agents run continuously — a single agent can execute hundreds of tasks in a month with no direct human-hours attached. Avakata prices what the engine produces each month, not the sessions or tasks it takes to get there.
Should agentic retainers be higher or lower than traditional agency retainers?
Agentic retainers should be higher in effective value delivered, and the floor is structurally higher than a traditional time-and-materials retainer — even if the headline number looks similar. A traditional retainer buys a fixed block of human hours. An agentic retainer includes continuous operation, replacement-cost-level output, interrupt management, and embedded institutional knowledge. Those components don't exist in legacy pricing models, so the cost basis is fundamentally different. The common mistake is discounting an agentic engagement to match what a client paid their last agency — that anchors to the wrong comparator and underprices the actual output being delivered.
What is replacement cost pricing for AI services?
Replacement cost pricing sets the retainer floor by asking: what would a client pay to hire a full-time employee or team to produce the same outcomes? If a senior operations hire costs $120k per year ($10k per month), that's the minimum an agentic engagement should cost — because the client is receiving equivalent or greater output. The method prevents underpricing by anchoring to a real market comparator (the labor market) rather than to hours worked or tasks completed. It's particularly useful when scoping agentic work because the output volume is high and the human-hours are low, which makes hourly-rate comparisons misleading.