BenchmarkFinanceDataAI-Native

What It Costs to Run GTM on AI: Our Per-Use-Case Math

N
Published by NativelyDrafted, reviewed, and edited by the team
· 8 min

The first month we ran all seven functions on our own company, the combined inference cost was $203. (Seven is what we ran internally at the time, covering areas like recruiting and finance that we do not sell; the product itself is eight go-to-market use cases.) The equivalent headcount (one person per function at U.S. salary medians) would have run $34,167 in base salary alone. Before benefits. Before tools. Before management overhead. This is what running a department this way actually costs, line by line.

$203

combined monthly inference, all 7 functions (our numbers)

$34,167

equivalent monthly headcount cost: one person per function, salary only

168×

ratio of salary to inference at current Sonnet 4.6 API pricing

~$400

all-in estimate (inference + orchestration infra) for all seven functions

Inference: Natively’s own dogfood numbers (not a study) at Claude Sonnet 4.6 standard pricing: $3.00/1M input, $15.00/1M output (Anthropic). Headcount: U.S. salary medians from BLS occupational surveys; 1.4× loaded multiplier is an industry-standard rule of thumb, not a sourced study.

What “per-utility math” means

Each function runs one business area: one utility. Sales runs outbound. Support answers tickets. Finance closes the books. The per-function cost is what each function spends in raw model inference per month: token count for every research pull, every draft, every classification, times the API price per token.

The reason to break it down by function instead of one blended number is that the economics differ wildly. A support resolution is a short, high-volume action: cheap per token, run thousands of times a month. A GTM strategy brief is a long, lower-frequency action: more tokens per run, but far fewer runs. The per-utility view shows which functions have the widest gap between inference cost and headcount cost, and which are closer than you’d expect.

These are Natively’s own numbers from running these use cases on ourselves. Not a benchmark study. Not a vendor projection. Actual token counts from our production logs, multiplied by the Anthropic Sonnet 4.6 standard API rate. Where we use prompt caching (reusing a shared research corpus across many actions in the same session), the cached read cost of $0.30/1M input applies to those tokens instead of the full $3.00/1M. We note where that matters.

The per-function breakdown

Seven functions, each running end to end. The table shows each function’s primary utility, approximate monthly action volume, and what it actually cost us in inference last month.

Per-utility inference cost: June 2026

Natively’s own dogfood numbers. Claude Sonnet 4.6 standard pricing.

FunctionFocusMonthly actionsInference / moHeadcount equiv.
SSalesOutbound sales~3,500 touchpoints$76$5,000 / mo
GGTMGTM strategy~220 briefs & digests$29$7,083 / mo
MMarketingContent marketing~150 posts & social$6$4,583 / mo
SSupportCustomer support~2,000 resolutions$33$3,750 / mo
RRecruitingRecruiting~800 sourcing actions$19$5,000 / mo
OOpsOps & knowledge~1,420 Q&A + runbooks$28$4,583 / mo
FFinanceFinance & admin~1,810 categorizations$12$4,167 / mo
Total, all functions$203$34,166 / mo

Headcount figures are U.S. salary medians (BLS); 1.4× loaded cost brings total to ~$47,800/month. Loaded ratio: $203 inference vs $47,800 loaded headcount ≈ 235:1. Action counts are our own logged production numbers.

Inference vs. headcount, by function

Each bar pair shows monthly inference cost (green) vs. headcount salary (gray), scaled to the same axis. The green bars are not too thin to see. They really are that small.

SalesOutbound sales$76$5,000 / mo
GTMGTM strategy$29$7,083 / mo
MarketingContent marketing$6$4,583 / mo
SupportCustomer support$33$3,750 / mo
RecruitingRecruiting$19$5,000 / mo
OpsOps & knowledge$28$4,583 / mo
FinanceFinance & admin$12$4,167 / mo
Monthly inference (our numbers)Monthly headcount (salary median, BLS)

What’s in the number, and what isn’t

The $203 is raw inference only. What it doesn’t include:

Add infrastructure and the all-in monthly estimate lands at $350–400 for all seven functions. Still $47,000 below the loaded headcount cost. Still a 120:1 ratio.

The finance function’s share of the inference bill (1,810 invoice categorizations, a month-end close digest, eight anomaly flags) came to $12 last month. Our Notion workspace costs more.

Why the math gets better at scale, not worse

Inference cost is linear with volume. Double the support ticket load and that function’s monthly inference roughly doubles, from $33 to $66. Hiring a second support rep adds about $3,750/month in salary before benefits. At 10× current volume, the gap between the inference line and the headcount line is ten times wider than it is today.

Cache hit rates also improve with scale. When the outbound sales function runs 3,500 touchpoints in a month, a large share of the research corpus (company context, product positioning, signal patterns) is already in cache. Those tokens cost $0.30/1M instead of $3.00/1M. At higher volumes the proportion of cached reads goes up, so the marginal cost per additional action is lower than the average cost, not higher.

Foundation Capital calls this the $4.6 trillion opportunity in their “Service as Software” thesis: the global services market, where responsibility for delivering an outcome shifts from the buyer to the vendor, and AI is the mechanism that makes the economics work. The numbers above are why that shift is possible. At $3 per million input tokens, the cost of running a department is no longer a constraint. It’s a rounding error.

What doesn’t narrow cleanly

The 235:1 ratio holds for the work each agent was specifically designed to do. It narrows fast outside that scope. A research team from Harvard Business School and BCG found that when you put AI on tasks just past the edge of what it does well, people using AI are about 19 percentage points less likely to reach the right answer than people using no AI at all. The model doesn’t signal the boundary. It answers with the same confidence whether it’s well inside its range or not.

That’s the reason the approval step exists: not to slow things down, but because the cheapness of the inference makes it tempting to hand over tasks the AI isn’t actually good at. Each function here has a defined scope. The work outside that scope routes to a person, and the per-utility math only makes sense inside the defined scope. NVIDIA’s research (June 2025) makes this concrete: for agentic workloads, well-scoped small models are 10–30× cheaper to serve than frontier ones: not because they’re worse, but because they’re matched to the task. The same logic applies here.

Frequently asked questions

What does it actually cost per month to run an AI agent department? At Claude Sonnet 4.6 standard pricing and our own logged production volumes, our seven functions run for $203/month in raw inference, or about $350–400/month all-in including orchestration infrastructure. By comparison, the equivalent headcount (one person per function, salary only) runs about $34,000/month.

Does the cost scale predictably as volume grows? Yes. Inference cost scales roughly linearly with action volume. Cache hit rates improve at scale, which lowers the marginal cost per action slightly. Infrastructure costs are largely fixed below a significant scale threshold. Headcount scales in steps: each additional person adds a full salary.

Is this what Natively charges customers? No. Natively prices by use case on the outcome delivered, not on token consumption, and exact pricing is shared up front in the first conversation while we are in private beta. The inference math above is how we account for the underlying cost, not how we bill. The favorable inference economics are what let us offer outcome pricing that still makes sense at the margin.

At $3 per million input tokens, the inference cost of running a department on agents is a rounding error in any budget where it makes sense. The real cost of going AI-native isn’t the compute. It’s the three to five weeks to scope, connect, and prove a single function, and then the month after that, as it runs.

Every figure in the table above is one we could re-run next month and show you again. That’s the point of publishing it. If the numbers shift (because API prices change, because volumes grow, because a function gets re-scoped), we’ll update the table. Natively’s finance function keeps these numbers, and the next version of this post will have six more months of production data behind it, and a more complete picture of what “cost-per-outcome” means once the comparison is apples to apples. We’re not there yet. But we’re publishing this now because the floor is already more interesting than the projections.

Sources & method

Inference pricing: Claude Sonnet 4.6, standard rate: $3.00/1M input, $15.00/1M output; cached reads $0.30/1M input. Source: Anthropic pricing page (July 2026). Introductory rate ($2/$10) applies through Aug 31, 2026; we cite standard to be conservative. Per-agent figures: Natively’s own production logs from running agents on our own company. Not a study; not a vendor projection. Token counts × API rate. Headcount: U.S. salary medians from BLS occupational employment statistics. Loaded cost (1.4× salary) is an industry rule of thumb for benefits + payroll taxes + T&E; our own estimate, not a sourced study. Foundation Capital $4.6T: “Service as Software” framework; primary URL foundationcapital.com/ai-service-as-software/. NVIDIA Research (June 2025): “Small Language Models are the Future of Agentic AI”: 10–30× cheaper to serve at comparable agentic task performance. HBS / BCG “Jagged Frontier”: Dell’Acqua et al., WP 24-013 (2023), 758 consultants, AI performance above vs. below the frontier.

Sources

  1. 1.Anthropic: Claude Sonnet 4.6 API pricing: $3.00/1M input, $15.00/1M output (Jul 2026)
  2. 2.Foundation Capital: AI leads a service as software paradigm shift: the $4.6T opportunity (Apr 2024)
  3. 3.Foundation Capital: The $4.6T Services-as-Software opportunity: Lessons from the first year (Jul 2025)
  4. 4.NVIDIA Research + Georgia Tech: Small Language Models are the Future of Agentic AI: 10–30× cheaper to serve for agentic tasks (Jun 2025)
  5. 5.Harvard Business School / BCG: Navigating the Jagged Technological Frontier (Dell'Acqua, McFowland, Mollick et al., Working Paper 24-013, 2023)
  6. 6.U.S. Bureau of Labor Statistics: Occupational Employment and Wage Statistics (OEWS)

See cold outreach campaigns running.Booked meetings, run end to end and stopped for your approval before anything is sent, published, or spent. Live in days, and the system stays in your account.

How cold outreach campaigns works →