【AIは“成果報酬型”時代へ】SaaSは終わるか?/企業価値2.4兆円「Sierra」共同創業者クレイ・バヴォア氏&日本統括・森川馨太氏/AI予算超過問題/日本での勝ち筋【PIVOT TALK】
PIVOT 公式チャンネルAcross the technology sector, the phrase "AI agent" has quickly become pervasive, yet its practical enterprise definition remains contentious. In an in-depth conversation with business media platform PIVOT, Clay Bavor, co-founder of Sierra and former 18-year Google executive, laid out the operational philosophy and commercial architecture behind the startup's rapid ascent. Rather than selling software seats or charging for raw token consumption, Sierra has staked its business on an outcome-based model, tying its compensation directly to solved customer problems and completed enterprise transactions.
Bavor discussed why building reliable production-grade agents is far more complex than wrapping frontier language models, how rising token expenses are reshaping corporate capital allocation, and why customer-facing AI may expand rather than merely eliminate human support roles.
Defining the Agent: From Simple Language Generation to Problem Resolution
When Sierra launched in early 2024, the concept of an AI agent required dedicated education—so much so that the company published an explanatory guide. Today, the term saturates San Francisco billboards, often with little distinction made between conversational bots and autonomous software.
Bavor defines an agent as a fundamentally new category of software that can reason, make decisions, and execute actions by leveraging the reasoning capabilities of large language models (LLMs). Rather than solely generating responses or drafting text, an agent connects directly to enterprise backends—such as order management databases, CRM systems, and transactional workflows—to execute end-to-end tasks on behalf of a company.
Sierra focuses specifically on customer-facing interactions across the entire consumer lifecycle, spanning customer service, technical troubleshooting, and revenue-generating sales advisory.
Regarding the widespread prediction that AI will cleanly replace human labor in support channels, Bavor notes that routine inquiries (such as order status checks or product returns) are already handled effectively by autonomous agents. However, he cautions against simple one-to-one labor replacement forecasts, invoking Jevons paradox: when the cost of executing a task drops dramatically, overall consumption tends to surge.
Freed from repetitive technical and transactional inquiries, enterprises may choose to redeploy staff toward higher-value, proactive outbound outreach and complex relational support that genuinely requires human empathy and judgment.
Expanding Beyond Support: Long-Horizon Workflows and Enterprise Scale
Sierra began in customer support—working with early clients like SiriusXM to troubleshoot satellite radio devices and transfer subscriptions across vehicles—primarily because language models were initially limited in reasoning depth. As model capabilities matured, the platform expanded into multi-step, multi-session customer journeys.
Under an initiative designated as "Horizon," Sierra's agents are deployed across extended timelines spanning weeks or months. These agents can guide consumers through complex financial workflows, such as home mortgage origination, insurance onboarding, car accident claim processing, or comprehensive travel planning.
The enterprise footprint behind this progression includes:
- Deployment across 40% of the Fortune 50.
- Adoption by one in three of the world’s leading banks.
- A client base where 30% of companies generate more than $10 billion in annual revenue.
- Practical volume milestones, such as originating over $1 billion in new mortgages each month for Rocket Mortgage in the United States, alongside powering conversational real estate search for Redfin.
Why Enterprise In-House Builds Falter: Voice Nuance, Guardrails, and Dialects
With foundational model providers like Anthropic and OpenAI introducing sophisticated coding tools, many enterprises consider building agents entirely in-house. Bavor contends that while building an initial demo is straightforward, taking an agent into production for millions of consumers requires addressing hundreds of granular technical hurdles that raw foundation models cannot solve out of the box.
Voice Activity and Ambient Noise Management
In conversational voice interfaces, latency and fluid conversational turn-taking are critical. Frontier models do not inherently distinguish between an active interruption (a customer interjecting to change the subject) and passive conversational affirmations (saying "uh-huh" or "yeah" while the agent is speaking). Sierra developed dedicated proprietary models specifically for voice activity detection to handle conversational pacing.
Additional real-world audio edge cases include:
- Filtering background television noise, barking dogs, and crying infants.
- Handling multi-party phone conversations—such as in healthcare environments where an elderly patient is assisted by an adult child—requiring the system to separate speakers and verify from whom legal medical consent is being obtained.
- Regional localization, such as equipping voice agents in Japan to understand and converse fluently in regional dialects like Kansai-ben alongside standard Japanese.
Safety, Guardrails, and Simulation Testing
Because enterprise agents are empowered to execute transactional actions—such as processing cash refunds or altering payment methods—they pose security and compliance risks if left unbounded. Beyond strict customer data compartmentalization, enterprise agents require behavioral guardrails ensuring they execute designated processes without deviating.
To validate an agent prior to deployment, Sierra constructs conversational simulators that subject the agent to 10,000 to 100,000 synthetic interactions. These simulations expose the agent to edge cases, contradictory instructions, and hostile inputs to verify adherence to enterprise policy before live traffic is enabled.
Dismantling the SaaS Model: Outcome-Based Pricing vs. Token Billing
A central divergence between Sierra and traditional enterprise software vendors lies in its commercial pricing structure. Bavor argues that conventional seat-based pricing is obsolete for autonomous software, while raw consumption metrics—such as charging per message or per API call—create misaligned incentives where the vendor profits from inefficient, prolonged exchanges.
Sierra operates primarily on outcome-based pricing:
- In customer service: Clients pay if and only if the agent completely and successfully resolves the customer's issue without human intervention.
- In sales advisory: Sierra collects a commission-style fee only when the agent directly drives a completed purchase or product upsell.
Addressing the practical difficulty of defining a "resolved" interaction—given that consumers rarely end phone calls with standardized declarations of satisfaction—Bavor explained that Sierra relies on straightforward resolution criteria. In ambiguous edge cases or gray areas, Sierra absorbs the cost rather than billing the customer, arguing that an imperfect outcome-based metric aligns enterprise incentives far better than charging for unvalidated conversational volume.
The Economics of Tokens and Internal Corporate Token Budgets
To insulate enterprise clients from escalating inference expenses, Sierra absorbs token volatility entirely; customers never receive an underlying token bill. To make this economically sustainable, the company avoids relying solely on expensive frontier models. Instead, it deploys a "constellation of models," combining major frontier LLMs with proprietary, fine-tuned, post-trained models engineered to perform specialized domain tasks faster, cheaper, and more accurately.
Internally, however, token economics have become a serious operational consideration. Sierra runs much of its own engineering and operational workflows through an internal agent named "Pine Cone," which currently generates approximately 70% of the company's code and assists executive staff with financial modeling and resource allocation.
This intensive usage revealed noticeable cost dynamics:
- Senior software engineers at Sierra routinely consume over $100,000 annually in inference tokens through coding agents.
- Runaway conversational contexts or looping API calls can quickly generate outsized expenses for minor tasks (Bavor cited an internal strategic query that inadvertently incurred $173 in context costs).
Bavor predicts that corporate finance will soon formalize token spending much like operational expenses or travel budgets. CFOs will likely restructure headcount planning to incorporate both base compensation and an allocated token budget per employee, alongside internal governance policies governing appropriate and cost-effective AI usage.
Market Traction, Deployment Speed, and Expansion in Japan
Sierra reported reaching $100 million in annual recurring revenue (ARR) within seven quarters of operation, progressing to $150 million in eight quarters, and hitting $200 million by its ninth quarter. In the US, Bavor reports that Sierra’s customer deployments touch roughly 90% of retail consumers and 50% of families in healthcare networks.
Enterprise deployment velocity has also shifted. Regulated global healthcare insurer Cigna went live on Sierra’s infrastructure in 56 days, contrasting sharply with legacy enterprise software rollouts that often span multiple quarters or years.
In Japan, where Sierra has established early deployments with SoftBank and mobile carrier LINEMO, the company recorded customer resolution rates reaching 97% and customer satisfaction (CSAT) scores of 93%. Bavor highlighted the Japanese service philosophy of omotenashi—delivering meticulous, non-transactional hospitality—as a vital benchmark for AI agent development. Because cultural interactions prioritize craft, quality, and the singular nature of each encounter (ichi-go ichi-e), meeting the expectations of Japanese consumers serves as a stringent test for the conversational fluency and nuance of their underlying technology.
Industry Consolidation: How Winners Will Be Decided
With artificial general intelligence (AGI) potentially approaching within the coming years, Bavor acknowledges that the conversational AI space is heavily contested. Competitors span hyperscalers (including his former employer, Google), venture-backed AI-native startups, and incumbent customer service software providers attempting to retool their platforms.
According to Bavor, long-term survival in this sector will not be determined by surface-level model demonstrations or raw architectural claims, but by two operational factors:
- Enterprise Trust: Operating agents that interface directly with an organization's most critical asset—its end customers—requires flawless security, rigorous guardrails, and consistent reliability under heavy traffic.
- Measurable Business Impact: Software vendors must prove tangible outcomes—demonstrable increases in completed loans, lower net resolution costs, improved retention, or concrete healthcare administrative outcomes.
In an enterprise environment increasingly fatigued by speculative AI initiatives, Bavor concludes that the companies that survive market consolidation will be those willing to stake their commercial success entirely on whether their systems actually resolve the underlying problem.
