Owning the outcome: Bessemer's AI-Native Services evaluation framework

How Bessemer evaluates which services markets are most ripe for AI disruption.

Great investments pair inventive technology with an inventive business model. In the cloud era, software won by becoming a system of engagement or a system of record: different models, both capable of compelling ROI. These models supported businesses to more efficiently do their work; but now, AI can deliver the work itself; the outcome is the product.
 
Professional services firms have always been measured on outcomes. A law firm doesn't sell you help understanding a contract; it sells you a redlined one. A TPA doesn't sell you claims software; it sells you a closed claim. That much hasn't changed and isn't going to. 
 
What's changing is what's doing the work behind that promise. Now it can be a machine instead of a person, sitting inside the delivery of the service itself rather than just its front end. Technology has always touched services at the experience layer (e.g., better funnels, better onboarding, better interfaces) but never the delivery layer. AI changes delivery itself, and that creates a wildly more attractive business.
 
The availability of agents to deliver services doesn't necessarily mean that every services market transforms into an autonomous, high-margin dominated by agents. So, to identify investable opportunities in AI-native services, we're exploring three core questions:
  1. Is this a market that's structurally ready to be taken — fragmented supply, incumbents who can't respond, essential work, and demand that expands rather than shrinks when AI collapses the price?
  2. Can AI actually do the work at software-like margins — and can you keep the surplus rather than competing it away?
  3. Once you've won the work, does anything stop it from leaving — recurring revenue, compounding data, or a regulatory moat that scales with you rather than against you?
TL;DR: We explore why this wave of services is structurally different from the last one, including the technical breakthroughs that finally changed the delivery economics of services, and why we believe so much of the value will accrue to whoever controls that delivery layer. From there, we share the framework our investment team uses to determine whether a given category can support a durable AI-native services firm with outsized value capture.
 

Why this wave of services is different

The cloud era produced many tech enabled services companies that made a real dent in their industries, but these companies largely focused on delivering exceptional customer experiences. LegalZoom digitized the front door of consumer legal services, but the work of law is still priced and delivered the way it always was. Lemonade won over customers but never proved it had fixed the carrier cost structure. Compass became the largest residential brokerage in the country and is still, at its core, a brokerage. Ultimately the cloud era never changed the underlying delivery economics of services, despite its digitization of the customer experience. Long horizon agents flip that equation, increasingly completing hours-long tasks with the only variable cost being that of inference. This has far reached effects beyond gross margin, which we explore in our framework below.

long horizon agents

Previously, Bessemer reported that professional services represent ~13% of US GDP,  making the sector 10x the size of the software industry, and that AI-powered workflows would compete for a meaningful share of those dollars while also enabling work that human labor could never scale to meet. 
 
Since then, the underlying technology has broken in favor of services businesses. As reasoning capabilities of agents inflect and they can complete increasingly long tasks, there are three additional underlying trends that help explain agents’ true autonomous capabilities.
 
  1. The first is multimodality, on both the action side and the input side. Browser agents can now operate portals and legacy systems they were never trained on, with real-desktop benchmark success rising from roughly 12% to within a few points of human performance. Frontier models are now pre-trained and fine-tuned to map pixels directly to precise UI coordinates, replacing the brittle DOM selectors and screen-scraping heuristics of RPA and the error-prone click estimation of 2024-era agents. Browser agents are also now trained end-to-end on verified task completion rather than next-token imitation, which instills self-correction if a click misfires or a page changes. On the voice side, speech-to-speech models are gaining frontier-level reasoning at conversational latency, skipping the transcribe-then-synthesize step entirely. Amperos, an AI biller for healthcare providers, runs on both tailwinds, navigating payer portals and phoning insurers to work denials no clinic can afford to staff. The same progress applies to inputs, since services work arrives as scanned faxes and hundred-page PDFs rather than clean API payloads. EvenUp, an AI demand package writer for plaintiff firms, ingests thousands of pages of scanned medical records and billing tables per case—work that improves in cost and accuracy with every gain in frontier vision models. Unlimited Industries, an AI-native civil engineering firm acting as engineer of record, benefits from the same multimodal input progress, turning the document-heavy front end of site design into stamped design packs.
  2. The second force is that the post-training and eval stack has matured into products sold off the shelf, letting small teams tune systems on their own production data. Much of the frontier's recent gain comes from reinforcement learning on verifiable rewards rather than bigger pretraining, and the scarce input has become the RL environment, a simulated workplace with a checkable reward that labs pay big bucks to acquire. Services firms hold a structural advantage here, because their output is verifiable (i.e., a claim pays or it doesn't), so every delivered unit of work doubles as a training signal, where the firm gets its environment for free (and can keep as alpha) as a byproduct of doing the work. Open frameworks (Trinity-RFT, Thinking Machines' Tinker) extend this to companies that want weights and IP in-house, wiring internal datasets and simulators into their own RL environments. Strala, an AI-native claims administrator, captures every adjuster correction as labeled data and tunes its system until entire classes of misses stop recurring. Crosby, an AI-native law firm for commercial contracts, has its own lawyers design the test sets and ship releases until the models beat their redlines.
  3. An emerging shift worth watching is how open-weight models now trail the closed frontier by roughly one model generation at a fraction of the cost. Few AI services firms have made this part of their story yet, as it’s still early days, but we expect the best firms to defend gross margin by routing routine volume to tuned open models and reserving frontier models for the hardest cases. 

Bessemer's framework for evaluating AI-services opportunities

AI will impact services work unevenly. These markets are changing quickly, and the old guard of TAM and market growth doesn’t sufficiently capture the dynamic nature of AI services. Some markets are structurally advantaged to incorporate AI through a high automation; others possess characteristics of Jevon’s Paradox, where the abundance of the cheap services will massively increase the market size; and others have incumbents whose innovator’s dilemma make it a more compelling market to win share. The best markets spike on all of these. 

To better assess these rapidly developing markets, we created a framework that scores these markets in three buckets: market structure and demand, delivery economics, and defensibility.
 
<iframe src=”https://bvp.com/embed/ai-native-services-criteria-1-11”></iframe> 
 
We think these scorecards are a helpful predictor of where real differentiated enterprise value can be created by using AI to deliver outcomes. To illustrate, we took a handful of examples to demonstrate how they we evaluate the “ripeness” of their respective markets for outcome native firms:
 
shape profiles ai services industry
 

Agent availability doesn’t necessarily determine the market opportunity

In the cloud era, the technology of “tech enabled services” sat in the workflow or experience layer, while the delivery layer stayed human, but in this wave the technology layer impacts the delivery mechanism itself. But the availability of agents is not the same thing as market opportunity. Some markets expand dramatically when the price of the service collapses, because enormous latent demand was always sitting there unserved at human price points.

Others are inelastic, and efficiency gains simply become a knife fight over a shrinking pie. Run two equally automatable markets through the 11 criteria in our framework and you will frequently find that one more easily supports a compounding firm and one supports a well-run services firm whose economics improve but whose ceiling remains the same.
 
One limitation of any scorecard including the ones we share above are that they can only grade markets that already exist. Some of the opportunities we're most excited about don't fit a category at all. These companies replace a stitched-together mess of software licenses, internal headcount, and third-party vendors with a delivered outcome, and because that bundle was never sold as one thing, what they do has no name today. We’ll have more to say about these in a future piece.
 
True agent delivery of professional services will be larger than anything software alone could have claimed. The founders who build enduring companies will pair the tech paradigm shift with the discipline to pick markets whose structure matches its ambition. Our hope is that this gives founders a useful way to pressure-test their services firms and, and if you see it differently, we'd love to hear from you. Email us at ekaplan@bvp.com, lfrost@bvp.com, and mmalik@bvp.com.