Owning the outcome: Bessemer's AI-Native Services evaluation framework
How Bessemer evaluates which services markets are most ripe for AI disruption.
In the cloud era, software won by becoming a system of engagement or a system of record. These models supported businesses to more efficiently do their work; but now, AI can deliver the work itself and the outcome is the product.
Professional services firms have always been measured on outcomes. A law firm doesn't sell you help understanding a contract; it sells you a redlined one. A TPA doesn't sell you claims software; it sells you a closed claim. That much hasn't changed and isn't going to.
What's changing is what's doing the work behind that promise. Now it can be a machine instead of a person, sitting inside the delivery of the service itself rather than just its front end. Technology has always touched services at the experience layer (e.g., better funnels, better onboarding, better interfaces) but never the delivery layer. AI changes delivery itself, and that creates a wildly more attractive business.
The availability of agents to deliver services doesn't necessarily mean that every services market transforms into an autonomous, high-margin market dominated by agents. So, to identify investable opportunities in AI-native services, we're exploring three core questions:
- Is this a market that's structurally ready to be taken given factors such as fragmented supply, incumbents who can't respond, essential work, and demand that expands rather than shrinks when AI collapses the price? Or is it a market that doesn't exist yet, where AI makes a completely new service category viable that previously wasn't possible or economical?
- Can AI actually do the work at software-like margins and can you keep the surplus rather than competing it away?
- Once you've won the work, does anything stop it from leaving i.e., is there recurring revenue, compounding data, or a regulatory moat that scales with you rather than against you?
| TL;DR: We explore why this wave of services is structurally different from the last one, including the technical breakthroughs that finally changed the delivery economics of services, and why we believe so much of the value will accrue to whoever controls that delivery layer. From there, we share the framework our investment team uses to determine whether a given category can support a durable AI-native services firm with outsized value capture. |
Why this wave of services is different
SaaS produced many tech enabled services companies that made a real dent in their industries, but these companies largely focused on delivering exceptional customer experiences. LegalZoom digitized the front door of consumer legal services, but the work of law is still priced and delivered the way it always was. Compass became the largest residential brokerage in the country and is still, at its core, a brokerage. Ultimately cloud never changed the underlying delivery economics of services, despite its digitization of the customer experience.
Long horizon agents flip that equation, increasingly completing hours-long tasks with the only variable cost being that of inference. This has far reached effects beyond gross margin, which we explore in our framework below.
Since then, the underlying technology has broken in favor of services businesses. As reasoning capabilities of agents inflect and they can complete increasingly long tasks, there are three additional underlying trends that help explain agents’ true autonomous capabilities.
- The first is that agents can now work through the same interfaces humans do: screens, phones, and paperwork. On screens, browser agents operate portals and legacy systems they were never trained on, with desktop benchmark success rising from ~35% fourteen months ago to ~85% today, surpassing the 72% human baseline. Frontier models are trained and post-trained to map pixels directly to precise UI coordinates, replacing the brittle DOM selectors of RPA, and trained end-to-end on verified task completion rather than next-token imitation, which instills self-correction when a click misfires or a page changes. On phones, speech-to-speech models are gaining frontier-level reasoning at conversational latency, skipping the transcribe-then-synthesize step entirely. Amperos, an AI biller for healthcare providers, runs on both browser and voice tailwinds, navigating payer portals and phoning insurers to work denials no clinic can afford to staff. Further, on paperwork, document reading has moved from OCR that simply transcribes characters to vision-language models that understand the whole page - parsing tables, charts, and handwriting in one pass, now at a few dollars per thousand pages. EvenUp, an AI demand package writer for plaintiff firms, ingests thousands of pages of scanned medical records and billing tables per case - work whose cost and accuracy rides directly from this curve. Unlimited Industries, an AI-native civil engineering firm acting as engineer of record, does the same for the document-heavy front end of site design, converting survey, soil, and hydrology PDFs into stamped design packs.
- The second force is that the post-training and eval stack has matured into products sold off the shelf, letting small teams tune systems on their own production data. Much of the frontier's recent gain comes from reinforcement learning on verifiable rewards rather than bigger pretraining, and the scarce input has become the RL environment, a simulated workplace with a checkable reward that labs pay big bucks to acquire. Services firms hold a structural advantage here, because their output is verifiable (i.e., a claim pays or it doesn't), so every delivered unit of work doubles as a training signal, where the firm gets its environment for free (and can keep as alpha) as a byproduct of doing the work. Open frameworks (Trinity-RFT, Thinking Machines' Tinker) extend this to companies that want weights and IP in-house, wiring internal datasets and simulators into their own RL environments. Strala, an AI-native claims administrator, captures every adjuster correction as labeled data and tunes its system until entire classes of misses stop recurring. Crosby, an AI-native law firm for commercial contracts, has its own lawyers design the test sets and ship releases until the models beat their redlines.
- An emerging shift worth watching is how open-weight models now trail the closed frontier by roughly one model generation at a fraction of the cost. Few AI services firms have made this part of their story yet, as it’s still early days, but we expect the best firms to defend gross margin by routing routine volume to tuned open models and reserving frontier models for the hardest cases.
Bessemer's framework for evaluating AI-services opportunities
AI will impact services work unevenly. These markets are changing quickly, and the old guard of TAM and market growth doesn’t sufficiently capture the dynamic nature of AI services. Some markets are structurally advantaged to incorporate AI through a high automation; others possess characteristics of Jevon’s Paradox, where the abundance of the cheap services will massively increase the market size; and others have incumbents whose innovator’s dilemma make it a more compelling market to win share. The best markets spike on all of these.
To better assess these rapidly developing markets, we created a framework that scores these markets in three buckets: market structure and demand, delivery economics, and defensibility.
We think these scorecards are a helpful predictor of where real differentiated enterprise value can be created by using AI to deliver outcomes. To illustrate, we took a handful of examples to demonstrate how they we evaluate the “ripeness” of their respective markets for outcome native firms:

Agent availability doesn’t necessarily determine the market opportunity
The technology behind "tech enabled services" sat in the workflow or experience layer while the delivery layer stayed human. In today's wave, the technology reaches delivery itself, which makes right now one of the most exciting moments ever to build a services firm from scratch. But some opportunities carry better structural advantages than others, and that is the driving force behind this framework. Market structure decides whether there's an opening, delivery economics decide whether value can be captured, and defensibility decides whether the moat compounds with every unit of work delivered. The markets we get most excited about spike on all three.
We also recognize the framework has limits. It mainly grades markets that already exist, and some of the most exciting opportunities are creating their own category. The status quo might be a messy stitch of outsourced third parties, internal hires, and software licenses that have never been sold as one thing, or the problem that a service solves might not have existed a few years ago at all. And even for a market that scores well on the scorecard, a strong market still has to be won. Selling work end-to-end sets a very high bar. Meeting that bar requires flawless execution, and exceeding it requires an added element of craftsmanship across technology, product, and brand, including in how context is gathered from clients, how the delivery of work is communicated, and how the final product is presented. The market determines what's possible, but the team determines how to do it.
True agentic delivery of professional services will be larger than anything software alone could have claimed. The founders who build enduring companies will pair the technology paradigm shift with the discipline to pick markets whose structure matches their ambition, and the craft to deliver a service clients love. Our hope is that this gives founders a useful way to pressure-test their services firms, and if you see it differently, we'd love to hear from you. Email us at ekaplan@bvp.com, lfrost@bvp.com, and mmalik@bvp.com.







