Answer

How should health systems evaluate AI vendors for revenue cycle?

Alex Oey Updated August 14, 2026

Short answer

Health systems should evaluate AI revenue cycle vendors against the administrative cost problem they are meant to solve, not against generic productivity claims. Two ideas from Sharlene Seidman, VP of Revenue Cycle at Johns Hopkins Medicine, anchor the answer: with decreasing reimbursement and unsustainable administrative burden between payers and providers, AI's value has to be measured by how much cost it removes; and vendors have to be vetted for proof of performance and shared values before they can function as partners. Building on that, GetixHealth's operator rubric adds four practical metrics buyers can pressure-test against their own data: cost per denial worked, touches per claim, hours saved per FTE, and - for denial and appeal use cases - appeal overturn rate.

Why It Matters

 Most AI vendor evaluations start with the wrong question. They ask "how much faster can your tool work?" when the operational question is "how much cost does your tool remove from the claim, the denial, or the authorization?"

Seidman put the pressure in plain terms: "Because in order for us to survive, we are going to have to figure out ways to reduce the cost and the administrative burdens that we're working with between payers and providers. It's not something that's sustainable."

That reframing matters because the current cost structure between payers and providers, as Seidman said, is not sustainable. In our experience working with providers, AI pilots that improve individual task speed without reducing touches per claim, cost per denial, or FTE demand rarely change financial performance in a material way - and rarely survive the second budget cycle.

Key Takeaways

  • Cost per denial worked is a truer measure than productivity. If an AI tool speeds up denial work but does not reduce the fully loaded cost to resolve each denial, it is productivity theater. Ask vendors to model the before-and-after cost per denial worked using your own volumes across specific denial classes (clinical, technical, coding, authorization).
  • Touches per claim reveals whether AI is removing labor or just relocating it. A claim that used to require six manual touches and now requires five with AI assistance is unlikely to change unit cost materially. Push vendors to demonstrate meaningful reductions and, where possible, zero-touch scenarios for defined segments across eligibility, claims editing, and coding support.
  • Hours saved per FTE must translate into redeployed capacity. Hours "saved" that never leave the P&L are not savings. In our experience, buyers should require the vendor to help measure how freed capacity is redeployed - higher-value work, reduced overtime, avoided hiring on prior auth or patient access teams.
  • Appeal overturn rate is the outcome metric that matters for denial and appeal tools. For AI applied to clinical denials and appeals specifically, the question is whether the tool actually wins more money back, not just whether it drafts letters faster. This metric does not apply cleanly to every AI category (eligibility, coding assist, patient access), so scope it to the use case.
  • Proof and shared values are legitimate evaluation criteria. Seidman was direct on how she thinks about vendor selection: "The vendor management side of it is very complicated because our industry has exploded and there are a lot of competing vendors out there. So to partner with the right people, you first have to vet them really well to make sure that what they're pitching to you is really proven and it's going to work for your organization. And the other thing that's really nice to find is that they share your same values. So I like to make sure that we partner with vendors that match our core value."
  • AI fluency will be table stakes for RCM leadership. Seidman's phrase was blunt: "That's going to be table stakes. There's no doubt that everyone who's in a revenue cycle leadership role in 10 years is going to have to understand the technology really well and how it can benefit all of us." Vendor evaluations should be led by leaders who can pressure-test claims, not delegated to buyers who cannot.

Expert Perspective

 Seidman ties the future of RCM leadership directly to the reimbursement environment. She frames the question as budget-critical rather than exploratory: administrative burden between payers and providers has to come down, and technology is one of the few remaining levers. In her words, future leaders "are going to have to understand not just the AI piece of it, but how it's going to be useful to drive down the cost."

Building on Seidman's point about unsustainable administrative cost, buyers can pressure-test vendors with concrete operator questions:

  • What is our current cost per denial worked in this book of business, and what will it be with your tool - by denial class?
  • How many manual touches does an average claim require today, and which of those touches will your tool eliminate, not augment?
  • If denials staff currently spend, say, 18 minutes per worked denial, what does that fall to with your tool, on which denial types, and how did you validate it?
  • For denial and appeal use cases, what appeal overturn rate has your tool produced in organizations with a payer mix similar to ours?
  • How do you define success in month three, month six, and month twelve?

Seidman's second filter - shared values - is not a soft criterion. In our experience working with providers, values alignment shows up in how a vendor behaves during a missed SLA, a payer policy change, or a compliance question. Moments that define whether a multi-year partnership holds together. Reference calls are more useful when they probe those moments than when they test general satisfaction.

A common mistake we see: organizations evaluate AI vendors the way they evaluate software features and are surprised when the tool does not change financial performance. The rubric above - grounded in Seidman's cost-pressure standard and applied across denials, prior auth, claims editing, eligibility, patient access, and coding support - forces the right framing into the evaluation.

Related episodes & articles