Real Talk About AI in Litigation Finance Underwriting

AI Can Transform Litigation Finance Underwriting—But Only When Human Judgment Remains Firmly In the Loop.

August 20, 2026

By Brenna Legaard

First came ChatGPT. Then came the breathless prognostication of the end of lawyering, or lawyers, or at least associates.

Then came the hallucinations, the embarrassment, and the sanctions.

Now the hallucination problem has been wrestled into semi-submission: the models are better and the humans have been warned. We can either happily embrace our new roles as professionally licensed “prompt engineers,” or we can grapple with a much more insidious problem: the way the structural deficiencies in AI outputs can dovetail with human frailties to hollow out your underwriting.

Let’s say you ask an LLM to tell you if a man who shows up at an ER in Mississippi with leg pain has a fracture. That LLM was fed immense swaths of human intellectual output during its “training,” like images of leg fractures, statistics about the incidence of leg fracture, information about Mississippi, narratives about leg pain, perhaps even the 1930 Southern Gothic novel As I Lay Dying by William Faulkner, which is a classic of American literature, concerns a leg fracture, and was not intended for diagnostic use.

That LLM may give you an answer about whether the Mississippi man’s leg is broken based solely on that training data. That answer would obviously be useless. But if you gave it the man’s X-ray, it could use what it learned in training to give you a diagnosis based on that X-ray.

A human doctor would get a medical history and do an examination, which would provide the context: did the man fall off the roof of a church, for example. The LLM won’t. In an edge case it’s important to know that the LLM might be reaching for context in training data, and that context may or may not be helpful and won’t be apparent from the output.

Both of our hypothetical answers—the one based on the Faulkner novel and the one based on the X-ray—will be delivered as clean, well-formatted, typo-free outputs in a format of your choosing, and may or may not include qualifications. To those of us who have been using clues like typos or caveats to judge reliability, both the garbage answer and the potentially useful answer may be indistinguishable, and they will both be clothed in all the signals that implicitly communicate reliability. Crucially, what they will not reveal is what information went into the determination of the outcome. And as we’ll discuss below, that, in a nutshell, is the problem.

Before now, lawyers and their clients haven’t needed to know that much about how legal tech works in order to make use of it. LLMs are different. Here’s what you need to know about that AI output that just landed smack in the middle of in your underwriting process.

What Data Was the Model Trained On?

The first thing to understand is the limitations of the data used to train an LLM. LLMs are trained on what exists: published opinions, filed documents, publicly available data that existed as of the model’s cut-off date. That’s a pretty good corpus for some things. If you want to know how the Eastern District of Texas is likely to handle obviousness at summary judgment, there’s a lot to work with.

But predicting how a case will actually resolve is a different matter entirely. In my work underwriting IP cases for potential funding, for instance, resolution turns on a mosaic of distinct but interwoven issues: that obviousness challenge, in light of a claim construction question, given a particular damages scenario, in the context of a charismatic inventor, tried in N.D. Cal. instead of E.D. Tex, against a given defendant. And here’s the deeper problem: the vast majority of case outcomes are reflected in confidential settlements. There is no dataset capturing how a particular combination of issues plays out in settlement, which means there is no dataset that can be used to train a model—any model—to predict case outcomes in any real-world sense. Any model that purports to do so was trained on something else, and unless you know what that something else was, you can’t understand its limitations.

Going back to our leg patient, if the model was fed X-rays and associated diagnoses during training, it can reliably read an X-ray. Otherwise, it can’t.

What Data Does the Model Have Access To?

In other words, was the model given the X-ray? Was it given case-specific data and told to derive the answer from that data using what it learned in training? If you know what information that model had, you can form an opinion about its adequacy, and you can use that opinion to assess the reliability of the output. Giving a model only information that supports infringement is different from also giving it all available information. Asking it to assess validity based on its “memory” of training data is different from giving it technical background plus the most relevant prior art. 

But either way, an expert human will have information that no model will have. I am an old-school patent lawyer. I have spent the last 25 years pursuing patents, analyzing portfolios, asserting and defending patent suits, and now underwriting potential IP investments for litigation funders and other investors. When I make assessments, I draw on what I’ve learned from all of that experience, including and especially the human dynamics I observed sitting in rooms with clients, going through mediations, watching jurors during trials, and working through issues with judges in hearings. And since legal underwriting is, at its core, the assessment of the likelihood that given humans will reach a given conclusion, those experiences are important. An LLM not only isn’t going to have those experiences, it isn’t going to know what it doesn’t have.

Research on integrating human and algorithmic judgment identifies this as a central issue: algorithms outperform humans in well-defined, data-rich environments, but human experts hold their advantage in complex, variable environments in part because they have access to information that isn’t encoded in any dataset.¹ LLMs can only understand what data exists:  they have known knowns, perhaps known unknowns, and no unknown unknowns whatsoever. Did you tell Claude that the inventor was particularly charismatic? Probably not. And even if you did, Claude wasn’t trained on inventor charisma data because there is no such thing as inventor charisma data. See above.

The weaker the data, the worse the LLM performance. The legal AI hallucination studies have made this vivid in an uncomfortable way. A 2024 peer-reviewed study of 15,000 U.S. federal cases found hallucination rates between 58% and 88% on factual legal questions — and critically, rates were highest for less prominent cases.²

Paradoxically, LLM performance also degrades with too much data. Researchers have established that all models degrade as input length grows, even on tasks that models can handle easily at shorter lengths.³ Cramming an LLM with data may just give you a different problem.

It Is Very Difficult to Identify Weaknesses From the Output Alone, By Design.

This is where the weaknesses of machines and the weaknesses of humans conspire to create blind spots.

When a model generates an assessment of patent validity, it is not starting from a blank slate, and its direction is not neutral. It is algorithmically aimed at producing a result that will satisfy us.⁴ It has also absorbed the patterns of how patent lawyers write about validity — our framings, our rhetorical conventions, our professional biases. It has learned that a strong validity argument looks a certain way, and it will produce output that looks exactly like one. An AI output is chock-full of everything that signals “stand down, I got this” to a human brain.

And here is where that calibration to please dovetails with our human frailties in a way that should give us genuine pause. We are trained to trust things that look like expertise. We are also — and this is the part nobody in the industry wants to say plainly — under relentless time pressure. An AI output may land in a review process that is already looking for reasons to move forward. The model’s confident framing then becomes an unwary underwriter’s confident framing. The unexamined assumptions baked into the prompt become that underwriter’s unexamined assumptions.

Research confirms this dynamic has a name: automation bias: the tendency of decision-makers to over-defer to algorithmic outputs, failing to detect errors even when they have access to information that would allow them to do so.⁵ Even nominally human-in-the-loop designs can degrade decision quality when oversight becomes mechanical rather than genuinely discretionary.

One study of LLMs on legal prediction tasks found something particularly telling: models tend to favor plaintiffs when plaintiffs have more evidence items in the record — not because they reasoned through the merits, but because quantity signals strength in the patterns they learned.⁶ That is not legal reasoning. It is a sophisticated form of the same availability bias we have always struggled with, now running at machine speed and dressed in authoritative language.

So Where Does This Leave Us?

Let me say this first: you would have to pry my LLMs out of my cold, dead hands. I am a better lawyer producing better work product more efficiently because I make informed use of AI. Particularly in data-rich situations, such as analyzing prior art and preparing claim charts, AI is a game changer, and for that reason, if you think you’re not already receiving AI-assisted work product from your law firms, claimants, and counterparties, you are wrong. It is everywhere, for good reason, and understanding what it can and cannot do well is no longer optional.

When the task is well-defined and the data is structured, the machine has real advantages: it doesn’t get tired, it doesn’t anchor on the first argument it reads, and it doesn’t have a relationship with the submitting broker.

But IP underwriting is not a routine, standardized, data-rich environment. It is almost definitionally the opposite. And in that environment, the evidence points in the other direction.

McKinsey’s research on commercial property and casualty underwriting found that companies that mandated black-box model outputs over human judgment entered a vicious cycle: the guidance failed to anticipate actual risk experience, underwriting performance deteriorated, staff lost faith in the models, and—because judgment had been systematically discouraged—underwriting skills atrophied.⁷ The highest-performing underwriters were those with structured, intentional approaches to analyzing exposures, using data-driven tools to supplement rather than replace that judgment.

What that means is that an AI assist in preparing a claim chart is an enormous efficiency booster. But for a human to evaluate that claim chart, she has to understand how the model interpreted the claims, what information it considered, and what else is baked into the reasoning that yielded this work product, including training and context. Perhaps most importantly, the human can’t surrender his or her judgment to the AI.

Three Things Worth Taking From All of This

First, it’s important/crucial to talk about AI use openly. Ask your firms and colleagues what they’re using and how they are using it, and focus the conversation on confirmation bias and gaps, not just hallucinations. The hallucination problem, as I said, is being managed. Gaps and confirmation bias are quieter and harder to detect. How do you do this? Ask them what models or products they’re using, whether they’re feeding their models data, letting them forage the web, or relying on training data, what their limitations are, how they’re trained and prompted, and how they’re interrogating outputs. You don’t need to have answers in mind; you just need to be curious. We’re all still learning; let’s learn together.

Second, respect the mental energy necessary to interrogate a polished, authoritative output. In my experience, the most effective use of AI is by the relentlessly curious, because curiosity provides the mental energy that is necessary to question deeply what presents as reliable. If your underwriters are approaching their work as an exercise in supporting a decision already made, you should know that an LLM can produce well-drafted, professionally credentialed, citation-studded support for almost any proposition, no matter how wrong. You want humans with border collie energy who will chew on those outputs until they’re fully understood. Hire curious people and reward their curiosity.

Third, and more to that point: nurture your own curiosity. One of the genuine pleasures of practicing law in the age of AI is the ease with which you can descend into a rabbit hole and climb back out. Use that. The lawyers who will be most dangerous in this environment — in the best sense — are the ones who use AI to ask harder questions, not just to answer easier ones faster.

‍ ‍

¹ Alur, Laine, Li, Shung, Raghavan & Shah, "Integrating Expert Judgment and Algorithmic Decision Making: An Indistinguishability Framework," MIT/Yale (2024) — https://arxiv.org/abs/2410.08783

² Dahl, Magesh, Suzgun & Ho, "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models," Journal of Legal Analysis, Vol. 16 (2024) — https://academic.oup.com/jla/article/16/1/64/7699227

³ Hong, Troynikov & Huber, "Context Rot: How Increasing Input Tokens Impacts LLM Performance," Chroma Research (2025) — https://research.trychroma.com/context-rot

⁴ Sharma et al., "Towards Understanding Sycophancy in Language Models," ICLR (2024) — https://arxiv.org/abs/2310.13548

⁵ Parasuraman & Manzey, "Complacency and Bias in Human Use of Automation: An Attentional Integration," Human Factors, Vol. 52(3) (2010) — https://doi.org/10.1177/0018720810376055; see also Green & Chen, "The Principles and Limits of Algorithm-in-the-Loop Decision Making," Proc. ACM on Human-Computer Interaction, Vol. 3, CSCW (2019) — https://doi.org/10.1145/3359152

⁶ "Legal Fact Prediction: The Missing Piece in Legal Judgment Prediction," arXiv (2024) — https://arxiv.org/abs/2409.07055

⁷ Chester, Ebert, Kauderer & McNeill, "From Art to Science: The Future of Underwriting in Commercial P&C Insurance," McKinsey & Company (2019) — https://www.mckinsey.com/industries/financial-services/our-insights/from-art-to-science-the-future-of-underwriting-in-commercial-p-and-c-insurance

Next
Next

When One Case in the Portfolio Changes the Disclosure Calculus