Legal skills in the genAI age - part 1: AI eats what it can measure
After I spoke at an annual conference about AI x Law a few months ago, lawyers and law students started reaching out with different versions of the same question. Is it still worth pursuing a legal career when every headline promises that AI will devalue legal work and replace the people who do it, especially the juniors? One very anxious student explained that she was extra worried her entire career would be snuffed out before it even began, because why should a 1L grind through case law when the cutting-edge models can clearly dissect, analyse, and summarise lengthy, complex legal texts in mere minutes?
I am by no means an expert in either AI or law, but I have been lucky enough to work at the forefront of both fields. On the law front, I have done well in law school, and seen litigation and solicitor work done at big law as well as the judiciary first hand. On the tech front, for the past year and a half I have been thrown into the deep end of AI/ML, wearing again my STEM hat to get hands on as much as I can about the science, math, and engineering behind the different approaches to building AI, to understand the VC and business landscape surrounding this technology as much as possible, and to build agent workflows to improve my personal and startup work. So, hopefully, my experience allows me to provide some clarity during this time of uncertainty.
As the title suggests, this is part one of a multi-part series. I have not organised my thoughts well enough to promise an exact count, but the map I am writing from has four stops. The first two parts concern existing human skills: what AI actually takes, and what that means for someone just entering the profession. The last two concern the skills of using AI itself: what genuine ability with these tools looks like beyond prompting, and how working alongside them reshapes the way our own abilities form.
Throughout, I will keep returning to skills, which I define as learned abilities to perform consistently a given set of tasks to a given standard while no longer having to calibrate each step of the process. The parts of a job that run without recalibration are, as we will see, exactly the parts AI absorbs first, and the parts that never stop demanding calibration are where everything else in this series sits. Another thread that runs through all four parts is the commodification of intelligence. As machine intelligence becomes plentiful and cheap, the scarce resource stops being how smart anyone is. It seems to me that what will separate people in this age is their relationship to mental effort, the willingness to keep thinking hard when a machine sings its siren's song of doing all the thinking for you. Intelligence is becoming a commodity, but the awareness and appetite to use your own when you should is not.
One final disclaimer before we begin. Everything I write in this series could age like milk if model capabilities jump again by leaps and bounds, like the jump from GPT-3 to GPT-4. While I have confidence this article will age more like wine, at least in the near future, please read this first part and its siblings as field reports about the right here and now rather than divinations about the AI and law future.
Part 1: AI eats what it can measure
Since I am prompted to write by those who came to me in fear, I start by analysing this fear. The fear tends to come in three versions: 1) law firms are dinosaurs about to be disrupted, so the legal industry's conservatism makes it especially vulnerable to technological change; 2) junior lawyers are about to be automated away, so the apprenticeship ladder will collapse before new lawyers can climb it; 3) lawyers as a profession are about to be completely upended, if not wholly replaced, so law is becoming a dying industry rather than a durable career.
No doubt each version contains a kernel of truth. The legal industry is risk-averse. Junior legal jobs have always been competitive, fragile, fluid, and unevenly distributed amongst the graduates. Plus, no matter how much one wants to kick and scream, legal tasks are plainly more automatable than we would have liked to admit even one year ago.
But these truths become misleading when they are inflated into a single story of inevitable collapse, and this is why I smell paltering. It is not false that law as an industry is not very tech-forward, that junior lawyers are exposed, or that legal work will change. The false conclusion is that slowness (not a good thing, but still) automatically leads to extinction, exposure to uselessness, and change to replacement.
It is also worth noticing that each of these half-truths has a seller behind it that benefits from the fear buy-in. Legal tech vendors benefit from the dinosaur story because it helps close enterprise deals. Consultants benefit from the obsolescence story because it sells transformation roadmaps. And the content economy benefits from the replacement story because doom travels further than nuance. The replacement story has recently acquired a seemingly authoritative number, a figure suggesting roughly 88% of legal tasks are automatable. That figure rests on a methodology its own authors called somewhat arbitrary, rated by non-occupationally-diverse annotators, and building on a threshold the underlying research itself acknowledged was contestable (the estimate traces to Anthropic's labour-market impact research, built on OpenAI's "GPTs are GPTs" methodology). Reading that number as a labour-market forecast, rather than as a sales instrument, is itself a kind of error.
Again, none of the above means the fear is baseless. Big changes are indeed happening both in and outside of the law. Stanford's Digital Economy Lab finds that entry-level employment in AI-exposed occupations fell roughly 13% since late 2022, while employment for experienced workers in the same occupations held steady or grew (via Jeff Raikes, "AI Is Capturing Cognition," Fortune/MSN, 2026, citing Stanford's Digital Economy Lab). More than half of business leaders surveyed by KPMG expect to reshape entry-level recruiting within the year (also cited in the same Raikes piece). Two grains of salt belong beside these numbers. Economy-wide studies, including recent European Central Bank research, still find AI's measured effect on employment and wages muted so far, so the squeeze is concentrated rather than general. And law carries a structural buffer most exposed industries lack, a mandatory licensing pipeline that stands between the technology and the seat, and no human practises law without passing through it, however capable the software becomes.
The pattern nonetheless deserves attention. The pressure lands on the young and inexperienced, while the protected tier is precisely the opposite cohort who has accumulated judgment through experience. I find this asymmetry to be the most interesting and important key that helps us examine the current landscape of AI and work.
To understand why the experience tier holds, it helps to have a filter before we reach the historical analogy. I think of it this way: AI hits hardest where output is cheaply verifiable by anyone, binary rule checks, document retrieval, pattern-matching against established legal texts, anything a second-year associate could mark right or wrong in seconds. Where quality is context-dependent, where the persuasiveness of an argument turns on a specific judge's dispositions or a client's unstated priorities, AI loses traction quickly. And where the very training data is inaccessible, privileged client files, private negotiation histories, local court intuitions not in any public corpus, the limitation is structural, not temporary. Damien Charlotin summarizes these two tests that together sort the automatable from the irreducible in any knowledge domain, with one focused on judgeability and the other on data accessibility (Damien Charlotin, "Artificial Authority" lecture series, HEC Paris). Notice what the first test quietly implies for drafting. Unlike code, which announces its own failures by refusing to compile or crashing a test suite, words cannot be run. A legal draft can now be generated rapidly and cheaply, but it cannot be verified at the same speed, and that gap between cheap generation and costly verification will follow us through the rest of this series.
History offers us a control group. In 2016, Professor Geoffrey Hinton declared we should stop training radiologists because image recognition AI would outperform them within five years. Nine years later, 1,208 residency positions were filled in 2025, up 4% year over year, vacancy rates remained at all-time highs, and average pay reached $520,000, up 48% since 2015. Radiology had in fact already run this experiment once, in its pre-AI automation wave, and the result explains why. When Vancouver General Hospital digitised its imaging workflow, radiologist productivity rose 27% on plain radiography and 98% on CT within a year, and US imaging utilisation rose 60% between 2000 and 2008. The efficiency did not shrink the profession. Faster scans created new diagnostic uses faster than old ones were automated, and whole-body trauma CTs went from exception to routine. Economists call this a Jevons paradox. When you make a resource cheaper, its consumption expands and consequently drives demand. A study further completes the picture, finding radiologists spend only 36% of their day on direct image interpretation, and giving the rest to oversight, clinical communication, teaching, and protocol review, exactly the meta and relational activities that fail both tests of judgeability and data accessibility. "The better the machines, the busier radiologists have become." (radiology data throughout this paragraph via Deena Mousa, "AI Isn't Replacing Radiologists," Works in Progress, September 2025)
I personally believe a legal Jevons paradox is already running in parallel. When AI makes legal analysis cheaper and faster, demand for that analysis tends to expand, with more scenarios getting stress-tested, more thorough due diligence becoming affordable, more clients who didn't even know they had a legal issue now seeking proper advice after learning more from chatbots. The productivity gain does not reduce the total volume of legal work, but instead shifts the concentration of tasks from one tier into the next, work that was probably previously priced out (Charlotin's Jevons-paradox argument, same lecture series). The early workplace evidence points the same way at the level of the individual worker. Studies tracking what happens after AI adoption keep finding that work becomes more intense rather than lighter, as adopters take back tasks they used to hand off, squeeze work into spare moments, and reinvest every saved hour into new output, all while their stretches of focused, uninterrupted work shrink. Nobody I know who works seriously with these tools reports a quieter day. The standard rejoinder is the fate of the horse, which the engine made permanently redundant. But I think the analogy fails on its own terms. A horse could do almost nothing an engine could not do better, whereas human lawyers keep specialising into exactly the work the machines handle worst (Charlotin again).
In law, I have yet to meet a lawyer whose day is mostly the benchmarked task either. Plus, a lawyer's job is probably more soft skills heavy than that of a radiologist, and the percentage of soft skill work increases as seniority grows. Be it client management, negotiation strategy, witness preparation, credibility assessment, oral advocacy, emotional intelligence, or business development, all of these require meta, relational, and human-interactive acumen in the physical world. Binary checks, retrieval, and short question answering can be easily measured and machine-optimised. Advising, persuading, managing the file, and absorbing responsibility remain grounded in relationships, judgment, and accountability, and these holistic tasks, with their ultra long horizon planning, complex context selection and switching, and emotional intelligence exercised with humans in the physical world, are ones no current benchmark comes close to measuring. Stanford's AI Index makes the measurement point bluntly, explaining that knowing a model scores 75% on a legal reasoning benchmark tells us little about how well it would fit into a law practice's activities (Stanford HAI, AI Index Report 2026).
I owe the readers the strongest evidence against my position before going further. In a recent Stanford experiment, sixteen contracts professors compared, blinded, the answers that human instructors and frontier models gave to first-year students' questions, and they preferred the model's answer roughly three times out of four (Alejandro Salinas et al., Stanford University, 2026, Law Professors Prefer AI Over Peer Answers). These were judgment-rich hypotheticals without answer keys, and the machine won anyway. But taking a closer look at the task shows that it was a bounded, single-turn tutoring exchange against a shared professional standard, with the relevant material fully public, no client, no record, and no consequence for being wrong. I would argue that it is the type of judgment work that passes both the judgeability and data accessibility tests. The study's own framing even concedes that most legal hypotheticals admit no single correct answer, which is precisely why the evaluation had to be a preference vote between candidates rather than a score against ground truth. I think the key takeaway from the experiment is that the measurable frontier now reaches deeper into cognition than most lawyers assumed, which should keep all of us on our toes for wanting to learn and experiment more with the technology. What it does not show is that there still is a great deal of work that is very difficult to measure.
Understanding this boundary also answers why more senior workers become more valuable while the measurable slice automates. For the legal industry, US judge Oliver Wendell Holmes gave the answer in 1881: "The life of the law has not been logic: it has been experience" (Oliver Wendell Holmes Jr., The Common Law, 1881). Large language models by default optimise toward the most probable answer, but the law usually does not have that one generalizable right answer. What a limitation defence is worth in front of a particular judge, what a specific client actually means when instructing you to "be aggressive," which of three plausible readings of a clause will the counterparty agree to and for what reason, none of this can be easily reducible to pattern retrieval over public data. The training data that could teach a model this kind of nuanced and situational decision-making is simply lacking, and mostly does not exist in collectable form at all with current data collection technology. Moreover, the heavily context-driven and human, subjective nature of law means persuasion and advisory in courthouses and negotiation rooms remain hard to automate with current technology. Hence, senior lawyers on average hold more of the valuable soft skills that resist automation, and thus they can survive and even thrive in the age of genAI.
The pattern beneath all of this is measurability. AI eats what it, and by it I mean also its creators, can (cheaply) measure and verify, and it stalls where quality ties more into context, judgment, and data not available publicly. This is why I think the skills we call soft are never really soft in terms of importance or difficulty. They are instead soft and shapeless to take a comfortable shape to be measured, analysed, and reproduced by machines.
So much for the senior, then, but the very anxious student from the opening of this series is still waiting for an answer. If the measurable work levels out, and AI lifts the weakest performers the most, why should a 1L grind through case law at all? That is the question I turn to in Part 2.
Stay tuned.
