The judgement paradox
Surveyors UK
- Risk, Legal & Insurance
- Technology & AI
I watched an AI creator’s breakdown of a recent OpenAI launch video last week. It opened with footage from 1979: a man at MIT, sitting in front of a screen, pointing at a shape and telling it where to go. The lab called it Put That There. The idea was simple. The computer should learn to understand you, not the other way round.
Forty-seven years later, OpenAI put that same clip at the front of its own launch. Not to show off a smarter model. To make a point. The demo that followed showed someone speaking an idea into existence with nothing in their hands. No keyboard. No mouse. No menus.
The creator’s argument was that the skill of operating software, the thing that got a generation of us paid, is dying. Knowing PowerPoint got you paid. Knowing Excel got you paid. That is going. What is left, she said, is judgement. Knowing what needs to get made, why it matters, and whether the finished result is any good.
I think she is right. I have some time researching professional judgement across surveying, audit, medicine and law, and there is a harder version of this argument sitting underneath the comfortable one.
The more essential judgement becomes, the less practice most people get at exercising it. And nobody is building that into their planning yet.
The comfortable version
You have heard some version of this sentence a dozen times this year. AI does the mundane tasks, humans keep the judgement. It shows up in media commentary, in LinkedIn posts, in every “future of work” panel discussion going. It is comfortable because it is true and because it asks nothing of you. You get to keep believing your judgement is safe simply because it belongs to you.
Here is what the research actually says about when that belief holds.
Daniel Kahneman spent much of his career arguing that human intuition is unreliable. Gary Klein spent his studying expert firefighters, nurses and pilots making brilliant split-second calls. In 2009 they did something unusual. They wrote a paper together to work out where they actually disagreed, and found they mostly did not. Confident, accurate intuition only shows up under two conditions. First, the environment has to be stable enough to learn from, with patterns that repeat in ways a person can pick up on. Second, the person needs a long run of practice with fast, honest feedback, so they find out quickly when they got it wrong.
Take either condition away and something different happens. Confident intuition does not become uncertain. It becomes confidently wrong. And feeling sure of yourself, they found, is not evidence that you are right. Subjective confidence is not a reliable signal of judgement accuracy.
This should matter enormously to surveyors, because it explains why your judgement has been trustworthy for as long as it has. Buildings are physical. Defects recur. Markets give feedback through transactions and claims. Surveying is precisely the kind of environment where expert intuition is supposed to work.
It also explains the trap. If AI starts doing the repetitions that built that intuition in the first place, the two conditions Kahneman and Klein identified stop being met without anyone noticing. Not because the environment changed. Because the person stopped getting the practice.
What is actually happening to judgement under AI assistance
In August 2025, the Lancet published the first real-world clinical evidence of AI deskilling. Researchers tracked more than 1,400 colonoscopies performed without AI assistance, across four centres, before and after doctors had spent months working alongside an AI detection tool. The doctors’ own unassisted detection rate fell from 28.4 per cent to 22.4 per cent. A six-point drop, in experienced specialists, after regular exposure to a tool that was supposed to be helping them.
A separate trial published in NEJM AI in 2026 went further. Forty-four physicians completed a twenty-hour course specifically designed to build AI literacy, the kind of training everyone assumes is the answer to this problem. They were then given AI diagnostic suggestions, some of them deliberately wrong. When the AI advice was wrong, diagnostic accuracy fell from 84.9 per cent to 73.3 per cent. Training in how AI works did not stop people from trusting it when it was confidently incorrect.
MIT’s Media Lab found something similar from a different angle. In a study using EEG scanning, participants who wrote essays with an LLM’s help showed markedly weaker neural engagement than those who wrote unaided, and the majority could not quote a single line from an essay they had just produced. Researchers at Wharton ran three large experiments and found people given an AI’s answer took it on most trials, even when it was wrong, a pattern they named cognitive surrender to distinguish it from ordinary, healthy delegation like using a calculator.
Put these together and a shape appears. Judgement is not simply a possession you keep once you have earned it. It behaves more like fitness. It needs load-bearing use to stay reliable, and AI is very good at removing the load.
Here is the uncomfortable extension of this argument, and it is the reason I think the comfortable version is doing more harm than good.
If judgement degrades through disuse, then telling a profession “judgement is what matters now” without also protecting the conditions that build judgement is not reassurance. It is a countdown. Every profession currently repeating that line, and every regulator writing it into standards, is assuming, without checking, that the humans involved will keep exercising real judgement rather than rubber-stamping AI output. The evidence above says that assumption does not hold by default. It has to be defended.
This is also, I think, why every regulated profession in the UK has landed on almost identical language within the space of a year, without any of them coordinating. The Financial Reporting Council told auditors in March 2026 that accountability sits with the audit partner, not the technology, because “you can’t blame it on the box.” ICAEW has treated professional scepticism as an erodible skill for years, long before AI, and requires it to be documented for exactly this reason. RICS’s standard, effective from March this year, defines professional judgement as knowledge, skill, experience and scepticism, and requires a named surveyor to produce a written decision on the reliability of any AI output they rely on. Architecture is heading the same way, with commentators already predicting that RIBA and ARB will follow RICS’s lead on documentation requirements.
That is not four professions independently discovering the same comforting idea. That is four professions independently discovering the same underlying risk, and reaching for the same defence, which is evidence.
Why documentation is the actual answer
If judgement erodes without use, and if confidence is not a reliable sign that judgement is intact, then the only honest response is to build friction back into the system deliberately. Not to slow things down for its own sake, but because the friction is where the judgement happens.
This looks different at different points in a career.
For a working surveyor, it means the written reliability decision the RICS standard already requires should not be treated as a compliance chore bolted on after the fact. It is the moment where judgement is actually exercised and made visible. A decision that exists only in someone’s head is indistinguishable, from the outside, from no decision at all. This is the distinction I have written about before between displayed compliance, a policy sitting in a folder, and demonstrated compliance, a timestamped record of an actual human decision at the point it was needed. My own framework for this, GUARD, treats documentation as the mechanism that turns judgement from a claim into evidence, but the principle matters more than the acronym.
For a junior surveyor, the stakes are longer term and less visible. If AI absorbs the routine inspections and write-ups that used to be how juniors built pattern recognition, on volume, over years, they do not simply learn more slowly. They may never build the tacit library at all. Researchers call this never-skilling, to distinguish it from the deskilling that happens to someone who already had the ability. A firm that lets AI do all of a junior’s early reps is not saving that person time. It may be removing the raw material their judgement was supposed to be built from, and nobody will see the gap until it matters.
For a firm as a whole, it means treating the gap between AI output and human sign-off as something to protect, not streamline away. A practical version of this: before accepting a high-consequence AI-assisted output, form your own view first, in writing, before you look at what the AI produced. This single habit, borrowed from clinical guidance on avoiding automation bias, forces the moment of judgement to happen before the AI’s confidence can shape it.
What this is not
This is not an argument that AI should be resisted, or that surveyors should distrust every tool that helps them. The Kahneman-Klein research is equally clear that in low-validity environments, where the data is thin or the situation genuinely unusual, algorithms can outperform human intuition, and pretending otherwise is its own kind of overconfidence. Automated valuation models already do a large share of mainstream residential valuation, and RICS’s own guidance acknowledges they work well precisely where the assets are common and the data is deep. Nobody sensible wants a surveyor’s ego standing in the way of a tool that is genuinely better in that narrow lane.
The judgement premium is not evenly spread. It concentrates in the messy, heterogeneous, low-data work: the unusual building, the contested valuation, the survey where nothing quite matches the textbook case. That is exactly where AI struggles and human pattern recognition, built over years, still wins. Protecting the conditions that build that pattern recognition is not sentimentality. It is protecting the one asset that cannot be bought off the shelf.
The real stakes
The claim that judgement is becoming the essential skill is correct. But it is not a fact you can simply hold onto. It is a capability you either actively maintain or slowly lose, and the loss does not announce itself.
Nobody wakes up one morning less capable.
They simply stop being tested, and confidence fills the gap where competence used to sit.
The profession that treats judgement as something to defend, with documentation, with deliberate reps for juniors, with a habit of forming a view before consulting the machine, will be the one still exercising real judgement in five years. The profession that treats it as a settled fact, something we simply have because we are human and AI is not, is the one that will discover the gap only when it costs them a client, a claim, or a career.
Judgement is not the safe harbour everyone is describing it as. It is the thing we now have to work to keep.
Litmus – independent insight on AI in surveying
If this raised more questions than it answered, that is the point of Litmus, the independent group I chair looking at what UK surveying is actually doing with AI, not what the marketing says. Our first report, Faster than we can explain: what surveying is actually doing with AI, is free to read at surveyors-uk.com/litmus-ai-in-surveying.
I also run a free monthly AI briefing for surveyors, a straightforward look at what is actually happening across surveying and AI from a practical viewpoint. Book here
Sources
Kahneman, D. and Klein, G. (2009), “Conditions for intuitive expertise: a failure to disagree”, American Psychologist. https://pubmed.ncbi.nlm.nih.gov/19739881/
Budzyń, K. et al. (2025), “Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study”, The Lancet Gastroenterology & Hepatology. https://pubmed.ncbi.nlm.nih.gov/40816301/
Qazi, I.A. et al. (2026), “Automation Bias in Large Language Model-Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy: A Randomized Clinical Trial”, NEJM AI. https://ai.nejm.org/doi/abs/10.1056/AIoa2501001
Kosmyna, N. et al. (2025), “Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task”, arXiv preprint. https://arxiv.org/abs/2506.08872
Shaw, S.D. and Nave, G. (2026), “Thinking – Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender”, Wharton School working paper. https://ssrn.com/abstract=6097646
Financial Reporting Council (2026), “Innovative new guidance supports audit firm adoption of emerging AI technologies”. https://www.frc.org.uk/news-and-events/news/2026/03/innovative-new-guidance-supports-audit-firm-adoption-of-emerging-ai-technologies/
ACCA (2025), “Accountants’ professional judgement critical to success in the age of AI”. https://www.accaglobal.com/gb/en/news/2025/February/AI-risk.html
RICS (2025), Responsible use of artificial intelligence in surveying practice, effective 9 March 2026. https://www.rics.org/content/dam/ricsglobal/documents/standards/Responsible-use-of-artificial-intelligence-in-surveying-practice_September-2025.pdf
RICS (2022), “Automated Valuation Models: A property market perspective”. https://www.rics.org/news-insights/wbef/automated-valuation-models-a-property-market-perspective
Lexology (2025), “New ARB Code – What Architects Need to Know”. https://www.lexology.com/library/detail.aspx?g=f1d62adb-6c9f-4aee-9347-269c4fda303d
RIBA (2025), RIBA AI Report 2025. https://www.riba.org/work/insights-and-resources/ai-report/riba-ai-report-2025/
Superpower Daily (2026), “OpenAI Teases Astra With 1979 Voice-and-Gesture Demo, but Leaves Product Details Out”. https://superpowerdaily.com/posts/openai-teases-astra-with-1979-voice-and-gesture-demo-but-leaves-product-details-out
Nina Young
CEO and Founder - Surveyors UK