Doctoral research | Trust in human-AI collaboration
Your AI works.
Your people still get it wrong.
Five years of research on one question: what makes people trust an AI the right amount. Not more, not less. The answer turns out to be uncomfortable for anyone rolling out AI right now, because the thing that drives trust is not how good the system is. It is how good it looks.
Engineering Trust-based Human-AI Collaboration
The problem worth your attention
Adoption is not the finish line.
Most AI programmes measure whether people use the system. That is the wrong finish line. What determines performance is whether people’s confidence tracks what the system can actually do. Researchers call this trust calibration, and it fails in both directions.
Capable systems get followed off a cliff
In one study, recruiters given a strong AI followed it uncritically and the combined team performed worse than the AI alone. Adoption went up. Outcomes went down.
Good systems get abandoned after five bad calls
As few as five confidently wrong predictions can trigger a lasting collapse in trust. Worse, people forgive an AI less than a human for the identical mistake, even when the AI is more accurate overall.
The stakes are highest where decisions are fast
In production and operations, a rescheduling call is made in minutes and cascades into delivery dates, utilisation, and cost. There is no time to audit the recommendation. Trust has to be right in advance.
The evidence base
Three studies, one narrowing question.
The research moves from a broad map of everything that shapes trust, to what actually matters in industrial decision-making, to a controlled test of whether it can be deliberately influenced.
What actually drives trust in intelligent systems
179 publications screened from 3,358 records across robotics, automation, information systems, and human factors.
479 separate drivers, sorted into the environment, the person, the system, and a fourth group earlier models had missed: the interaction itself. The honest conclusion is that the list is too long to design from. The useful question is which few drivers do the work.
Which drivers matter in real operations
Expert interviews in production management, followed by a survey and structural equation modelling.
One driver dominates everything else: how capable the AI is perceived to be. Comprehensibility matters, but only as a fitting second. Objective factors such as task predictability and the cost of an error never reach trust directly, they arrive through perception. One driver surfaced that no earlier trust model had included: digital affinity, the personal pull toward working with new technology.
Do explanations actually help
253 manufacturing professionals and engineers scheduling a shop floor. A black-box AI with short written rationales against a fully transparent rule, with identical recommendations in both conditions.
Short rationales made people follow the AI more. They did not raise considered trust directly, and they did not improve understanding. Trust moved only downstream, by making the system look more capable. The explanations worked as a signal of competence, not as an explanation.
What the research establishes
Four findings that change how you roll out AI.
These are not opinions about AI. They are measured effects, replicated across an observational study and a controlled experiment using the same instruments.
The thesis condenses them into three insights: perceived capability is the mechanism that carries trust, trust and reliance have to be treated as two different things, and calibration is a property of the arrangement rather than of the individual user.
Trust is bought with appearances, not accuracy
Perceived capability is by far the strongest driver of trust in AI decision support. It held across three different model specifications and was then replicated under experimental control. Everything else, including how risky the decision is and how much an error costs, works through that perception rather than around it.
The uncomfortable corollary is that making a system look capable already moves trust. Vendors know this. Your people cannot tell the difference from the interface alone.
Study II, replicated under experimental control in Study IIIExplanations persuade more than they explain
The standard assumption in explainable AI is that showing the reasoning improves understanding, and that understanding produces appropriate trust. In an early-interaction setting that assumption did not hold. Brief rationales lifted the sense that the system was capable while barely touching comprehension.
Trust did rise, but only downstream of that impression, never from the explanation itself. Explanations are therefore an adoption lever. Treating them as a safety mechanism is a category error.
Study IIIBehaviour and belief come apart
People acted on the AI’s advice more often without their considered trust rising to match. That gap is the dangerous one. It looks like successful adoption in every dashboard you have, while the critical judgment that was supposed to catch the bad recommendation quietly disengages.
Evidence from outside the factory points the same way. In a chess study, people relied more on a partner they knew was an AI while reporting more trust in one they believed was human, and the weaker performers were the ones most swayed by who the partner was.
For a single decision, behaviour is what matters. Over months of decisions, the underlying belief decides whether people stay alert or drift into compliance.
Study III, consistent with wider literatureExplanations help experts more than novices
The expected pattern was that explanations help most where knowledge is thinnest. The data pointed the other way. The trust benefit was smaller for people who found the task hard and for those without domain experience, and larger for experts, who read the same cues as credible signals of competence. Reliance was not moderated at all, which suggests the novices were not simply following blindly.
These boundary effects sit below the study’s own sensitivity threshold and should be read as directional rather than settled. The robust part is the asymmetry: trust is fragile and context-dependent, while reliance is stubborn.
Study III, moderation analysisThe meta-finding. Trust calibration is a lever and a vulnerability at the same time. The perception pathway shows exactly where design can intervene. The gap between behaviour and belief shows exactly where that intervention can backfire. Which means the thing you need to design is not the model. It is the arrangement around it: who decides what, who checks what, and what the interface is allowed to imply.
That arrangement is the third insight and the one that survives the move from research to practice. Calibration is not a property of the model, and it cannot be left to the individual user. The thesis proposes building a discipline around it under the working title Human-AI Collaboration Engineering. At this stage that names the unit of design and its first focus. The methods and reusable patterns still have to be built and tested in the field.
The stress test
Generative AI makes all of this harder.
The studies examined task-specific AI. Large language models change five conditions at once, and each one attacks a mechanism the research identified. Hover or tap a card for what it does to your organization.
The competence boundary disappears
One system drafts, summarises, codes, and plans, and is excellent at some of it and unreliable at the rest.
With narrow AI, people learned where the system was strong. A jagged capability profile removes that learning curve entirely.
Your people cannot calibrate against a target that has no stable edges. Expect confident use in exactly the areas where the model is weakest.
Attacks Finding 01Fluency reads as competence
Confident, articulate, human-sounding output regardless of whether the reasoning underneath is sound.
If perceived capability is the gateway to trust, then a system that manufactures the impression of capability has direct access to that gateway.
The Clever Hans problem returns in linguistic form. The performance is convincing and the mechanism behind it stays hidden.
Attacks Finding 01The system writes its own explanations
Explanation stops being something you design and becomes something the model generates.
Self-generated rationales can be persuasive, tailored, or simply disconnected from what the model actually did. Prompt injection and data poisoning let third parties steer them.
You lose your independent reference point. Routing validated content through a model only stacks one opaque layer on another.
Attacks Finding 02The system agrees with you
Sycophantic responses raise the user’s sense of being right, their trust, and their intention to come back.
The interaction becomes a mirror that confirms the position you arrived with, removing the friction that recalibration depends on. It adapts to your language and preferences in real time.
Modelling work shows this can escalate false confidence even for a perfectly rational user. It is not a problem you can train out of people.
Attacks Finding 03The judgment doing the checking erodes
Heavy AI use correlates with self-overestimation, cognitive offloading, and weaker critical thinking.
More strikingly, higher AI literacy correlates with higher overestimation of one’s own performance, not lower. Delegation makes it worse, since people commission more questionable decisions through goals and examples than through explicit instructions.
Inflated signals and self-written explanations arrive at users whose capacity to evaluate them is quietly declining. This is the compounding one.
Attacks Finding 03A second-order problem
The thing to be judged gets harder to judge, while the ability to judge it gets weaker.
Neither half can be solved by better models. Both are properties of how the work is arranged around the model.
What stays human is the class of decisions that cannot be computed at all, where values conflict, evidence runs out, and responsibility cannot be delegated without the outcome losing its legitimacy.
Where design has to step inFrom evidence to practice
Five strategies for AI people can judge.
Existing guidance locates the fix in model design, organizational practice, and regulation. The research adds a fourth layer that is usually skipped: the design of the interaction itself.
Their evidential weight varies. Some translate a mechanism measured in the studies. Others extend it with the wider literature. The distinction is stated rather than hidden.
Treat opacity as a cost, not a default
For a wide range of tasks, interpretable models match black-box accuracy. Use the simplest model that clears the performance bar, and make anyone proposing an opaque one show what the extra accuracy is worth against the calibration risk it introduces.
Engineer honesty into how the system presents itself
If perceived capability is the gateway to trust, the system must not oversell itself at that gateway. Communicate confidence boundaries, surface uncertainty, and flag misbehaviour. Necessary, but not sufficient, since an honest system can still be opaque and still invite people to read more into it than is there.
Build in cognitive distance
Cognitive distance means deliberate reflection on the advice before accepting it. Where opacity is unavoidable, that reflection has to be built into the interaction, because you cannot delegate vigilance to the user and expect it to hold. The risk is sharpest where AI shapes what people notice before they have started thinking.
Practical moves are small and cheap: drop the first-person voice, turn assertions into questions, and require a position before the AI offers one.
Put verification in the process, not in the person
Self-calibration varies with expertise and may itself degrade with AI use. Use independent checks, escalation routines, and role-based oversight. Notably, control mechanisms tend to increase trust and adoption rather than slow them down. Keep the high-stakes, long-horizon, hard-to-reverse decisions with humans.
Use regulation as a backstop, not a strategy
Regulation shapes market conditions, for instance by forcing disclosure that deflates inflated capability claims. It can produce trustworthy systems. It cannot produce calibrated trust, because compliance does not make users infer capability correctly.
Where to start
Six questions worth asking this quarter.
If your AI initiative cannot answer these, calibration is being left to chance. That is not a technology gap. It is a design gap.
Can your users state where the system is unreliable? If they cannot name a weakness, they have not calibrated. They have simply accepted.
Are you measuring adoption, or appropriate override? Usage rates rise in both healthy and unhealthy scenarios. The rate at which people correctly reject a bad recommendation does not.
Do your explanations survive contact with a wrong answer? An explanation that sounds equally convincing when the output is wrong is a persuasion device, not a calibration aid.
Who is accountable for a decision the human could not realistically have checked? Holding people responsible for outcomes they never meaningfully controlled destroys trust in the arrangement, not just the tool.
Does your interface claim more than the model can deliver? Confident phrasing, a first-person voice, and a polished rationale all raise perceived capability without raising actual capability.
Which decisions have you deliberately decided not to delegate? A list of what stays human is a stronger governance artefact than a policy document about responsible AI.
Working together
Research is only useful when it moves something.
I work with leaders and teams who are past the pilot stage and now face the harder question of whether people are using AI well. Advisory, keynotes, team sessions, and hands-on sparring on live initiatives.