Engineering Trust-based Human-AI Collaboration

EN DE

Doctoral research | Trust in human-AI collaboration

Your AI works.
Your people still get it wrong.

Five years of research on one question: what makes people trust an AI the right amount. Not more, not less. The answer turns out to be uncomfortable for anyone rolling out AI right now, because the thing that drives trust is not how good the system is. It is how good it looks.

Engineering Trust-based Human-AI Collaboration

3,358records screened across four research disciplines
479distinct drivers of trust identified and mapped
253professionals in a controlled scheduling experiment
3peer-reviewed publications behind the argument

The problem worth your attention

Adoption is not the finish line.

Most AI programmes measure whether people use the system. That is the wrong finish line. What determines performance is whether people’s confidence tracks what the system can actually do. Researchers call this trust calibration, and it fails in both directions.

01

Capable systems get followed off a cliff

In one study, recruiters given a strong AI followed it uncritically and the combined team performed worse than the AI alone. Adoption went up. Outcomes went down.

02

Good systems get abandoned after five bad calls

As few as five confidently wrong predictions can trigger a lasting collapse in trust. Worse, people forgive an AI less than a human for the identical mistake, even when the AI is more accurate overall.

03

The stakes are highest where decisions are fast

In production and operations, a rescheduling call is made in minutes and cascades into delivery dates, utilisation, and cost. There is no time to audit the recommendation. Trust has to be right in advance.

The evidence base

Three studies, one narrowing question.

The research moves from a broad map of everything that shapes trust, to what actually matters in industrial decision-making, to a controlled test of whether it can be deliberately influenced.

Study I · Systematic review

What actually drives trust in intelligent systems

179 publications screened from 3,358 records across robotics, automation, information systems, and human factors.

479 separate drivers, sorted into the environment, the person, the system, and a fourth group earlier models had missed: the interaction itself. The honest conclusion is that the list is too long to design from. The useful question is which few drivers do the work.

Study II · Interviews and survey

Which drivers matter in real operations

Expert interviews in production management, followed by a survey and structural equation modelling.

One driver dominates everything else: how capable the AI is perceived to be. Comprehensibility matters, but only as a fitting second. Objective factors such as task predictability and the cost of an error never reach trust directly, they arrive through perception. One driver surfaced that no earlier trust model had included: digital affinity, the personal pull toward working with new technology.

Study III · Controlled experiment

Do explanations actually help

253 manufacturing professionals and engineers scheduling a shop floor. A black-box AI with short written rationales against a fully transparent rule, with identical recommendations in both conditions.

Short rationales made people follow the AI more. They did not raise considered trust directly, and they did not improve understanding. Trust moved only downstream, by making the system look more capable. The explanations worked as a signal of competence, not as an explanation.

What the research establishes

Four findings that change how you roll out AI.

These are not opinions about AI. They are measured effects, replicated across an observational study and a controlled experiment using the same instruments.

The thesis condenses them into three insights: perceived capability is the mechanism that carries trust, trust and reliance have to be treated as two different things, and calibration is a property of the arrangement rather than of the individual user.

01

Trust is bought with appearances, not accuracy

Perceived capability is by far the strongest driver of trust in AI decision support. It held across three different model specifications and was then replicated under experimental control. Everything else, including how risky the decision is and how much an error costs, works through that perception rather than around it.

The uncomfortable corollary is that making a system look capable already moves trust. Vendors know this. Your people cannot tell the difference from the interface alone.

Study II, replicated under experimental control in Study III
02

Explanations persuade more than they explain

The standard assumption in explainable AI is that showing the reasoning improves understanding, and that understanding produces appropriate trust. In an early-interaction setting that assumption did not hold. Brief rationales lifted the sense that the system was capable while barely touching comprehension.

Trust did rise, but only downstream of that impression, never from the explanation itself. Explanations are therefore an adoption lever. Treating them as a safety mechanism is a category error.

Study III
03

Behaviour and belief come apart

People acted on the AI’s advice more often without their considered trust rising to match. That gap is the dangerous one. It looks like successful adoption in every dashboard you have, while the critical judgment that was supposed to catch the bad recommendation quietly disengages.

Evidence from outside the factory points the same way. In a chess study, people relied more on a partner they knew was an AI while reporting more trust in one they believed was human, and the weaker performers were the ones most swayed by who the partner was.

For a single decision, behaviour is what matters. Over months of decisions, the underlying belief decides whether people stay alert or drift into compliance.

Study III, consistent with wider literature
04

Explanations help experts more than novices

The expected pattern was that explanations help most where knowledge is thinnest. The data pointed the other way. The trust benefit was smaller for people who found the task hard and for those without domain experience, and larger for experts, who read the same cues as credible signals of competence. Reliance was not moderated at all, which suggests the novices were not simply following blindly.

These boundary effects sit below the study’s own sensitivity threshold and should be read as directional rather than settled. The robust part is the asymmetry: trust is fragile and context-dependent, while reliance is stubborn.

Study III, moderation analysis

The meta-finding. Trust calibration is a lever and a vulnerability at the same time. The perception pathway shows exactly where design can intervene. The gap between behaviour and belief shows exactly where that intervention can backfire. Which means the thing you need to design is not the model. It is the arrangement around it: who decides what, who checks what, and what the interface is allowed to imply.

That arrangement is the third insight and the one that survives the move from research to practice. Calibration is not a property of the model, and it cannot be left to the individual user. The thesis proposes building a discipline around it under the working title Human-AI Collaboration Engineering. At this stage that names the unit of design and its first focus. The methods and reusable patterns still have to be built and tested in the field.

The stress test

Generative AI makes all of this harder.

The studies examined task-specific AI. Large language models change five conditions at once, and each one attacks a mechanism the research identified. Hover or tap a card for what it does to your organization.

Risk 01

The competence boundary disappears

One system drafts, summarises, codes, and plans, and is excellent at some of it and unreliable at the rest.

With narrow AI, people learned where the system was strong. A jagged capability profile removes that learning curve entirely.

Your people cannot calibrate against a target that has no stable edges. Expect confident use in exactly the areas where the model is weakest.

Attacks Finding 01
Risk 02

Fluency reads as competence

Confident, articulate, human-sounding output regardless of whether the reasoning underneath is sound.

If perceived capability is the gateway to trust, then a system that manufactures the impression of capability has direct access to that gateway.

The Clever Hans problem returns in linguistic form. The performance is convincing and the mechanism behind it stays hidden.

Attacks Finding 01
Risk 03

The system writes its own explanations

Explanation stops being something you design and becomes something the model generates.

Self-generated rationales can be persuasive, tailored, or simply disconnected from what the model actually did. Prompt injection and data poisoning let third parties steer them.

You lose your independent reference point. Routing validated content through a model only stacks one opaque layer on another.

Attacks Finding 02
Risk 04

The system agrees with you

Sycophantic responses raise the user’s sense of being right, their trust, and their intention to come back.

The interaction becomes a mirror that confirms the position you arrived with, removing the friction that recalibration depends on. It adapts to your language and preferences in real time.

Modelling work shows this can escalate false confidence even for a perfectly rational user. It is not a problem you can train out of people.

Attacks Finding 03
Risk 05

The judgment doing the checking erodes

Heavy AI use correlates with self-overestimation, cognitive offloading, and weaker critical thinking.

More strikingly, higher AI literacy correlates with higher overestimation of one’s own performance, not lower. Delegation makes it worse, since people commission more questionable decisions through goals and examples than through explicit instructions.

Inflated signals and self-written explanations arrive at users whose capacity to evaluate them is quietly declining. This is the compounding one.

Attacks Finding 03
The verdict

A second-order problem

The thing to be judged gets harder to judge, while the ability to judge it gets weaker.

Neither half can be solved by better models. Both are properties of how the work is arranged around the model.

What stays human is the class of decisions that cannot be computed at all, where values conflict, evidence runs out, and responsibility cannot be delegated without the outcome losing its legitimacy.

Where design has to step in

From evidence to practice

Five strategies for AI people can judge.

Existing guidance locates the fix in model design, organizational practice, and regulation. The research adds a fourth layer that is usually skipped: the design of the interaction itself.

Their evidential weight varies. Some translate a mechanism measured in the studies. Others extend it with the wider literature. The distinction is stated rather than hidden.

01
Model design

Treat opacity as a cost, not a default

For a wide range of tasks, interpretable models match black-box accuracy. Use the simplest model that clears the performance bar, and make anyone proposing an opaque one show what the extra accuracy is worth against the calibration risk it introduces.

02
Model design

Engineer honesty into how the system presents itself

If perceived capability is the gateway to trust, the system must not oversell itself at that gateway. Communicate confidence boundaries, surface uncertainty, and flag misbehaviour. Necessary, but not sufficient, since an honest system can still be opaque and still invite people to read more into it than is there.

03
Interaction design

Build in cognitive distance

Cognitive distance means deliberate reflection on the advice before accepting it. Where opacity is unavoidable, that reflection has to be built into the interaction, because you cannot delegate vigilance to the user and expect it to hold. The risk is sharpest where AI shapes what people notice before they have started thinking.

Practical moves are small and cheap: drop the first-person voice, turn assertions into questions, and require a position before the AI offers one.

04
Organization

Put verification in the process, not in the person

Self-calibration varies with expertise and may itself degrade with AI use. Use independent checks, escalation routines, and role-based oversight. Notably, control mechanisms tend to increase trust and adoption rather than slow them down. Keep the high-stakes, long-horizon, hard-to-reverse decisions with humans.

05
Governance

Use regulation as a backstop, not a strategy

Regulation shapes market conditions, for instance by forcing disclosure that deflates inflated capability claims. It can produce trustworthy systems. It cannot produce calibrated trust, because compliance does not make users infer capability correctly.

Where to start

Six questions worth asking this quarter.

If your AI initiative cannot answer these, calibration is being left to chance. That is not a technology gap. It is a design gap.

Can your users state where the system is unreliable? If they cannot name a weakness, they have not calibrated. They have simply accepted.

Are you measuring adoption, or appropriate override? Usage rates rise in both healthy and unhealthy scenarios. The rate at which people correctly reject a bad recommendation does not.

Do your explanations survive contact with a wrong answer? An explanation that sounds equally convincing when the output is wrong is a persuasion device, not a calibration aid.

Who is accountable for a decision the human could not realistically have checked? Holding people responsible for outcomes they never meaningfully controlled destroys trust in the arrangement, not just the tool.

Does your interface claim more than the model can deliver? Confident phrasing, a first-person voice, and a polished rationale all raise perceived capability without raising actual capability.

Which decisions have you deliberately decided not to delegate? A list of what stays human is a stronger governance artefact than a policy document about responsible AI.

Working together

Research is only useful when it moves something.

I work with leaders and teams who are past the pilot stage and now face the harder question of whether people are using AI well. Advisory, keynotes, team sessions, and hands-on sparring on live initiatives.

Scroll to Top
Cookie Consent with Real Cookie Banner