Testing AI Under Contradiction and Uncertainty

SWANK AI Guidance Note 15
Operational Testing · Uncertainty · AI Governance

Core Standard

TEST THE SYSTEM WHERE REALITY BECOMES DIFFICULT

AI should not be evaluated only by how well it performs with clean, complete and consistent information.

Operational quality becomes visible when:

evidence conflicts

information is incomplete

context changes

users communicate unpredictably

and

reassessment becomes necessary.


Purpose

Artificial intelligence systems are often demonstrated under favourable conditions.

The information may be:

  • complete;
  • clearly structured;
  • internally consistent;
  • easy to interpret;
  • and directly relevant to the task.

Real institutional environments are rarely so tidy.

Public services, education, healthcare, governance and complex organisations routinely operate with:

  • incomplete information;
  • contradictory accounts;
  • missing records;
  • ambiguous language;
  • disputed evidence;
  • changing circumstances;
  • and time pressure.

AI should therefore be tested under the conditions in which it will actually operate.


Ideal-Condition Testing Is Not Enough

A system may perform impressively when:

  • all information is available;
  • all sources agree;
  • terminology is consistent;
  • questions are clearly written;
  • the expected answer is obvious;
  • and no material evidence is disputed.

That can demonstrate technical capability.

It does not necessarily demonstrate operational reliability.

The more consequential the intended use, the more important it becomes to test what happens when the information environment is difficult.


Contradiction Testing

AI systems should be tested using scenarios where reliable sources disagree.

For example:

Source A records X.

Source B records Y.

The system should not automatically resolve the contradiction merely because one account:

  • is longer;
  • is more recent;
  • is written more confidently;
  • uses more formal language;
  • or fits an existing narrative more neatly.

An appropriate output may instead be:

The sources conflict and the available evidence does not establish a reliable resolution.

The ability to preserve contradiction can be a sign of higher-quality reasoning.


Contradiction Is Not Noise

In institutional systems, disagreement may be materially relevant evidence.

A model designed to produce one clean answer may treat contradiction as something to remove.

That can be operationally dangerous.

Where evidence is genuinely unresolved, the system should be capable of preserving:

  • competing sources;
  • evidential status;
  • uncertainty;
  • chronology;
  • and the need for human review.

A coherent answer should not be preferred over an accurate representation of uncertainty.


Testing Missing Information

Real-world systems often operate with incomplete records.

Testing should therefore include situations where:

  • key documents are unavailable;
  • chronology contains gaps;
  • an expected record is missing;
  • a user omits important context;
  • the system receives only part of a case;
  • or data exists in another system that has not been retrieved.

Relevant questions include:

  • Does the AI recognise the gap?
  • Does it state that important information is missing?
  • Does it fabricate a plausible explanation?
  • Does it overstate confidence?
  • Does it request additional information?
  • Does it escalate appropriately for human review?

A system that performs well only when all necessary information is present may be unreliable in practice.


Appropriate Uncertainty

A well-governed AI system should sometimes be able to say:

The available information does not establish a reliable conclusion.

Other appropriate outputs may include:

Insufficient information.

Sources conflict.

Further verification is required.

Human review is required.

This conclusion should not be relied upon without additional evidence.

The ability to refuse false certainty can itself be a safety feature.


Confidence Should Change With Evidence Quality

An AI system should not communicate the same level of certainty regardless of the quality of the evidence.

Testing should examine whether confidence changes when:

  • source quality declines;
  • evidence becomes contradictory;
  • records are missing;
  • chronology becomes unclear;
  • or critical variables are unavailable.

An output that remains equally confident under materially weaker evidence may mislead human users.


Testing Reassessment

Systems should also be tested for their ability to respond when new information arrives.

A realistic scenario may proceed like this:

  1. the system reaches an initial conclusion;
  2. new evidence becomes available;
  3. the new evidence materially contradicts the earlier position;
  4. the system is asked to reassess.

Relevant questions include:

  • Does the system revise its earlier assessment?
  • Does it preserve the earlier record appropriately?
  • Does the new evidence genuinely influence the output?
  • Does the model continue repeating an outdated classification?
  • Are downstream users alerted to material change?
  • Can the revised conclusion be traced to the new evidence?

A system that cannot meaningfully change its position when the evidence changes is not fully adaptive.


Testing Corrections

A related test concerns factual correction.

Suppose an earlier record contains an error and the organisation later supplies verified corrected information.

The system should be tested to determine whether:

  • the original error continues to appear;
  • the correction is recognised;
  • later summaries incorporate the corrected position;
  • the historical record remains appropriately visible;
  • and affected downstream conclusions are reconsidered.

A correction pathway is only useful if the operational system can actually respond to it.


Chronology Testing

Sequence matters.

Testing should include cases where the same events appear in different temporal orders.

For example:

Decision → new evidence → review

and

New evidence → decision → no review

contain the same three concepts.

They do not describe the same institutional process.

AI systems that summarise or reconstruct events should therefore be tested for their ability to preserve:

  • event date;
  • record date;
  • date information became known;
  • decision date;
  • review date;
  • and correction date.

Chronology is part of the evidence.


Ambiguous Language

Real users do not always communicate in perfectly structured prompts.

Testing should include:

  • ambiguous wording;
  • indirect requests;
  • mixed topics;
  • unclear terminology;
  • spelling or grammar errors;
  • second-language communication;
  • emotional language;
  • highly formal language;
  • very long inputs;
  • and unusual communication styles.

The relevant question is not whether the model can understand an ideal prompt.

It is whether the surrounding system behaves safely and usefully when communication is imperfect.


Emotionally Charged Language

Emotionally charged communication can affect both human and machine interpretation.

Testing should examine whether strong language causes a system to:

  • exaggerate apparent risk;
  • misclassify urgency;
  • infer motivation;
  • ignore substantive content;
  • or over-escalate.

The opposite should also be tested.

Calm language should not automatically cause a genuine urgent issue to be treated as insignificant.

Communication tone and substantive risk are not necessarily the same thing.


Unusual Communication Styles

Systems may be exposed to communication shaped by:

  • disability;
  • neurodevelopment;
  • cultural differences;
  • translation;
  • stress;
  • professional assistance;
  • AI assistance;
  • or individual writing style.

Testing should examine whether unusual structure or language produces inappropriate:

  • risk classifications;
  • credibility assumptions;
  • priority decisions;
  • or accessibility barriers.

Operational robustness requires more than performance on average users.


Long-Input Testing

Large language models may receive substantial documentation.

Testing should examine whether, as input length increases:

  • early information is forgotten;
  • contradictions are missed;
  • chronology is flattened;
  • central issues become buried;
  • unsupported details become more influential;
  • or summaries drift away from the source.

A system that performs well on short examples may behave differently when asked to process real institutional records.


Duplicated Records

Institutional data often contains repetition.

The same proposition may appear in:

  • an original record;
  • a later summary;
  • a report quoting the summary;
  • a dashboard;
  • and a subsequent assessment.

Testing should examine whether the AI treats repeated copies as independent corroboration.

The system should ideally distinguish:

multiple independent sources

from

multiple reproductions of one source.

Otherwise repetition can create artificial confidence.


Misleading Information

Operational testing may include information that is plausible but wrong.

For example:

  • an inaccurate date;
  • a fabricated citation;
  • a misleading summary;
  • a confident but unsupported assertion;
  • or an outdated record.

Relevant questions include:

  • Does the model detect the inconsistency?
  • Does it verify the claim?
  • Does it repeat it?
  • Does the surrounding workflow require human verification?
  • Can a later correction change the result?

The objective is not merely to test whether the model can detect deception.

It is to understand how the system handles unreliable information.


Fabricated Citations

AI systems may generate citations that appear plausible.

Testing should therefore include:

  • nonexistent sources;
  • incorrect quotations;
  • wrong page references;
  • misleading attribution;
  • and real sources cited for propositions they do not support.

Where citations materially influence institutional decisions, verification should be part of the workflow.

A citation that looks professional is not necessarily accurate.


Contradictory Instructions

AI systems may also receive conflicting instructions.

For example:

  • one policy says X;
  • another procedure says Y;
  • an internal email provides different guidance;
  • or a user prompt conflicts with a system rule.

Testing should assess whether the system:

  • identifies the inconsistency;
  • explains the conflict;
  • follows an appropriate hierarchy;
  • or simply selects one instruction without acknowledging the contradiction.

Institutional systems often contain inconsistencies.

AI governance should assume that reality rather than design only for perfect documentation.


Adversarial Operational Testing

Operational testing may include scenarios designed to stress the system.

These may involve:

  • misleading prompts;
  • irrelevant but persuasive information;
  • emotionally charged language;
  • fabricated citations;
  • duplicated records;
  • contradictory instructions;
  • incomplete evidence;
  • very long documents;
  • or attempts to induce overconfident conclusions.

The objective is broader than cybersecurity.

It is to understand whether the system remains reliable when the information environment becomes difficult.


Human Interaction Testing

Testing the model alone is not sufficient.

The people using it are part of the operational system.

A technically accurate system may still create poor outcomes if:

  • users misunderstand the output;
  • uncertainty is displayed poorly;
  • staff over-trust recommendations;
  • source material is difficult to access;
  • override pathways are too burdensome;
  • or time pressure eliminates meaningful review.

Operational testing should therefore include realistic human workflows.


Test the Reviewer, Not Just the Model

Useful questions include:

  • Does the reviewer understand what the AI output means?
  • Can they identify uncertainty?
  • Can they inspect the source?
  • Will they challenge an implausible result?
  • Can they override the system?
  • Do they know when not to rely upon it?
  • Does the interface encourage passive acceptance?

A model may be technically robust while the human-system interaction remains weak.


Testing Failure Pathways

A system should also be tested for what happens when it fails.

Questions include:

  • Is the error detected?
  • Who is alerted?
  • Can the workflow be paused?
  • Can a human intervene?
  • Can the affected output be withdrawn?
  • Can incorrect downstream records be corrected?
  • Is the incident documented?
  • Does the organisation learn from it?

Failure testing should examine the entire institutional response, not only the initial technical error.


Testing High-Stakes Systems

Higher-consequence uses require stronger testing.

This may apply where AI contributes to:

  • healthcare;
  • safeguarding;
  • education;
  • employment;
  • public services;
  • benefits;
  • disciplinary action;
  • financial access;
  • regulatory decisions;
  • or legal rights.

Testing in these environments should pay particular attention to:

  • false certainty;
  • source integrity;
  • human override;
  • correction;
  • bias;
  • contestability;
  • reversibility;
  • and the consequences of both false positives and false negatives.

The cost of error should influence the depth of testing.


Test Both Over-Escalation and Under-Recognition

Risk-detection systems may be tuned to favour sensitivity.

That may reduce missed concerns.

It may also create more false positives.

Testing should therefore examine both directions:

Over-Escalation

Does the system identify serious concern where the evidence does not justify it?

Under-Recognition

Does the system fail to identify significant concern because communication is calm, indirect or unusual?

A robust system should not be evaluated against only one type of error.


Testing for Bias

Testing should consider whether performance differs across relevant groups or communication contexts.

Depending upon the use case, this may include:

  • language;
  • writing style;
  • disability-related communication;
  • demographic differences;
  • age;
  • cultural context;
  • or other relevant variables.

The purpose is not merely to calculate average accuracy.

It is to identify whether certain people or situations experience systematically weaker outcomes.


Testing Before and After Deployment

Testing should not end at launch.

An AI-assisted system may change because:

  • the model changes;
  • vendor behaviour changes;
  • data changes;
  • users change how they use it;
  • workflows expand;
  • new integrations are added;
  • or institutional conditions evolve.

Organisations should consider testing:

before deployment

after material model change

after significant workflow change

after material incidents

and

periodically where consequence justifies it.


Operational Drift

A use case may gradually expand beyond what was originally tested.

For example:

Original use: summarising internal meeting notes.

Later use:

  • summarising disputed records;
  • supporting complaints decisions;
  • prioritising people;
  • or generating recommendations.

The technical system may be the same.

The consequence and governance requirements have changed.

Testing should follow the actual operational use, not merely the original approval.


What Good Testing Should Reveal

Operational testing should help an organisation understand:

  • where the system is reliable;
  • where uncertainty increases;
  • what types of evidence create difficulty;
  • when human review becomes necessary;
  • how errors propagate;
  • where users over-trust outputs;
  • whether correction works;
  • and what use cases may be inappropriate.

Testing should produce governance information, not merely a performance score.


Questions for Organisations

When evaluating an AI-assisted system, organisations may ask:

  1. Has the system been tested with contradictory evidence?
  2. Has it been tested with missing information?
  3. Can it express uncertainty appropriately?
  4. Does it revise conclusions when new evidence appears?
  5. Does correction change downstream outputs?
  6. Does it preserve chronology?
  7. Can it distinguish repetition from independent corroboration?
  8. How does it handle unusual communication styles?
  9. Has it been tested with misleading or fabricated information?
  10. Can human reviewers understand and override it?
  11. What happens when the system fails?
  12. Has it been tested under the conditions in which it will actually be used?

SIAAF Relevance

This Guidance Note principally relates to:

Domain 02 — Decision Integrity & Human Oversight

Can human reviewers exercise genuine judgment when the system encounters difficult evidence?

Domain 03 — Evidence & Traceability

Does the system preserve sources, contradictions, chronology and correction?

Domain 05 — Escalation, Challenge & Contestability

Can outputs be challenged and revised when evidence changes?

Domain 06 — Risk, Harm & Operational Resilience

How does the institution respond when the AI, human or workflow fails?

It may also engage:

Domain 07 — AI Literacy & Organisational Readiness

Where users need sufficient understanding to recognise uncertainty, contradiction and system limitations.


SWANK AI Standard

TEST THE SYSTEM WHERE REALITY BECOMES DIFFICULT

Do not test only:

clean prompts

complete data

agreeing sources

ideal users

and

expected answers.

Test:

contradiction

missing information

ambiguity

uncertainty

new evidence

correction

long inputs

unusual communication

human override

and

failure.

A system’s operational quality becomes visible when reality stops cooperating with the demonstration.


Related SWANK AI Guidance

Guidance Note 02 — AI Summarisation and Record Integrity

Guidance Note 08 — Reassessment Alongside Escalation

Guidance Note 11 — Meaningful Human Review

Guidance Note 13 — Automation Bias in Professional Decision-Making

Guidance Note 14 — AI Incident Reporting and Error Correction

Guidance Note 20 — Governing AI in High-Stakes Environments


SWANK AI

Independent AI & Institutional Assurance

We do not just review AI. We review the institutional systems responsible for governing it.

Scroll to Top