SWANK AI Guidance Note 15
Operational Testing · Uncertainty · AI Governance
Core Standard
TEST THE SYSTEM WHERE REALITY BECOMES DIFFICULT
AI should not be evaluated only by how well it performs with clean, complete and consistent information.
Operational quality becomes visible when:
evidence conflicts
information is incomplete
context changes
users communicate unpredictably
and
reassessment becomes necessary.
Purpose
Artificial intelligence systems are often demonstrated under favourable conditions.
The information may be:
- complete;
- clearly structured;
- internally consistent;
- easy to interpret;
- and directly relevant to the task.
Real institutional environments are rarely so tidy.
Public services, education, healthcare, governance and complex organisations routinely operate with:
- incomplete information;
- contradictory accounts;
- missing records;
- ambiguous language;
- disputed evidence;
- changing circumstances;
- and time pressure.
AI should therefore be tested under the conditions in which it will actually operate.
Ideal-Condition Testing Is Not Enough
A system may perform impressively when:
- all information is available;
- all sources agree;
- terminology is consistent;
- questions are clearly written;
- the expected answer is obvious;
- and no material evidence is disputed.
That can demonstrate technical capability.
It does not necessarily demonstrate operational reliability.
The more consequential the intended use, the more important it becomes to test what happens when the information environment is difficult.
Contradiction Testing
AI systems should be tested using scenarios where reliable sources disagree.
For example:
Source A records X.
Source B records Y.
The system should not automatically resolve the contradiction merely because one account:
- is longer;
- is more recent;
- is written more confidently;
- uses more formal language;
- or fits an existing narrative more neatly.
An appropriate output may instead be:
The sources conflict and the available evidence does not establish a reliable resolution.
The ability to preserve contradiction can be a sign of higher-quality reasoning.
Contradiction Is Not Noise
In institutional systems, disagreement may be materially relevant evidence.
A model designed to produce one clean answer may treat contradiction as something to remove.
That can be operationally dangerous.
Where evidence is genuinely unresolved, the system should be capable of preserving:
- competing sources;
- evidential status;
- uncertainty;
- chronology;
- and the need for human review.
A coherent answer should not be preferred over an accurate representation of uncertainty.
Testing Missing Information
Real-world systems often operate with incomplete records.
Testing should therefore include situations where:
- key documents are unavailable;
- chronology contains gaps;
- an expected record is missing;
- a user omits important context;
- the system receives only part of a case;
- or data exists in another system that has not been retrieved.
Relevant questions include:
- Does the AI recognise the gap?
- Does it state that important information is missing?
- Does it fabricate a plausible explanation?
- Does it overstate confidence?
- Does it request additional information?
- Does it escalate appropriately for human review?
A system that performs well only when all necessary information is present may be unreliable in practice.
Appropriate Uncertainty
A well-governed AI system should sometimes be able to say:
The available information does not establish a reliable conclusion.
Other appropriate outputs may include:
Insufficient information.
Sources conflict.
Further verification is required.
Human review is required.
This conclusion should not be relied upon without additional evidence.
The ability to refuse false certainty can itself be a safety feature.
Confidence Should Change With Evidence Quality
An AI system should not communicate the same level of certainty regardless of the quality of the evidence.
Testing should examine whether confidence changes when:
- source quality declines;
- evidence becomes contradictory;
- records are missing;
- chronology becomes unclear;
- or critical variables are unavailable.
An output that remains equally confident under materially weaker evidence may mislead human users.
Testing Reassessment
Systems should also be tested for their ability to respond when new information arrives.
A realistic scenario may proceed like this:
- the system reaches an initial conclusion;
- new evidence becomes available;
- the new evidence materially contradicts the earlier position;
- the system is asked to reassess.
Relevant questions include:
- Does the system revise its earlier assessment?
- Does it preserve the earlier record appropriately?
- Does the new evidence genuinely influence the output?
- Does the model continue repeating an outdated classification?
- Are downstream users alerted to material change?
- Can the revised conclusion be traced to the new evidence?
A system that cannot meaningfully change its position when the evidence changes is not fully adaptive.
Testing Corrections
A related test concerns factual correction.
Suppose an earlier record contains an error and the organisation later supplies verified corrected information.
The system should be tested to determine whether:
- the original error continues to appear;
- the correction is recognised;
- later summaries incorporate the corrected position;
- the historical record remains appropriately visible;
- and affected downstream conclusions are reconsidered.
A correction pathway is only useful if the operational system can actually respond to it.
Chronology Testing
Sequence matters.
Testing should include cases where the same events appear in different temporal orders.
For example:
Decision → new evidence → review
and
New evidence → decision → no review
contain the same three concepts.
They do not describe the same institutional process.
AI systems that summarise or reconstruct events should therefore be tested for their ability to preserve:
- event date;
- record date;
- date information became known;
- decision date;
- review date;
- and correction date.
Chronology is part of the evidence.
Ambiguous Language
Real users do not always communicate in perfectly structured prompts.
Testing should include:
- ambiguous wording;
- indirect requests;
- mixed topics;
- unclear terminology;
- spelling or grammar errors;
- second-language communication;
- emotional language;
- highly formal language;
- very long inputs;
- and unusual communication styles.
The relevant question is not whether the model can understand an ideal prompt.
It is whether the surrounding system behaves safely and usefully when communication is imperfect.
Emotionally Charged Language
Emotionally charged communication can affect both human and machine interpretation.
Testing should examine whether strong language causes a system to:
- exaggerate apparent risk;
- misclassify urgency;
- infer motivation;
- ignore substantive content;
- or over-escalate.
The opposite should also be tested.
Calm language should not automatically cause a genuine urgent issue to be treated as insignificant.
Communication tone and substantive risk are not necessarily the same thing.
Unusual Communication Styles
Systems may be exposed to communication shaped by:
- disability;
- neurodevelopment;
- cultural differences;
- translation;
- stress;
- professional assistance;
- AI assistance;
- or individual writing style.
Testing should examine whether unusual structure or language produces inappropriate:
- risk classifications;
- credibility assumptions;
- priority decisions;
- or accessibility barriers.
Operational robustness requires more than performance on average users.
Long-Input Testing
Large language models may receive substantial documentation.
Testing should examine whether, as input length increases:
- early information is forgotten;
- contradictions are missed;
- chronology is flattened;
- central issues become buried;
- unsupported details become more influential;
- or summaries drift away from the source.
A system that performs well on short examples may behave differently when asked to process real institutional records.
Duplicated Records
Institutional data often contains repetition.
The same proposition may appear in:
- an original record;
- a later summary;
- a report quoting the summary;
- a dashboard;
- and a subsequent assessment.
Testing should examine whether the AI treats repeated copies as independent corroboration.
The system should ideally distinguish:
multiple independent sources
from
multiple reproductions of one source.
Otherwise repetition can create artificial confidence.
Misleading Information
Operational testing may include information that is plausible but wrong.
For example:
- an inaccurate date;
- a fabricated citation;
- a misleading summary;
- a confident but unsupported assertion;
- or an outdated record.
Relevant questions include:
- Does the model detect the inconsistency?
- Does it verify the claim?
- Does it repeat it?
- Does the surrounding workflow require human verification?
- Can a later correction change the result?
The objective is not merely to test whether the model can detect deception.
It is to understand how the system handles unreliable information.
Fabricated Citations
AI systems may generate citations that appear plausible.
Testing should therefore include:
- nonexistent sources;
- incorrect quotations;
- wrong page references;
- misleading attribution;
- and real sources cited for propositions they do not support.
Where citations materially influence institutional decisions, verification should be part of the workflow.
A citation that looks professional is not necessarily accurate.
Contradictory Instructions
AI systems may also receive conflicting instructions.
For example:
- one policy says X;
- another procedure says Y;
- an internal email provides different guidance;
- or a user prompt conflicts with a system rule.
Testing should assess whether the system:
- identifies the inconsistency;
- explains the conflict;
- follows an appropriate hierarchy;
- or simply selects one instruction without acknowledging the contradiction.
Institutional systems often contain inconsistencies.
AI governance should assume that reality rather than design only for perfect documentation.
Adversarial Operational Testing
Operational testing may include scenarios designed to stress the system.
These may involve:
- misleading prompts;
- irrelevant but persuasive information;
- emotionally charged language;
- fabricated citations;
- duplicated records;
- contradictory instructions;
- incomplete evidence;
- very long documents;
- or attempts to induce overconfident conclusions.
The objective is broader than cybersecurity.
It is to understand whether the system remains reliable when the information environment becomes difficult.
Human Interaction Testing
Testing the model alone is not sufficient.
The people using it are part of the operational system.
A technically accurate system may still create poor outcomes if:
- users misunderstand the output;
- uncertainty is displayed poorly;
- staff over-trust recommendations;
- source material is difficult to access;
- override pathways are too burdensome;
- or time pressure eliminates meaningful review.
Operational testing should therefore include realistic human workflows.
Test the Reviewer, Not Just the Model
Useful questions include:
- Does the reviewer understand what the AI output means?
- Can they identify uncertainty?
- Can they inspect the source?
- Will they challenge an implausible result?
- Can they override the system?
- Do they know when not to rely upon it?
- Does the interface encourage passive acceptance?
A model may be technically robust while the human-system interaction remains weak.
Testing Failure Pathways
A system should also be tested for what happens when it fails.
Questions include:
- Is the error detected?
- Who is alerted?
- Can the workflow be paused?
- Can a human intervene?
- Can the affected output be withdrawn?
- Can incorrect downstream records be corrected?
- Is the incident documented?
- Does the organisation learn from it?
Failure testing should examine the entire institutional response, not only the initial technical error.
Testing High-Stakes Systems
Higher-consequence uses require stronger testing.
This may apply where AI contributes to:
- healthcare;
- safeguarding;
- education;
- employment;
- public services;
- benefits;
- disciplinary action;
- financial access;
- regulatory decisions;
- or legal rights.
Testing in these environments should pay particular attention to:
- false certainty;
- source integrity;
- human override;
- correction;
- bias;
- contestability;
- reversibility;
- and the consequences of both false positives and false negatives.
The cost of error should influence the depth of testing.
Test Both Over-Escalation and Under-Recognition
Risk-detection systems may be tuned to favour sensitivity.
That may reduce missed concerns.
It may also create more false positives.
Testing should therefore examine both directions:
Over-Escalation
Does the system identify serious concern where the evidence does not justify it?
Under-Recognition
Does the system fail to identify significant concern because communication is calm, indirect or unusual?
A robust system should not be evaluated against only one type of error.
Testing for Bias
Testing should consider whether performance differs across relevant groups or communication contexts.
Depending upon the use case, this may include:
- language;
- writing style;
- disability-related communication;
- demographic differences;
- age;
- cultural context;
- or other relevant variables.
The purpose is not merely to calculate average accuracy.
It is to identify whether certain people or situations experience systematically weaker outcomes.
Testing Before and After Deployment
Testing should not end at launch.
An AI-assisted system may change because:
- the model changes;
- vendor behaviour changes;
- data changes;
- users change how they use it;
- workflows expand;
- new integrations are added;
- or institutional conditions evolve.
Organisations should consider testing:
before deployment
after material model change
after significant workflow change
after material incidents
and
periodically where consequence justifies it.
Operational Drift
A use case may gradually expand beyond what was originally tested.
For example:
Original use: summarising internal meeting notes.
Later use:
- summarising disputed records;
- supporting complaints decisions;
- prioritising people;
- or generating recommendations.
The technical system may be the same.
The consequence and governance requirements have changed.
Testing should follow the actual operational use, not merely the original approval.
What Good Testing Should Reveal
Operational testing should help an organisation understand:
- where the system is reliable;
- where uncertainty increases;
- what types of evidence create difficulty;
- when human review becomes necessary;
- how errors propagate;
- where users over-trust outputs;
- whether correction works;
- and what use cases may be inappropriate.
Testing should produce governance information, not merely a performance score.
Questions for Organisations
When evaluating an AI-assisted system, organisations may ask:
- Has the system been tested with contradictory evidence?
- Has it been tested with missing information?
- Can it express uncertainty appropriately?
- Does it revise conclusions when new evidence appears?
- Does correction change downstream outputs?
- Does it preserve chronology?
- Can it distinguish repetition from independent corroboration?
- How does it handle unusual communication styles?
- Has it been tested with misleading or fabricated information?
- Can human reviewers understand and override it?
- What happens when the system fails?
- Has it been tested under the conditions in which it will actually be used?
SIAAF Relevance
This Guidance Note principally relates to:
Domain 02 — Decision Integrity & Human Oversight
Can human reviewers exercise genuine judgment when the system encounters difficult evidence?
Domain 03 — Evidence & Traceability
Does the system preserve sources, contradictions, chronology and correction?
Domain 05 — Escalation, Challenge & Contestability
Can outputs be challenged and revised when evidence changes?
Domain 06 — Risk, Harm & Operational Resilience
How does the institution respond when the AI, human or workflow fails?
It may also engage:
Domain 07 — AI Literacy & Organisational Readiness
Where users need sufficient understanding to recognise uncertainty, contradiction and system limitations.
SWANK AI Standard
TEST THE SYSTEM WHERE REALITY BECOMES DIFFICULT
Do not test only:
clean prompts
complete data
agreeing sources
ideal users
and
expected answers.
Test:
contradiction
missing information
ambiguity
uncertainty
new evidence
correction
long inputs
unusual communication
human override
and
failure.
A system’s operational quality becomes visible when reality stops cooperating with the demonstration.
Related SWANK AI Guidance
Guidance Note 02 — AI Summarisation and Record Integrity
Guidance Note 08 — Reassessment Alongside Escalation
Guidance Note 11 — Meaningful Human Review
Guidance Note 13 — Automation Bias in Professional Decision-Making
Guidance Note 14 — AI Incident Reporting and Error Correction
Guidance Note 20 — Governing AI in High-Stakes Environments
SWANK AI
Independent AI & Institutional Assurance
We do not just review AI. We review the institutional systems responsible for governing it.
