SWANK AI Guidance Note 06
AI Governance · Evidential Reliability · Human Review
Core Standard
INDICATOR IS NOT DETERMINATION
AI-detection technology may contribute information.
It should not be treated as infallible proof of authorship.
Where consequences matter, conclusions should remain proportionate to the total evidence available.
Purpose
As generative AI becomes more widely used, organisations increasingly want to determine whether text, images, documents or assessed work were produced with AI assistance.
Detection tools may sometimes provide useful indicators.
They may also produce:
- false positives;
- false negatives;
- uncertain probability scores;
- inconsistent classifications;
- and overconfident interpretations of stylistic patterns.
An AI-detection result is therefore an analytical signal, not a factual determination of authorship.
The Evidential Problem
AI-detection systems generally attempt to identify statistical, linguistic or stylistic characteristics associated with generated material.
Those characteristics may also appear in human-produced work.
Formal writing may be:
- highly structured;
- grammatically consistent;
- repetitive;
- concise;
- unusually polished;
- or statistically predictable.
None of those features proves that AI was used.
Conversely, AI-assisted material may be edited, reorganised or substantially rewritten by a human and therefore evade detection.
Detection systems can therefore be wrong in both directions.
Probability Is Not Proof
A numerical output may appear authoritative.
For example:
“92% likely AI-generated”
can easily be interpreted as:
“There is a 92% probability that this person used AI.”
Those statements are not necessarily equivalent.
A detector score may instead reflect how strongly certain characteristics of the material resemble patterns associated with generated text under that particular system.
Organisations should understand:
- what the score actually measures;
- what validation evidence exists;
- what error rate applies;
- what population the tool was tested on;
- and whether the result is appropriate for the decision being made.
Numerical precision should not create false evidential certainty.
False Positives
A false positive occurs when human-produced work is classified as AI-generated.
This may have serious consequences where detection is used in:
- education;
- employment;
- disciplinary processes;
- complaints handling;
- authorship disputes;
- professional assessment;
- or other consequential settings.
Certain forms of human writing may be especially vulnerable to misclassification, including:
- formal writing;
- second-language writing;
- highly edited work;
- formulaic academic writing;
- technical prose;
- accessibility-assisted writing;
- and structured administrative correspondence.
A high detector score should therefore not silently become a factual finding.
False Negatives
The opposite problem also exists.
AI-assisted work may not be detected where it has been:
- heavily edited;
- rewritten;
- combined with human work;
- generated in small sections;
- paraphrased;
- or created using systems that produce less statistically distinctive output.
A low detector score therefore does not prove that AI was not used.
Detection technology may indicate patterns.
It cannot necessarily establish authorship with certainty.
Wider Evidential Context
Where determining AI assistance genuinely matters, organisations should consider the wider evidence.
Relevant material may include:
- previous work;
- drafting history;
- document metadata;
- version history;
- research notes;
- source materials;
- contemporaneous records;
- assessment conditions;
- communication with the author;
- ability to explain reasoning;
- ability to identify and discuss sources;
- and other relevant evidence.
No single technical indicator should silently become the entire evidential basis for a consequential conclusion.
Education
AI detection is particularly sensitive in educational environments.
An allegation of prohibited AI use may affect:
- grades;
- progression;
- disciplinary records;
- qualifications;
- academic reputation;
- confidence;
- and future opportunities.
Students should therefore understand the institution’s AI rules before assessment takes place.
Institutions should define:
- what AI use is permitted;
- what use requires disclosure;
- what use is prohibited;
- what evidence may be considered where misuse is suspected;
- and how a student can challenge an allegation.
Detection software should support an evidential process.
It should not replace one.
Authorship Assessment
Where authorship is disputed, an appropriate review may consider whether the author can:
- explain the argument;
- identify sources;
- describe the drafting process;
- reproduce relevant reasoning;
- discuss earlier versions;
- explain stylistic choices;
- and demonstrate knowledge consistent with the work submitted.
This may provide context that a detection score alone cannot capture.
The objective should be to assess the evidence fairly rather than force a technical indicator to carry more weight than it can support.
Administrative Environments
The same principle applies outside education.
AI detection should not automatically determine whether:
- correspondence is genuine;
- an individual understands a submission;
- a complaint deserves consideration;
- a document is unreliable;
- a person has acted deceptively;
- or communication should receive lower priority.
An organisation may legitimately identify that AI appears to have assisted drafting.
That does not establish whether the substantive content is true, false, relevant or adequately evidenced.
Substance still requires review.
AI-Assisted Accessibility
Some individuals may use generative AI to support:
- writing;
- speech-related communication;
- language;
- cognitive organisation;
- information structuring;
- fatigue management;
- or navigation of complex administrative processes.
An institution should therefore take care before treating apparent AI assistance as evidence of:
- deception;
- lack of understanding;
- lack of authenticity;
- or improper participation.
The question should remain whether the particular use was permitted and whether the substantive information can be supported.
Human Review
Consequential findings involving suspected AI use should ordinarily remain subject to meaningful human review.
The reviewer should be able to consider:
- the detector result;
- its limitations;
- the surrounding evidence;
- relevant policy;
- the author’s explanation;
- alternative explanations;
- and the consequence of an incorrect finding.
Human review should not consist merely of approving the detector’s output.
Detector Governance
Before using an AI-detection tool operationally, organisations should ask:
- What exactly does the system claim to detect?
- What evidence supports its accuracy?
- What are its false-positive and false-negative rates?
- Has it been validated on relevant populations?
- Does performance vary by language or writing style?
- How often is the model updated?
- Can results be independently reviewed?
- Does the system provide meaningful explanation?
- What information is retained?
- What happens when a result is challenged?
- Can a human override the classification?
Approval should attach to a defined tool and use case rather than to the general idea of “AI detection.”
Consequence-Sensitive Use
The evidential standard should increase with consequence.
A detector used informally to identify material worth reviewing presents a different risk from a detector used to support:
- disciplinary action;
- loss of academic credit;
- exclusion;
- employment consequences;
- allegations of dishonesty;
- or adverse administrative findings.
The more consequential the outcome, the weaker the case for relying on one probabilistic indicator without supporting evidence.
Uncertainty Should Be Visible
Where a detector cannot establish a reliable conclusion, the appropriate institutional position may be:
AI assistance is suspected but not established.
or:
The detection result is one indicator and requires further review.
That is analytically stronger than presenting uncertain evidence as settled fact.
A refusal to overstate certainty is part of good governance.
Questions for Organisations
Where AI-detection technology is used, organisations may ask:
- What exactly does the tool measure?
- What is the validated error rate?
- Could human writing produce the same result?
- Could AI-assisted writing avoid detection?
- What other evidence is available?
- Is the consequence proportionate to the strength of the evidence?
- Has the individual been given an opportunity to respond?
- Can a human reviewer override the result?
- Are accessibility or communication factors relevant?
- Is the institution treating an indicator as though it were proof?
- Can the decision later be reconstructed?
- Is the uncertainty clearly documented?
SIAAF Relevance
This Guidance Note principally relates to:
Domain 02 — Decision Integrity & Human Oversight
Where people must exercise independent judgment rather than simply approve technical classifications.
Domain 03 — Evidence & Traceability
Where a detector result forms part of the evidential basis for a consequential finding.
Domain 05 — Escalation, Challenge & Contestability
Where individuals must be able to challenge or correct an adverse classification.
It may also engage:
Domain 06 — Risk, Harm & Operational Resilience
Where incorrect detection may produce significant institutional consequences.
Domain 07 — AI Literacy & Organisational Readiness
Where staff need to understand what AI-detection tools can and cannot reliably establish.
SWANK AI Standard
INDICATOR IS NOT DETERMINATION
Detection technology may contribute evidence.
It should not replace evidential judgment.
Where consequences matter:
verify
contextualise
permit challenge
preserve uncertainty
require meaningful human review
and
keep conclusions proportionate to the total evidence available.
Related SWANK AI Guidance
Guidance Note 01 — Handling AI-Assisted Complaints and Correspondence
Guidance Note 04 — AI Literacy and Academic Integrity
Guidance Note 07 — Public-Sector AI and Accessibility
Guidance Note 12 — AI and Administrative Fairness
Guidance Note 13 — Automation Bias in Professional Decision-Making
SWANK AI
Independent AI & Institutional Assurance
We do not just review AI. We review the institutional systems responsible for governing it.
