Source Traceability by Design

SWANK AI Guidance Note 17
Source Traceability · Record Integrity · AI Governance

Core Standard

EVERY MATERIAL CLAIM SHOULD HAVE A PATH BACK TO SOURCE

The more consequential the output, the more important that pathway becomes.

AI may help organisations summarise, classify, retrieve and interpret information.

It should not make the evidential foundation invisible.


Purpose

As AI systems increasingly process organisational information, source traceability should be treated as a design requirement rather than an afterthought.

A useful AI output should allow a material proposition to be connected back to the information from which it arose.

Without traceability, a reviewer may encounter a statement such as:

“The record establishes that…”

without being able to determine:

  • which record;
  • which author;
  • which date;
  • which passage;
  • whether the information was disputed;
  • whether it was later corrected;
  • or whether later evidence changed the position.

The output may appear clear while the evidence beneath it becomes increasingly difficult to inspect.


Traceability Is Part of Reliability

An AI-generated conclusion is not strengthened merely because it is written clearly.

Reliability depends partly upon whether a reviewer can determine:

where the information came from

what the source actually said

what status the information had

and

whether the AI interpretation remained faithful to it.

A system that produces strong summaries without preserving source pathways may create administrative efficiency at the cost of evidential visibility.


From Output Back to Evidence

A well-designed AI-assisted information system should ideally allow a reviewer to move through a visible pathway:

output → proposition → source

For example:

AI summary:
“The organisation had previously identified a material implementation risk.”

The system should, where consequence requires it, allow a reviewer to identify:

  • the relevant document;
  • the date;
  • the author or originating function;
  • the exact section;
  • whether the statement represented fact or opinion;
  • and whether later material modified it.

This turns the AI output into an entry point to evidence rather than a replacement for evidence.


Different Uses Require Different Levels of Traceability

Not every AI output requires the same documentation burden.

A low-consequence brainstorming tool may require little source attribution.

A system supporting consequential administrative or professional decisions may require substantially more.

Depending upon the context, useful traceability may include:

  • document identifier;
  • source title;
  • author or origin;
  • date;
  • page;
  • paragraph;
  • section;
  • version;
  • source status;
  • confidence;
  • relevant qualification;
  • and whether the material remains disputed.

The appropriate level should remain proportionate to consequence.


Material Claims Require Stronger Traceability

The stronger the institutional reliance on an AI-generated proposition, the stronger the case for preserving its source pathway.

For example, traceability becomes particularly important where AI output contributes to:

  • safeguarding decisions;
  • healthcare;
  • education;
  • employment;
  • disciplinary processes;
  • public services;
  • financial access;
  • complaints outcomes;
  • regulatory action;
  • or other consequential institutional decisions.

A material claim should not become difficult to verify merely because AI made it easier to summarise.


Generated Citations

AI-generated citations should not be assumed accurate merely because they appear plausible.

AI systems may produce:

  • nonexistent references;
  • incorrect page numbers;
  • inaccurate quotations;
  • wrong authors;
  • misattributed material;
  • or sources that do not support the proposition claimed.

Where citations materially support a decision or recommendation, they should be verified against the underlying source.

A citation is useful only if the path actually leads to the evidence.


Retrieved Sources and Generated Sources Are Different

Organisations should understand whether an AI system is:

retrieving an existing source reference

or

generating a citation from model output.

Those are different processes.

A retrieved citation may still require checking.

A generated citation may require even greater caution.

System design should make that distinction visible where possible.


Summaries Should Preserve Source Connections

AI summarisation is especially dependent upon traceability.

A summary may contain several propositions drawn from different documents.

Good design should allow a reviewer to identify which source supports each material statement.

Without that structure, an apparently concise summary can force later users to reconstruct the entire evidence base manually.

The better pathway is:

summary

material proposition

identified source

original context


Source Traceability and Disagreement

Traceability becomes particularly important when sources conflict.

For example:

Source A records X.

Source B records Y.

The available material does not resolve the difference.

That is analytically stronger than producing a single blended statement.

If a system preserves source attribution, disagreement remains visible.

If source attribution disappears, later readers may incorrectly assume that a single institutional position has been established.


Traceability Prevents False Corroboration

Repeated information can create the appearance of independent support.

For example:

  1. an original source records a proposition;
  2. an AI summary repeats it;
  3. a later report quotes the summary;
  4. another AI system retrieves the later report;
  5. the proposition appears in several records.

Without source traceability, a reviewer may believe several sources independently support the conclusion.

In reality, they may all originate from the same initial record.

Traceability helps distinguish:

independent corroboration

from

repetition of one source.


Traceability and Chronology

Source identification should preserve temporal context where relevant.

A reviewer may need to know:

  • when the event occurred;
  • when the record was created;
  • when the information became known;
  • when it was relied upon;
  • when a decision followed;
  • and whether later evidence changed the position.

The source alone may not be enough.

The relationship between source and time may also be material.


Traceability and Corrections

Where a source is later corrected, systems should be able to identify which downstream outputs may have relied upon the earlier version.

Relevant questions include:

  • Which summaries used the incorrect source?
  • Which decisions relied upon them?
  • Were later AI prompts influenced by the original error?
  • Has the corrected material propagated?
  • Can a reviewer distinguish the historical record from the corrected position?

Traceability enables correction to travel through the system.

Without it, an institution may know that a source was wrong while remaining unable to identify where the error went.


Source Provenance

Traceability is not only about finding a document.

It may also require understanding provenance.

Relevant questions include:

  • Who created the source?
  • What system produced it?
  • Is it original or derived?
  • Has it been edited?
  • Was it generated by AI?
  • Was it verified?
  • Is it a primary or secondary record?
  • Does it quote another source?
  • Is the version current?

Two records may contain identical text while having very different evidential significance.


Fact, Inference and AI-Generated Material

A traceable system should preserve distinctions between:

  • source fact;
  • reported statement;
  • professional interpretation;
  • analytical inference;
  • AI-generated synthesis;
  • recommendation;
  • and unresolved information.

A generated summary should not make all of these categories appear equivalent.

The user should be able to understand not only:

where did this statement come from?

but also:

what kind of statement is it?


Source Traceability in Retrieval Systems

Retrieval-augmented AI systems can improve access to organisational information.

Their governance should still examine:

  • what records are eligible for retrieval;
  • whether source references remain visible;
  • whether outdated material can dominate;
  • whether corrections are indexed;
  • whether permissions are respected;
  • and whether retrieved passages are presented in sufficient context.

Retrieval is not automatically traceability.

A system may retrieve a source while still obscuring why a particular proposition was produced.


Context Around the Source

A correct citation can still mislead if the surrounding context is omitted.

For example, a paragraph may state:

“X was reported.”

while the following sentence states:

“This could not be independently verified.”

If the AI retrieves only the first statement, the citation may be technically accurate but substantively incomplete.

Source traceability should therefore be accompanied by enough context to preserve meaning.


System Procurement

Traceability should be considered before an AI system is purchased or deployed.

Organisations may ask vendors:

  • Can outputs link to underlying records?
  • Are source references preserved automatically?
  • Can users inspect retrieved passages?
  • Are citations generated or retrieved?
  • Does the system preserve page or section references?
  • Can multiple sources be shown separately?
  • Can disputed information remain identified?
  • Are source corrections reflected in future outputs?
  • Can administrators reconstruct how an output was produced?

A system that cannot preserve appropriate traceability may be unsuitable for consequential use even if its outputs are highly fluent.


Traceability and Human Review

Meaningful human review depends upon source access.

A reviewer cannot adequately challenge AI-generated analysis if the evidential pathway ends at the output.

Where consequences are significant, reviewers should be able to:

  • inspect source material;
  • verify citations;
  • identify conflicting records;
  • distinguish source from inference;
  • recognise uncertainty;
  • and determine whether the AI omitted material context.

Traceability enables human judgment.


Auditability

Source traceability supports later review.

An independent reviewer should ideally be able to determine:

  • what information was available;
  • which sources were relied upon;
  • what AI-generated interpretation was produced;
  • what the human reviewer saw;
  • what decision followed;
  • and whether later correction occurred.

This does not require preserving every technical detail in every context.

It requires enough evidence to reconstruct material decisions proportionately.


High-Stakes Systems

The need for traceability increases where incorrect output could materially affect:

  • rights;
  • safety;
  • opportunities;
  • essential services;
  • professional decisions;
  • or significant institutional action.

In high-stakes environments, statements such as:

“the system indicates”

or

“the record shows”

should not become evidential dead ends.

The pathway should remain visible.


Questions for Organisations

Where AI is used to process organisational information, organisations may ask:

  1. Can every material claim be traced to an identifiable source?
  2. Can users inspect the original material?
  3. Are citations retrieved or generated?
  4. Are generated references verified?
  5. Is source provenance visible?
  6. Are disputed sources distinguishable?
  7. Can the system preserve contradictory evidence?
  8. Can repeated records be traced back to a common origin?
  9. Can corrections be linked to downstream outputs?
  10. Does the system preserve chronology where relevant?
  11. Can a human reviewer reconstruct the evidence pathway?
  12. Is the level of traceability proportionate to the consequence of the output?

SIAAF Relevance

This Guidance Note principally relates to:

Domain 03 — Evidence & Traceability

Can the organisation prove how an important conclusion was reached?

It also engages:

Domain 02 — Decision Integrity & Human Oversight

Where human reviewers need access to source material before relying upon AI-generated outputs.

Domain 04 — Communication & Feedback Integrity

Where information must retain meaning and provenance as it moves through an institution.

Domain 05 — Escalation, Challenge & Contestability

Where affected people or reviewers need to challenge inaccurate or unsupported material.

Domain 06 — Risk, Harm & Operational Resilience

Where weak traceability may allow errors to propagate through consequential systems.


SWANK AI Standard

EVERY MATERIAL CLAIM SHOULD HAVE A PATH BACK TO SOURCE

Where consequence requires it, organisations should preserve:

source

provenance

date

context

evidential status

contradiction

correction

and

the relationship between source and output.

AI should make institutional information easier to navigate.

It should not make the evidence supporting a conclusion harder to find.


Related SWANK AI Guidance

Guidance Note 02 — AI Summarisation and Record Integrity

Guidance Note 05 — Preserving Disagreement in AI-Assisted Records

Guidance Note 14 — AI Incident Reporting and Error Correction

Guidance Note 15 — Testing AI Under Contradiction and Uncertainty

Guidance Note 18 — AI, Chronology and Organisational Memory

Guidance Note 19 — Correction Rights in AI-Assisted Systems


SWANK AI

Independent AI & Institutional Assurance

We do not just review AI. We review the institutional systems responsible for governing it.

Scroll to Top