In brief
- Start with frequent, well-defined questions rather than the entire data estate.
- Authoritative policies and maintained manuals are better sources than unreviewed shared folders.
- Duplicates, outdated versions and missing ownership are not fixed by RAG.
- Sensitive content needs a defined purpose, permission, minimisation and controlled lifecycle.
Which questions should the data answer?
Selection starts with the work of the target audience. Collect concrete questions, required decisions and current search paths. A source is valuable when it answers those questions with sufficient authority and detail.
- Frequency and effort of current information search
- Impact of a correct or incorrect answer
- Intended audience and required languages
- Need for a citation, version date or live value
What makes a source ready?
A suitable source has a subject owner, traceable validity and structure that makes relevant sections findable. The examples show why technical readability alone does not make a source ready.
| Real source example | Readiness | Why | Next step |
|---|---|---|---|
| Approved HR policy with owner, version and role-based access | Ready | Authority, validity and audience are traceable. | Import passages, permissions and citations and test typical questions. |
| Current product manual as a readable PDF but without a revision date | Conditionally ready | The content is useful, but freshness cannot be assessed reliably. | Add owner, version and replacement process before import. |
| Team wiki containing approved pages, drafts and obsolete guidance | Prepare first | Results may sound plausible without being authoritative. | Label status, archive old versions and identify authoritative pages. |
| Personal email archives and meeting notes | Do not include by default | Permissions, purpose, completeness and domain approval are unclear. | Move only specifically curated findings into an owned source. |
| Inventory from an ERP export | Wrong data path | The value is stale after the next transaction. | Retrieve current inventory through an authorised API from the system of record. |
Which data should not be included without review?
A comprehensive repository can reduce both quality and security. Personal notes, old exports, email archives and documents without owners often contain contradictions or information not intended for the broad audience.
- Drafts without clear labelling or approval
- Several versions with no identifiable current edition
- Personal data or trade secrets without a necessary purpose
- Content with unclear licensing or reuse rights
- Tables and values that should come live from a business system
How should the first data scope be prioritised?
A small, well-maintained domain is a better foundation than an uncontrolled full import. The team assesses value, source readiness, integration effort and failure risk and starts with a clearly owned scope.
- Step 1
Collect questions
Capture real information needs from the target process.
- Step 2
Map sources
Identify the authoritative and permitted source for each question.
- Step 3
Assess readiness
Review quality, freshness, structure and permissions.
- Step 4
Pilot a bounded domain
Start with an owned topic area and measurable questions.
- Step 5
Feed back gaps
Use unanswered questions to improve the source material.
Example: an assistant for HR policies
Instead of indexing the entire shared drive, the company starts with approved employment policies, official form guidance and a list of responsible teams. Old drafts and individual case files stay out. For pay or leave balances, the application retrieves personal values only after authentication from the HR system. General knowledge and personal data are deliberately separated.
What to remember
Start with a useful, owned and authorised source domain. A knowledge assistant reveals existing information quality; it does not replace it.
Sources and further reading
These primary sources provide further detail on definitions, technical foundations or responsible use.
Content reviewed