> For the complete documentation index, see [llms.txt](https://docs.cherryai.com.cn/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.cherryai.com.cn/docs/zhong-wen-fan-ti/knowledge-base/sources.md).

# 加入同整理資料

知識庫支援文件、Cherry Studio 筆記、本地目錄同單個網頁地址。導入之後仲要檢查處理狀態、正文同 Chunks，仲有資料更新時重新索引。

{% hint style="info" %}
完成標準唔係「文件已經出現喺列表入面」，而係資料可以讀、Chunks 完整，而且真實問題可以召回到正確來源。
{% endhint %}

## 揀正確入口

<figure><img src="/files/024a74247461ed4b2b0089f937db2f9f5571906b" alt="知识库中的文件、笔记、目录和链接四种资料入口"><figcaption><p>按資料來源揀入口：少量文件用【文件】，同類文件集合用【目錄】，Cherry Studio 內容用【筆記】，公開網頁用【連結】。</p></figcaption></figure>

| 入口 | 適合咩資料                    | 導入後嘅關係      | 主要注意事項               |
| -- | ------------------------ | ----------- | -------------------- |
| 檔案 | PDF、Office、Markdown、文本等  | 保存託管副本      | 單次最多揀 20 項           |
| 筆記 | Cherry Studio 入面已經整理好嘅內容 | 導入當時嘅內容快照   | 原筆記之後改咗都唔會自動同步       |
| 目錄 | 同一主題底下一批本地文件             | 按目錄內容建立資料條目 | 唔好將唔相關目錄成個導入         |
| 連結 | 單個可以公開訪問嘅網頁              | 保存抓取時嘅網頁快照  | 登入頁、腳本渲染或者受限制網頁可能唔完整 |

{% hint style="warning" %}
支援嘅文件包括 PDF、DOCX、DOC、PPTX、XLSX、XLS、MD、TXT、CSV、HTML 同 EPUB。掃描版 PDF 或圖片型內容仲需要檢查 OCR。
{% endhint %}

## 添加同驗收資料

{% stepper %}
{% step %}

### 1. 揀資料來源

打開知識庫，撳添加資料按鈕，揀【文件】、【筆記】、【目錄】或【連結】。
{% endstep %}

{% step %}

### 2. 確認揀中嘅內容

文件同筆記可以批量揀；單次互動式添加最多 20 項。資料多嘅時候分批添加，或者用目錄入口。
{% endstep %}

{% step %}

### 3. 處理同名衝突

新資料同現有條目同名時，揀【全部保留】或者【替換】。更新制度、手冊同筆記快照時通常揀【替換】。

{% hint style="warning" %}
揀【全部保留】會令新舊內容同時參與召回。只係真係需要並行查詢唔同版本時先咁做，並且喺名稱入面標明日期或者版本。
{% endhint %}
{% endstep %}

{% step %}

### 4. 等待處理完成

資料會經過複製、讀取、切分同索引等階段。未配置嵌入模型時唔會建立向量，但仍然會建立關鍵詞索引。

<figure><img src="/files/605329854681bd110f8a7a74d7e588d92af4cf74" alt="包含多条已处理资料的员工差旅制度知识库"><figcaption><p>資料進入可用狀態之後，再抽查正文同 Chunks。</p></figcaption></figure>
{% endstep %}

{% step %}

### 5. 抽查正文同 Chunks

打開資料睇正文，或者喺資料行菜單睇 Chunks。重點檢查標題順序、表格、OCR 文字同關鍵句有冇被錯誤拆開。
{% endstep %}

{% step %}

### 6. 完成召回測試

用一條答案明確嘅問題檢查正確來源同片段。資料更新之後，都要用同一組問題重新測試。

<figure><img src="/files/65ce03c7ec8a7c321ec30cb89dc17f5b9a06494e" alt="召回测试中的来源、相关度、片段内容和排序"><figcaption><p>最終驗收要睇來源、片段完整度同排序，唔係只睇有冇返回結果。</p></figcaption></figure>
{% endstep %}
{% endstepper %}

## 資料狀態同處理方法

| 現象          | 可能原因                | 處理方法               |
| ----------- | ------------------- | ------------------ |
| 長時間處理中      | 文件較大、解析器或者模型唔可用     | 檢查原文件、文檔處理同嵌入模型    |
| 顯示錯誤        | 複製、讀取、切分或者索引失敗      | 打開錯誤信息，按失敗階段處理     |
| 正文缺失或者亂碼    | 文件處理器唔適配、掃描內容未做 OCR | 更換文檔處理方式或者配置 OCR   |
| Chunk 缺少關鍵句 | 分塊邊界唔合適             | 調整分塊之後執行【重新索引】     |
| 新舊版本同時命中    | 同名資料揀咗【全部保留】        | 刪除舊條目，或者重新導入再揀【替換】 |

## 重新索引同刪除

切分、解析器或者嵌入設定改咗之後，舊條目唔會自動套用新設定。對單條資料用【重新索引】，或者批量揀資料後重新索引。

{% hint style="danger" %}
刪除條目會移除而家知識庫入面嘅託管副本同索引。佢唔會刪除原始文件或者原筆記，但刪除之前仍然要確認知識庫入面係咪有唯一副本。
{% endhint %}

## 配置說明

| 配置項    | 產品預設值   | 建議起點       | 作用              | 適用場景       | 注意事項              |
| ------ | ------- | ---------- | --------------- | ---------- | ----------------- |
| 單次添加數量 | 最多 20 項 | 先添加少量代表性資料 | 控制一次導入規模        | 第一次建庫或者排錯  | 大批量導入前先驗證解析同召回    |
| 同名處理   | 發生衝突時揀  | 更新資料優先【替換】 | 決定新舊條目會唔會並存     | 制度、手冊、筆記更新 | 【全部保留】可能會令舊內容參與召回 |
| 重新索引   | 手動執行    | 設定改咗之後執行   | 令舊資料用新解析、分塊或者模型 | 調優或者修復資料   | 完成之後一定要重新做召回測試    |

## 用戶案例

小林每月更新差旅制度。佢將新文件用相同名稱導入，並揀【替換】，等資料處理完成後抽查正文同 Chunks，再用固定問題測試住宿、交通同審批規則。

完成標準係：舊規則唔再出現喺召回結果入面，新規則嘅條件同金額可以穩定命中。

## 常見問題

<details>

<summary>修改原筆記後，知識庫會自動更新嗎？</summary>

唔會。筆記導入嘅係當時內容嘅快照。修改之後需要重新添加並揀【替換】，或者對對應資料執行【重新索引】。

</details>

<details>

<summary>網頁點解只抓到部分內容？</summary>

需要登入、依賴腳本渲染或者有訪問限制嘅網頁可能無法完整抓取。可以改用文件或者筆記保存正文之後再導入。

</details>

<details>

<summary>刪除知識庫條目會刪除原文件嗎？</summary>

唔會刪除原始文件或者原筆記，但會移除知識庫入面嘅託管副本同索引。

</details>

## 繼續閱讀

<table data-view="cards"><thead><tr><th></th><th></th><th data-hidden data-card-target data-type="content-ref"></th></tr></thead><tbody><tr><td><strong>文件解析同 OCR</strong></td><td>處理正文缺失、亂碼同掃描內容。</td><td><a href="/pages/4b9563c434830d8aac5df782c734b2ddcd8509f0">/pages/4b9563c434830d8aac5df782c734b2ddcd8509f0</a></td></tr><tr><td><strong>檢查資料同召回</strong></td><td>用固定問題驗收檢索質量。</td><td><a href="/pages/c63f889f83b9213cb096e7e37b06ede1b3b4a986">/pages/c63f889f83b9213cb096e7e37b06ede1b3b4a986</a></td></tr><tr><td><strong>數據、私隱同維護</strong></td><td>了解備份、刪除同服務邊界。</td><td><a href="/pages/c717d714b290fe2e0838ea72d335a09917e42a89">/pages/c717d714b290fe2e0838ea72d335a09917e42a89</a></td></tr></tbody></table>


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.cherryai.com.cn/docs/zhong-wen-fan-ti/knowledge-base/sources.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
