Document Processing
In short:This is Cherry Studio's central configuration for "recognizing text in PDFs / images / scans".
For example, the following all depend on it:
You drag a scanned contract PDF into the chat box and want the AI to understand its contents
You put a pile of invoice images intoKnowledge Base, hoping to be able to search them later
your Agent needs to open a screenshot in the local folder for analysis
Behind these scenarios, you first need to turn "text in images" into "text the AI can read"; this step is technically called OCR(Optical Character Recognition, optical character recognition).
Cherry Studio centralizes OCR configuration ina single settings page: configure it once here, and all places that use OCR will share the same configuration.
Configuration entry
Open Settings → Document Processing:

The panel is divided into two parts, handling "image text recognition" and "PDF parsing" respectively.
1. OCR Service — recognize text from images
Applicable to: images (screenshots, scans), content that must be recognized as text before AI can read it.
macOS: choose "System OCR" and you're good to go,no configuration required, using the system's built-in image recognition capability, offline and free ✅
Windows: choose "System OCR" out of the box; if you need to recognize languages other than English/Chinese, you need to download the corresponding language pack in Windows
Linux / Advanced: optional Tesseract, Paddle OCR, OpenVINO, etc.
OCR engine comparison
System OCR
The simplest, no configuration, usually good enough
Tesseract
Classic open-source OCR, built into Cherry Studio, supports custom languages
Paddle OCR
Better Chinese recognition (open-sourced by Baidu), requires "Star River Community access token + API URL"
OpenVINO
Intel graphics can accelerate it
If unsure, use the default System OCR; switch only if recognition is poor.
2. Document Processing provider — structured parsing for PDFs / complex documents
Applicable to: PDFs with tables / multi-column layouts / scanned pages, long documents. Ordinary plain-text PDFs can be read directly, no need to go through here.
MinerU(default)
Free cloud service, specialized in complex-layout PDFs (academic papers, contracts, etc.), requires going to mineru.net to register and obtain an API Key
Paddle OCR
Offline solution, requires configuring a Star River Community access token
Third-party provider
Uses the vision model of a configured AI service provider to recognize them (smarter results but paid)
Configure MinerU (default option)
in API Key Fill in the key obtained from MinerU in the field
API Host Keep default
https://mineru.netWhen switching to the knowledge base or Agent, no additional configuration is needed; it will automatically use the settings here
Relationship with the knowledge base
Document processing only handles the step "non-text → text"
The converted text then continues to embedding model vectorization, storage into the database
See the detailed "enable in knowledge base" process at Knowledge base document preprocessing
When configuration is not needed
You only import plain text into the knowledge base (
.md/.txt/.docxplain text paragraphs in) → does not go through document processing at allYou only use chat, without uploading files → same as above
Tips and tricks
MinerU is significantly better than Tesseract for PDFs with tables / multi-column layouts; it is the first choice for academic papers and the like
For offline scenarios, use Paddle OCR or Tesseract (works without internet)
After switching processors, previously vectorized materials will not be automatically reprocessed — need to re-import manually
💡 Get help and submit feedback
If you encounter any questions, bugs, or have suggestions for feature improvements during configuration or use, please refer to Feedback and Suggestions the official channels provided there.
Last updated
Was this helpful?