The Japanese PDF gets garbled when registering it as knowledge. #7570

Closed
opened 2026-02-21 18:21:14 -05:00 by yindo · 9 comments
Owner

Originally created by @jin413413 on GitHub (Jan 7, 2025).

Originally assigned to: @JohnJyong on GitHub.

Self Checks

  • This is only for bug report, if you would like to ask a question, please head to Discussions.
  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

Dify version

0.15.0

Cloud or Self Hosted

Self Hosted (Docker)

Steps to reproduce

Ver. 0.15.0, Japanese PDF files display garbled text when registered as knowledge and previewed. In 0.14.2, the same files did not have garbled text when registered as knowledge. However, some Japanese PDFs do not display garbled text, so the cause is unknown.
スクリーンショット 2025-01-07 232900

✔️ Expected Behavior

Please ensure that the text is displayed correctly without any garbling.

Actual Behavior

No response

Originally created by @jin413413 on GitHub (Jan 7, 2025). Originally assigned to: @JohnJyong on GitHub. ### Self Checks - [X] This is only for bug report, if you would like to ask a question, please head to [Discussions](https://github.com/langgenius/dify/discussions/categories/general). - [X] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [X] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [X] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [X] Please do not modify this template :) and fill in all the required fields. ### Dify version 0.15.0 ### Cloud or Self Hosted Self Hosted (Docker) ### Steps to reproduce Ver. 0.15.0, Japanese PDF files display garbled text when registered as knowledge and previewed. In 0.14.2, the same files did not have garbled text when registered as knowledge. However, some Japanese PDFs do not display garbled text, so the cause is unknown. ![スクリーンショット 2025-01-07 232900](https://github.com/user-attachments/assets/d928db62-84ca-47fb-8347-4dcb6ff9a401) ### ✔️ Expected Behavior Please ensure that the text is displayed correctly without any garbling. ### ❌ Actual Behavior _No response_
yindo added the 🐞 bug label 2026-02-21 18:21:14 -05:00
yindo closed this issue 2026-02-21 18:21:14 -05:00
Author
Owner

@dosubot[bot] commented on GitHub (Jan 7, 2025):

Hey there, @jin413413! 👋 I'm here to help you with any bugs, questions, or contributions you have for the project. I'll assist you while we wait for a human maintainer.

I found a similar open issue regarding garbled text when using a Japanese PDF as a knowledge base in Dify version 0.6.11. The suggested solution involves ensuring that the text extraction process supports Japanese characters. Here are some steps that might help:

  1. PDF Encoding: Ensure the PDF is properly encoded with Japanese text.
  2. Text Extraction Tool: Verify that the text extraction tool or library used by Dify supports Japanese characters.

Additionally, you can adjust configuration settings within Dify to improve text extraction accuracy for Japanese PDFs. Specifically, you can use custom segmentation and preprocessing rules:

  • Segmentation Rules:

    • Set a segmentation identifier such as "\n".
    • Define the maximum segment length.
    • Set the number of overlapping characters between segments.
  • Preprocessing Rules:

    • Replace continuous spaces, newlines, and tabs.
    • Delete all URLs and email addresses.

For enhanced preprocessing capabilities, consider using optional ETL solutions like Dify ETL or Unstructured ETL [1].

To continue talking to Dosu, mention @dosu.


Help Dosu learn! Give it feedback: Great Response | Irrelevant Answer | Incorrect Sources | Too Verbose | Hallucination | Other

@dosubot[bot] commented on GitHub (Jan 7, 2025): <!-- Greeting --> Hey there, @jin413413! 👋 I'm here to help you with any bugs, questions, or contributions you have for the project. I'll assist you while we wait for a human maintainer. <!-- Answer --> I found a similar open issue regarding garbled text when using a Japanese PDF as a knowledge base in Dify version 0.6.11. The suggested solution involves ensuring that the text extraction process supports Japanese characters. Here are some steps that might help: 1. **PDF Encoding**: Ensure the PDF is properly encoded with Japanese text. 2. **Text Extraction Tool**: Verify that the text extraction tool or library used by Dify supports Japanese characters. Additionally, you can adjust configuration settings within Dify to improve text extraction accuracy for Japanese PDFs. Specifically, you can use custom segmentation and preprocessing rules: - **Segmentation Rules**: - Set a segmentation identifier such as "\n". - Define the maximum segment length. - Set the number of overlapping characters between segments. - **Preprocessing Rules**: - Replace continuous spaces, newlines, and tabs. - Delete all URLs and email addresses. For enhanced preprocessing capabilities, consider using optional ETL solutions like **Dify ETL** or [**Unstructured ETL**](https://unstructured.io/) <sup>[[1]](https://github.com/langgenius/dify/issues/6012)</sup>. <!-- Dosu Comment Footer --> *To continue talking to [Dosu](https://dosu.dev), mention @dosu.* --- <sup>Help Dosu learn! Give it feedback: [Great Response](https://app.dosu.dev/response-feedback/4f0e5434-bb4d-4f95-9b9a-1c94a7f45a49?feedback_type=great_response) | [Irrelevant Answer](https://app.dosu.dev/response-feedback/4f0e5434-bb4d-4f95-9b9a-1c94a7f45a49?feedback_type=irrelevant_answer) | [Incorrect Sources](https://app.dosu.dev/response-feedback/4f0e5434-bb4d-4f95-9b9a-1c94a7f45a49?feedback_type=incorrect_sources) | [Too Verbose](https://app.dosu.dev/response-feedback/4f0e5434-bb4d-4f95-9b9a-1c94a7f45a49?feedback_type=too_verbose) | [Hallucination](https://app.dosu.dev/response-feedback/4f0e5434-bb4d-4f95-9b9a-1c94a7f45a49?feedback_type=hallucination) | [Other](https://app.dosu.dev/response-feedback/4f0e5434-bb4d-4f95-9b9a-1c94a7f45a49?feedback_type=other)</sup>
Author
Owner

@crazywoola commented on GitHub (Jan 7, 2025):

Can you provide some example PDF files for testing?

@crazywoola commented on GitHub (Jan 7, 2025): Can you provide some example PDF files for testing?
Author
Owner

@jin413413 commented on GitHub (Jan 8, 2025):

Although version 0.15.0 could read many Japanese PDF files, the PDF files on this specific site appear as gibberish when previewed in the knowledge feature.

https://www.nta.go.jp/publication/pamph/gensen/0024001-021.pdf

@jin413413 commented on GitHub (Jan 8, 2025): Although version 0.15.0 could read many Japanese PDF files, the PDF files on this specific site appear as gibberish when previewed in the knowledge feature. https://www.nta.go.jp/publication/pamph/gensen/0024001-021.pdf
Author
Owner

@XiaoBa-Yu commented on GitHub (Jan 13, 2025):

It may be a language problem
image
image

@XiaoBa-Yu commented on GitHub (Jan 13, 2025): It may be a language problem ![image](https://github.com/user-attachments/assets/28ff7fcc-8d65-4fd0-9328-4e41ccc26d15) ![image](https://github.com/user-attachments/assets/879560d4-a9ce-4f29-a098-4b2cb7539697)
Author
Owner

@nb75km commented on GitHub (Feb 12, 2025):

I don't know if this will be helpful, but...
After experimenting with PDFs created using Python’s ReportLab, I found that when I explicitly registered and embedded a TrueType (TTF) font in the PDF, it did not work correctly. However, when I embedded the text using a CID font, the knowledge registration worked correctly.

Below is the sample code.

• Code that worked correctly

from reportlab.pdfgen import canvas
from reportlab.pdfbase.cidfonts import UnicodeCIDFont
from reportlab.pdfbase import pdfmetrics

pdfmetrics.registerFont(UnicodeCIDFont('HeiseiKakuGo-W5'))

def create_pdf_cid():
    c = canvas.Canvas("test_japanese_cid.pdf")
    c.setFont('HeiseiKakuGo-W5', 12)
    c.drawString(100, 750, "こんにちは、世界!")
    c.showPage()
    c.save()

create_pdf_cid()

• Code that did not work correctly

from reportlab.pdfgen import canvas
from reportlab.pdfbase import pdfmetrics
from reportlab.pdfbase.ttfonts import TTFont

# 1. Register the font
pdfmetrics.registerFont(TTFont('MyJapaneseFont', '/path/to/your_japanese_font.ttf'))

def create_pdf():
    c = canvas.Canvas("test_japanese.pdf")
    
    # 2. Set the registered font
    c.setFont('MyJapaneseFont', 12)
    
    # 3. Draw a Japanese string (Unicode string)
    c.drawString(100, 750, "こんにちは、世界!")

    c.showPage()
    c.save()

create_pdf()
@nb75km commented on GitHub (Feb 12, 2025): I don't know if this will be helpful, but... After experimenting with PDFs created using Python’s ReportLab, I found that when I explicitly registered and embedded a TrueType (TTF) font in the PDF, it did not work correctly. However, when I embedded the text using a CID font, the knowledge registration worked correctly. Below is the sample code. **• Code that worked correctly** ``` from reportlab.pdfgen import canvas from reportlab.pdfbase.cidfonts import UnicodeCIDFont from reportlab.pdfbase import pdfmetrics pdfmetrics.registerFont(UnicodeCIDFont('HeiseiKakuGo-W5')) def create_pdf_cid(): c = canvas.Canvas("test_japanese_cid.pdf") c.setFont('HeiseiKakuGo-W5', 12) c.drawString(100, 750, "こんにちは、世界!") c.showPage() c.save() create_pdf_cid() ``` **• Code that did not work correctly** ``` from reportlab.pdfgen import canvas from reportlab.pdfbase import pdfmetrics from reportlab.pdfbase.ttfonts import TTFont # 1. Register the font pdfmetrics.registerFont(TTFont('MyJapaneseFont', '/path/to/your_japanese_font.ttf')) def create_pdf(): c = canvas.Canvas("test_japanese.pdf") # 2. Set the registered font c.setFont('MyJapaneseFont', 12) # 3. Draw a Japanese string (Unicode string) c.drawString(100, 750, "こんにちは、世界!") c.showPage() c.save() create_pdf() ```
Author
Owner

@iwaim commented on GitHub (Feb 14, 2025):

This problem also occurs in "Doc Extractor" block. I confirmed that it does not occur if I change back to "pypdfium2 ver.4.30.0".

Image

sample file: https://www.nta.go.jp/publication/pamph/gensen/0024001-021.pdf
Dify version: 0.15.3 with Self Hosted (Docker)

@iwaim commented on GitHub (Feb 14, 2025): This problem also occurs in "Doc Extractor" block. I confirmed that it does not occur if I change back to "pypdfium2 ver.4.30.0". ![Image](https://github.com/user-attachments/assets/8cc68bd1-2da4-4e10-9e53-ddeed042a542) sample file: https://www.nta.go.jp/publication/pamph/gensen/0024001-021.pdf Dify version: 0.15.3 with Self Hosted (Docker)
Author
Owner

@iwaim commented on GitHub (Feb 14, 2025):

sample DSL file.

app:
  description: ''
  icon: 🤖
  icon_background: '#FFEAD5'
  mode: workflow
  name: Japanese_PDF
  use_icon_as_answer_icon: false
kind: app
version: 0.1.5
workflow:
  conversation_variables: []
  environment_variables: []
  features:
    file_upload:
      allowed_file_extensions:
      - .JPG
      - .JPEG
      - .PNG
      - .GIF
      - .WEBP
      - .SVG
      allowed_file_types:
      - image
      allowed_file_upload_methods:
      - local_file
      - remote_url
      enabled: false
      fileUploadConfig:
        audio_file_size_limit: 50
        batch_count_limit: 5
        file_size_limit: 15
        image_file_size_limit: 10
        video_file_size_limit: 100
        workflow_file_upload_limit: 10
      image:
        enabled: false
        number_limits: 3
        transfer_methods:
        - local_file
        - remote_url
      number_limits: 3
    opening_statement: ''
    retriever_resource:
      enabled: true
    sensitive_word_avoidance:
      enabled: false
    speech_to_text:
      enabled: false
    suggested_questions: []
    suggested_questions_after_answer:
      enabled: false
    text_to_speech:
      enabled: false
      language: ''
      voice: ''
  graph:
    edges:
    - data:
        isInIteration: false
        sourceType: start
        targetType: document-extractor
      id: 1739520066113-source-1739520083193-target
      source: '1739520066113'
      sourceHandle: source
      target: '1739520083193'
      targetHandle: target
      type: custom
      zIndex: 0
    - data:
        isInIteration: false
        sourceType: document-extractor
        targetType: end
      id: 1739520083193-source-1739520115807-target
      source: '1739520083193'
      sourceHandle: source
      target: '1739520115807'
      targetHandle: target
      type: custom
      zIndex: 0
    nodes:
    - data:
        desc: ''
        selected: false
        title: 開始
        type: start
        variables:
        - allowed_file_extensions: []
          allowed_file_types:
          - document
          allowed_file_upload_methods:
          - local_file
          label: f
          max_length: 48
          options: []
          required: true
          type: file
          variable: f
      height: 90
      id: '1739520066113'
      position:
        x: 80
        y: 282
      positionAbsolute:
        x: 80
        y: 282
      selected: false
      sourcePosition: right
      targetPosition: left
      type: custom
      width: 244
    - data:
        desc: ''
        is_array_file: false
        selected: false
        title: テキスト抽出ツール
        type: document-extractor
        variable_selector:
        - '1739520066113'
        - f
      height: 92
      id: '1739520083193'
      position:
        x: 80
        y: 415
      positionAbsolute:
        x: 80
        y: 415
      selected: false
      sourcePosition: right
      targetPosition: left
      type: custom
      width: 244
    - data:
        desc: ''
        outputs:
        - value_selector:
          - '1739520083193'
          - text
          variable: text
        selected: false
        title: 終了
        type: end
      height: 90
      id: '1739520115807'
      position:
        x: 80
        y: 543
      positionAbsolute:
        x: 80
        y: 543
      selected: true
      sourcePosition: right
      targetPosition: left
      type: custom
      width: 244
    viewport:
      x: 364
      y: -224.5
      zoom: 1
@iwaim commented on GitHub (Feb 14, 2025): sample DSL file. ```yaml app: description: '' icon: 🤖 icon_background: '#FFEAD5' mode: workflow name: Japanese_PDF use_icon_as_answer_icon: false kind: app version: 0.1.5 workflow: conversation_variables: [] environment_variables: [] features: file_upload: allowed_file_extensions: - .JPG - .JPEG - .PNG - .GIF - .WEBP - .SVG allowed_file_types: - image allowed_file_upload_methods: - local_file - remote_url enabled: false fileUploadConfig: audio_file_size_limit: 50 batch_count_limit: 5 file_size_limit: 15 image_file_size_limit: 10 video_file_size_limit: 100 workflow_file_upload_limit: 10 image: enabled: false number_limits: 3 transfer_methods: - local_file - remote_url number_limits: 3 opening_statement: '' retriever_resource: enabled: true sensitive_word_avoidance: enabled: false speech_to_text: enabled: false suggested_questions: [] suggested_questions_after_answer: enabled: false text_to_speech: enabled: false language: '' voice: '' graph: edges: - data: isInIteration: false sourceType: start targetType: document-extractor id: 1739520066113-source-1739520083193-target source: '1739520066113' sourceHandle: source target: '1739520083193' targetHandle: target type: custom zIndex: 0 - data: isInIteration: false sourceType: document-extractor targetType: end id: 1739520083193-source-1739520115807-target source: '1739520083193' sourceHandle: source target: '1739520115807' targetHandle: target type: custom zIndex: 0 nodes: - data: desc: '' selected: false title: 開始 type: start variables: - allowed_file_extensions: [] allowed_file_types: - document allowed_file_upload_methods: - local_file label: f max_length: 48 options: [] required: true type: file variable: f height: 90 id: '1739520066113' position: x: 80 y: 282 positionAbsolute: x: 80 y: 282 selected: false sourcePosition: right targetPosition: left type: custom width: 244 - data: desc: '' is_array_file: false selected: false title: テキスト抽出ツール type: document-extractor variable_selector: - '1739520066113' - f height: 92 id: '1739520083193' position: x: 80 y: 415 positionAbsolute: x: 80 y: 415 selected: false sourcePosition: right targetPosition: left type: custom width: 244 - data: desc: '' outputs: - value_selector: - '1739520083193' - text variable: text selected: false title: 終了 type: end height: 90 id: '1739520115807' position: x: 80 y: 543 positionAbsolute: x: 80 y: 543 selected: true sourcePosition: right targetPosition: left type: custom width: 244 viewport: x: 364 y: -224.5 zoom: 1 ```
Author
Owner

@iwaim commented on GitHub (Feb 14, 2025):

The text via copy and paste is correct.

Image

@iwaim commented on GitHub (Feb 14, 2025): The text via copy and paste is correct. ![Image](https://github.com/user-attachments/assets/c0d58d5f-fd76-4d5a-b901-adc64d37c195)
Author
Owner

@dosubot[bot] commented on GitHub (Mar 17, 2025):

Hi, @jin413413. I'm Dosu, and I'm helping the Dify team manage their backlog. I'm marking this issue as stale.

Issue Summary:

  • Japanese PDF files become garbled in version 0.15.0, but not in 0.14.2.
  • Suggested solutions include checking PDF encoding and using tools supporting Japanese characters.
  • Example PDFs were requested for testing.
  • Insights on font embedding were shared, suggesting CID fonts over TrueType fonts.
  • The issue also affects the "Doc Extractor" block; reverting to "pypdfium2 ver.4.30.0" was suggested.

Next Steps:

  • Please confirm if this issue is still relevant to the latest version of Dify. If so, you can keep the discussion open by commenting here.
  • If there is no further activity, this issue will be automatically closed in 15 days.

Thank you for your understanding and contribution!

@dosubot[bot] commented on GitHub (Mar 17, 2025): Hi, @jin413413. I'm [Dosu](https://dosu.dev), and I'm helping the Dify team manage their backlog. I'm marking this issue as stale. **Issue Summary:** - Japanese PDF files become garbled in version 0.15.0, but not in 0.14.2. - Suggested solutions include checking PDF encoding and using tools supporting Japanese characters. - Example PDFs were requested for testing. - Insights on font embedding were shared, suggesting CID fonts over TrueType fonts. - The issue also affects the "Doc Extractor" block; reverting to "pypdfium2 ver.4.30.0" was suggested. **Next Steps:** - Please confirm if this issue is still relevant to the latest version of Dify. If so, you can keep the discussion open by commenting here. - If there is no further activity, this issue will be automatically closed in 15 days. Thank you for your understanding and contribution!
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#7570