support full Firecrawl crawlerOptions and pageOptions when add a url document in datasets page when Extract web content with 🔥Firecrawl #5308

Closed
opened 2026-02-21 18:10:21 -05:00 by yindo · 0 comments
Owner

Originally created by @dolonfly on GitHub (Aug 29, 2024).

Self Checks

  • I have searched for existing issues search for existing issues, including closed ones.
  • I confirm that I am using English to submit this report (我已阅读并同意 Language Policy).
  • [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:)
  • Please do not modify this template :) and fill in all the required fields.

1. Is this request related to a challenge you're experiencing? Tell me about your story.

support full Firecrawl crawlerOptions and pageOptions when add a url document in datasets page

Refer to the Firecrawl api reference for detailed information.

Missing Options List:

  • crawlerOptions:

    • ignoreSitemap
    • allowBackwardCrawling
    • allowExternalContentLinks
    • limit
  • pageOptions:

    • onlyIncludeTags
    • removeTags
image

2. Additional context or comments

in my use case: when use removeTags option can clean lot unnecessary html tag and convert 2 markdown cleaner

3. Can you help us with this feature?

  • I am interested in contributing to this feature.
Originally created by @dolonfly on GitHub (Aug 29, 2024). ### Self Checks - [X] I have searched for existing issues [search for existing issues](https://github.com/langgenius/dify/issues), including closed ones. - [X] I confirm that I am using English to submit this report (我已阅读并同意 [Language Policy](https://github.com/langgenius/dify/issues/1542)). - [X] [FOR CHINESE USERS] 请务必使用英文提交 Issue,否则会被关闭。谢谢!:) - [X] Please do not modify this template :) and fill in all the required fields. ### 1. Is this request related to a challenge you're experiencing? Tell me about your story. support full Firecrawl crawlerOptions and pageOptions when add a url document in datasets page Refer to the [Firecrawl api reference](https://docs.firecrawl.dev/api-reference/endpoint/crawl-post) for detailed information. ### Missing Options List: - **`crawlerOptions`**: - `ignoreSitemap` - `allowBackwardCrawling` - `allowExternalContentLinks` - `limit` - **`pageOptions`**: - `onlyIncludeTags` - `removeTags` <img width="1190" alt="image" src="https://github.com/user-attachments/assets/fc1ca67e-c012-4759-bf52-390e3d13d9d4"> ### 2. Additional context or comments in my use case: when use `removeTags` option can clean lot unnecessary html tag and convert 2 markdown cleaner ### 3. Can you help us with this feature? - [X] I am interested in contributing to this feature.
yindo added the 💪 enhancement label 2026-02-21 18:10:21 -05:00
yindo closed this issue 2026-02-21 18:10:21 -05:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: langgenius/dify#5308