[GH-ISSUE #363] URL Scraping for document sourcing #209

Closed
opened 2026-02-22 18:18:21 -05:00 by yindo · 3 comments
Owner

Originally created by @shatfield4 on GitHub (Nov 13, 2023).
Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/363

Originally assigned to: @shatfield4 on GitHub.

A user should be able to input a URL to a website and we should scrape the site for them and embed it for them as a document.

Originally created by @shatfield4 on GitHub (Nov 13, 2023). Original GitHub issue: https://github.com/Mintplex-Labs/anything-llm/issues/363 Originally assigned to: @shatfield4 on GitHub. A user should be able to input a URL to a website and we should scrape the site for them and embed it for them as a document.
yindo added the stage: specificationsfeature request labels 2026-02-22 18:18:21 -05:00
yindo closed this issue 2026-02-22 18:18:21 -05:00
Author
Owner

@cope commented on GitHub (May 10, 2024):

Does this work for protected websites, like for example, those that requires token, or any other, authorization?

@cope commented on GitHub (May 10, 2024): Does this work for protected websites, like for example, those that requires token, or any other, authorization?
Author
Owner

@timothycarambat commented on GitHub (May 10, 2024):

@cope it would not work for any deep web site since that would extremely tedious to setup or configure for any site/portal/auth and no to mention them just blocking headless browsers anyway

@timothycarambat commented on GitHub (May 10, 2024): @cope it would not work for any [deep web site](https://www.google.com/search?q=deep+web+definition) since that would extremely tedious to setup or configure for any site/portal/auth and no to mention them just blocking headless browsers anyway
Author
Owner

@cope commented on GitHub (May 10, 2024):

@timothycarambat gotcha, no worries, was just curious. Once I was able to spin it up I saw the Confluence Data Connector and that is what I was actually looking for.

@cope commented on GitHub (May 10, 2024): @timothycarambat gotcha, no worries, was just curious. Once I was able to spin it up I saw the Confluence Data Connector and that is what I was actually looking for.
yindo changed title from URL Scraping for document sourcing to [GH-ISSUE #363] URL Scraping for document sourcing 2026-06-05 14:34:04 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: Mintplex-Labs/anything-llm#209