|
|
|
@@ -27,25 +27,25 @@ They can be divided into two groups.
|
|
|
|
|
|
|
|
|
|
- `apiKey` is required. Can be set as an environment variable `LLAMA_CLOUD_API_KEY`
|
|
|
|
|
- `checkInterval` is the interval in seconds to check if the parsing is done. Default is `1`.
|
|
|
|
|
- `maxTimeout` is the maximum timout to wait for parsing to finish. Default is `2000`
|
|
|
|
|
- `maxTimeout` is the maximum timeout to wait for parsing to finish. Default is `2000`
|
|
|
|
|
- `verbose` shows progress of the parsing. Default is `true`
|
|
|
|
|
- `ignoreErrors` set to false to get errors while parsing. Default is `true` and returns an empty array on error.
|
|
|
|
|
|
|
|
|
|
#### Advanced params:
|
|
|
|
|
|
|
|
|
|
- `resultType` can be set to `markdown`, `text` or `json`. Defaults to `text`. More information about `json` mode on the next pages.
|
|
|
|
|
- `language` primarly helps with OCR recognition. Defaults to `en`. Click [here](../../../api/type-aliases/Language.md) for a list of supported languages.
|
|
|
|
|
- `language` primarily helps with OCR recognition. Defaults to `en`. Click [here](../../../api/type-aliases/Language.md) for a list of supported languages.
|
|
|
|
|
- `parsingInstructions?` Optional. Can help with complicated document structures. See this [LlamaIndex Blog Post](https://www.llamaindex.ai/blog/launching-the-first-genai-native-document-parsing-platform) for an example.
|
|
|
|
|
- `skipDiagonalText?` Optional. Set to true to ignore diagonal text. (Text that is not rotated 0, 90, 180 or 270 degrees)
|
|
|
|
|
- `invalidateCache?` Optional. Set to true to ignore the LlamaCloud cache. All document are kept in cache for 48hours after the job was completed to avoid processing the same document twice. Can be useful for testing when trying to re-parse the same document with, e.g. different `parsingInstructions`.
|
|
|
|
|
- `doNotCache?` Optional. Set to true to not cache the document.
|
|
|
|
|
- `fastMode?` Optional. Set to true to use the fast mode. This mode will skip OCR of images, and table/heading reconstruction. Note: Non-compatible with `gpt4oMode`.
|
|
|
|
|
- `doNotUnrollColumns?` Optional. Set to true to keep the text according to document layout. Reduce reconstruction accuracy, and LLM's/embedings performances in most cases.
|
|
|
|
|
- `pageSeperator?` Optional. The page seperator to use. Defaults is `\\n---\\n`.
|
|
|
|
|
- `doNotUnrollColumns?` Optional. Set to true to keep the text according to document layout. Reduce reconstruction accuracy, and LLMs/embeddings performances in most cases.
|
|
|
|
|
- `pageSeparator?` Optional. The page separator to use. Defaults is `\\n---\\n`.
|
|
|
|
|
- `gpt4oMode` set to true to use GPT-4o to extract content. Default is `false`.
|
|
|
|
|
- `gpt4oApiKey?` Optional. Set the GPT-4o API key. Lowers the cost of parsing by using your own API key. Your OpenAI account will be charged. Can also be set in the environment variable `LLAMA_CLOUD_GPT4O_API_KEY`.
|
|
|
|
|
- `boundingBox?` Optional. Specify an area of the document to parse. Expects the bounding box margins as a string in clockwise order, e.g. `boundingBox = "0.1,0,0,0"` to not parse the top 10% of the document.
|
|
|
|
|
- `targetPages?` Optional. Specify which pages to parse by specifying them as a comma-seperated list. First page is `0`.
|
|
|
|
|
- `targetPages?` Optional. Specify which pages to parse by specifying them as a comma-separated list. First page is `0`.
|
|
|
|
|
- `numWorkers` as in the python version, is set in `SimpleDirectoryReader`. Default is 1.
|
|
|
|
|
|
|
|
|
|
### LlamaParse with SimpleDirectoryReader
|
|
|
|
|