LiteParse
by LlamaIndex
Parses PDFs, Office files, and images locally into markdown, JSON with bounding boxes, or plain text for agents and RAG pipelines, from a CLI or Rust, Node.js, Python, and WASM libraries.
Skills
Markdown Output
Renders documents to structured Markdown with headings, tables, lists, images, and links, ready to pass to an LLM or RAG pipeline.
Selective OCR
Runs bundled Tesseract OCR only where native text is missing, or sends pages to an HTTP OCR server such as EasyOCR or PaddleOCR.
Complexity Check
Checks cheaply whether a document needs OCR or heavier parsing, so a pipeline can route, reject, or estimate cost before a full parse.
Page Screenshots
Generates PNG screenshots of all or selected pages at a custom DPI so vision-capable agents can look at the rendered page itself.
Batch And Worker Parsing
Parses whole directories with lit batch-parse, or many files in Python or Node.js worker pools with hard per-parse timeouts.
Related Agents
Latitude
Traces, scores, and evaluates production AI agents via OpenTelemetry and SDKs, groups failures into signals, and dispat…
MinerU
Converts PDFs, Office documents, and scanned pages into LLM-ready Markdown or JSON, preserving tables, formulas, and re…
AgentOps for Apify Builders Bundle
Apify actor bundle for agent builders: normalize run traces for QA, control costs, and guard tool calls with a firewall…
AnythingLLM
Runs an all-in-one, self-hosted RAG and agents app: chat with documents, build agents, and connect tools via MCP. Deskt…