feat(coding-agent): replaced markitdown CLI with native markit-ai library for document conversion

- Replaced Python-based markitdown CLI with native markit-ai library for document and notebook conversion.
- Added support for converting Jupyter notebooks (.ipynb) to markdown via markit-ai integration.
- Created markit utility module with file/buffer conversion wrappers and standardized error handling.
- Removed markitdown from Python tools manager and eliminated external CLI dependency.
- Updated fetch, read, and scraper tools to use new markit conversion API across codebase.
- Added test coverage for ipynb file conversion through markit utility.
This commit is contained in:
can1357
2026-04-01 20:51:29 +02:00
parent 9d5508980f
commit d8f5bb4be2
14 changed files with 297 additions and 125 deletions
+2 -2
View File
@@ -1006,7 +1006,7 @@ Core pipeline in `renderUrl(url, timeout, raw, signal)`:
1. Normalize URL (`normalizeUrl`) and short-circuit `pi-internal://`.
2. If not `raw`, run `handleSpecialUrls(...)` over `specialHandlers` from `../web/scrapers`.
3. Fetch with `loadPage(...)`.
4. Detect convertible binaries (`isConvertible`) and attempt `fetchBinary(...)` + `convertWithMarkitdown(...)`.
4. Detect convertible binaries (`isConvertible`) and attempt `fetchBinary(...)` + `convertWithMarkit(...)`.
5. Handle structured/non-HTML content directly:
- JSON: `formatJson`
- RSS/Atom/XML feed: `parseFeedToMarkdown`
@@ -1017,7 +1017,7 @@ Core pipeline in `renderUrl(url, timeout, raw, signal)`:
- `/.well-known/llms.txt`, `/llms.txt`, `/llms.md` (`tryLlmEndpoints`)
- `Accept: text/markdown, text/plain...` negotiation (`tryContentNegotiation`)
- HTML-to-text conversion (`renderHtmlToText`)
7. If HTML conversion is low-quality (`isLowQualityOutput`), try document-link extraction (`extractDocumentLinks`) + markitdown.
7. If HTML conversion is low-quality (`isLowQualityOutput`), try document-link extraction (`extractDocumentLinks`) + markit.
`renderHtmlToText(...)` fallback order is explicit in code: Jina reader endpoint (`https://r.jina.ai/<url>`), then `trafilatura` (via `ensureTool`), then `lynx`, then native `htmlToMarkdown`.