What it is
Kreuzberg extracts content from nearly any document format you throw at it—PDFs, Word files, Excel sheets, PowerPoint decks, images (with OCR), emails, HTML, and archives. It's built for AI agents to read files on your behalf so you can ask 'What's in this contract?' or 'Summarize this scanned receipt' without manually opening anything. Perfect for founders, PMs, and operators who need AI to digest documents quickly.
What it does
- Extracts text, tables, metadata, and images from 90+ file formats including PDF, Word, Excel, PowerPoint, images, emails, and HTML
- Runs OCR on scanned PDFs and photos so AI can read text from images or receipts
- Processes batches of files at once—upload a folder of contracts and get all the data out
- Outputs clean Markdown or JSON for easy reading and further AI processing
- Handles password-protected PDFs and multi-language documents
Use it when
- You need AI to read contracts, invoices, or research papers and pull out key info
- You're processing receipts, scanned documents, or photos with text and need OCR
- You want to bulk-extract data from a folder of mixed file types (PDFs, Word docs, spreadsheets)
- You're building a knowledge base and need to ingest documents automatically
Add this skill to your AI
Copy the source URL and ask your AI to add it. Most coding agents — Claude Code, Cursor, Codex, Gemini CLI, OpenCode, Windsurf — will fetch and install it for you.
Try this prompt:
Add this skill to my project: https://github.com/kreuzberg-dev/kreuzberg/tree/main/skills/kreuzberg