A Python tool to extract highlights and annotations from PDF files, particularly useful for PDFs highlighted in Apple Books.
- Extract highlights, underlines, and other annotations from PDF files
- Display highlighted text with page numbers
- Include annotation notes/comments
- Export to text or Markdown format
- Save to file or display in console
- Python 3.6+
- pymupdf (PyMuPDF)
- Clone or download this repository
- Install dependencies:
pip install pymupdfOr if using the included virtual environment:
source .venv/bin/activate # On macOS/Linux
pip install pymupdfpython extract_highlights.py path/to/your/file.pdfpython extract_highlights.py path/to/your/file.pdf -o highlights.txtpython extract_highlights.py path/to/your/file.pdf -o highlights.md -f markdownpdf_file- Path to the PDF file (required)-o, --output- Output file path (optional, prints to console if not specified)-f, --format- Output format:textormarkdown(default: text)
Extract highlights and display them:
python extract_highlights.py ~/Documents/mybook.pdfSave highlights to a text file:
python extract_highlights.py ~/Documents/mybook.pdf -o mybook_highlights.txtSave highlights as Markdown:
python extract_highlights.py ~/Documents/mybook.pdf -o mybook_highlights.md -f markdownThe tool uses PyMuPDF (fitz) to:
- Open the PDF file
- Iterate through all pages
- Find annotation objects (highlights, underlines, etc.)
- Extract the text from the highlighted regions
- Collect any notes or comments attached to the highlights
- Format and output the results
- Highlights
- Underlines
- Strikeouts
- Squiggly underlines
- This tool works with standard PDF annotations
- Apple Books stores highlights as PDF annotations, making them compatible
- The quality of extracted text depends on the PDF structure
- Some PDFs may have image-based text that cannot be extracted
MIT
Feel free to submit issues or pull requests for improvements!