A Python-based utility for automatically scanning, renaming, and organizing document files (PDFs and images) based on their content using OCR. The system watches for new files, processes them to extract reference numbers, incident numbers and dates, and organizes them into appropriate folders.
- Automatic file watching and processing
- OCR-based document scanning for reference numbers and dates
- PDF and image file support
- Automatic file organization into dated folders
- PDF splitting capabilities
- Batch file processing
- File renaming utilities
Before running this project, you need to install the following dependencies:
pip install opencv-python
pip install pillow
pip install paddleocr
pip install pdf2image
pip install watchdog
pip install pypdfThese packages are also available in the requirements.txt file and can be installed using:
pip install -r requirements.txtYou'll also need to install system dependencies for pdf2image:
For Ubuntu/Debian:
sudo apt-get install poppler-utilsFor macOS:
brew install popplerfile_handling.py: Core utilities for file operations and organizationscan_number.py: OCR processing and document scanning functionalitywatch.py: File monitoring and automatic program activationDockerfile: Used to create a docker container that can run the project independently starting from watch.py
To start automatic file monitoring and processing:
python watch.pyThis will:
- Monitor the current directory for new files
- Automatically process new documents
- Organize files into dated folders
You can also process and handle files manually using the menu interface:
python file_handling.pyMenu options include:
- Folderize - Organize scanned and renamed files into dated folders
- Move all PDFs in the current working directory into specified location
- Prepend PDFs with 'scan' - necessary for scanning the file (mainly used for rescanning)
- Split PDF into multiple PDFs - for splitting different pages of PDFs into separate files
- Exit
The system expects and processes files with the following patterns:
- Files starting with "scan" (for listing)
- Files starting with "OD" (for organization)
- Supported formats: .pdf, .jpeg, .png (although major updates have been made for PDFs specifically)
Files are organized into folders based on their dates:
- Regular files go into MM-DD folders
- Files without reference numbers go to "Reference_not_found"
- Files without dates go to "Date_not_found"
- Files without incident numbers are ignored
listscans(): List all scan files in current directoryrename_scans(): Rename scan filesfolderize(): Organize files into dated folderssplit_pdf(): Split PDFs into smaller filesmove_pdfs_to_folder(): Batch move/copy PDF files
scan_image(): Process and extract info from imagesscan_pdf(): Process and extract info from PDFsprocess(): Image processing for better OCR resultsocr(): Extract reference numbers and dates using OCR
- Continuous monitoring of directory for new files
- Automatic processing of new documents
- Multi-process handling of file watching and manual scanning
- Files without reference numbers are moved to "Reference_not_found" directory
- Files without dates are moved to "Date_not_found" directory
- Invalid files and processing errors are logged to console
To contribute to this project:
- Fork the repository
- Create a feature branch
- Commit your changes
- Push to the branch
- Create a Pull Request
MIT License