This Python script is designed to extract transaction data from PDF files and save the data to an Excel file. It utilizes the pdfplumber library to read PDFs and pandas to manage and export the data.
- Transaction Line Detection: Uses regular expressions to accurately identify transaction lines within the PDF text.
- Data Extraction: Breaks down each transaction line into components like date, transaction type, description, amount, and balance.
- Export to Excel: Saves the extracted data into an Excel file using the
openpyxlengine for easy analysis and record-keeping.
To run this script, you need to have Python installed along with the following libraries:
pdfplumberpandasre(regular expression library, which is included by default in Python distributions)openpyxl
Install the required Python libraries using pip:
- Place the PDF file you want to process in the same directory as the script or specify the path to the file.
- Run the script by executing: