Usually, yes, and the data is the business's own even when the way in is the vendor's. There are four routes, and they differ in how reliable they are. The most complete is a read-only copy of the ERP's database, taken from a replica or a restored backup where the licence allows. Next come the report exports the system can already produce on a schedule, such as CSV or print files. Third is reading the printed or PDF reports it generates. The least reliable is driving the screens with bots that click and type like a user. Whichever route is used, check each extract against a total the old system prints itself.
What the usual answer says
The pages that rank for this question focus on APIs, database connectors and extraction tools. One analytics vendor's guide warns that "Even standard SQL queries or traditional ETL (extract, transform, load) processes don't work because of data duplication, security considerations and the risk of breaking core business operations." That is a fair warning about querying a live system. None of the pages compares the routes by how often they break, or mentions the reports the ERP already prints.
What the evidence says about screens
Screen automation has the best-documented failures, though not from ERPs. Michael Wornow and colleagues, in a 2024 study of enterprise workflow automation, examined robotic process automation in two case studies, a hospital and a business-to-business invoice process. They found "high set-up costs (12-18 months), unreliable execution (60% initial accuracy), and burdensome maintenance (requiring multiple FTEs)". In the full text, the invoice bot "took 6 months of improvement to reach 95%", and 2 FTEs were assigned to monitor it. At the hospital, "Quarterly updates to the hospital's electronic health record and constant changes to payers' websites would break the bot", and the hospital eventually had "to develop a custom API that the bot could use". Summarising earlier work, the authors note that bots "cannot adapt to slight variations in input (e.g., a button changing location on a screen, or a form field being renamed)".
The same study tried AI agents that read the screen visually. From a plain-language description of each workflow, their system completed 40% of the authors' test workflows end to end.
Reports: stable, but summarised
In our experience, an old ERP's standard financial reports are the part of it least likely to change, and they carry the totals the business trusts. That makes them a good fallback when the database is off-limits. They have limits of their own. Reports summarise and round, can carry filters nobody remembers setting, and may not have the transaction-line grain an analysis needs. A patch that shifts a column, or a subtotal line read as a record, breaks a parser as surely as a moved button breaks a bot.
How hard they are to read is also unproven for this case. A 2025 benchmark of table extraction from scientific PDFs by Marijan Soric and colleagues, covering 86,000 pages, found that "current methods suffer from a lack of generalizability when facing heterogeneous data". A separate 2026 benchmark, ParseBench, tested 14 methods on about 2,000 enterprise pages and found "no method is consistently strong across all five dimensions", with the best overall score at 84.9%. It was published by the company whose product scored highest. Neither tests a single report layout repeated thousands of times. Our reading is that such a report is a narrower problem than the varied documents these benchmarks measure, and a page total makes each extract checkable.
The order to try
First, ask for a read-only copy of the database. Query a replica or a restored backup rather than the live system, check the licence and the vendor's terms, and never write to it. Second, look for scheduled exports at the line level. Third, read the printed reports, and tie every page to its total. Fourth, and last, automate the screens, for data that exists nowhere else.
The part no route solves
Old systems are full of codes whose meaning lives in someone's head: branch numbers, product classes, customer types, flags. Lumber systems add their own, such as units stored per line rather than per item, tallies held in free-text fields, and price codes that changed meaning years ago. The first extract raises more questions about what fields mean than about how to reach them. Write the answers down as they come, because that list is what makes the extract usable.
When it doesn't apply
Some vendor contracts forbid direct database access, and some charge for exports; then reports are the practical route. Systems hosted by the vendor may expose nothing but screens and reports. And where a migration is already under way, extracting history once may be all that is needed.
How Quarri works explains the platform as a layer over existing systems, not a migration.
Sources
- Phocas, "How to extract data from ERP systems": phocassoftware.com
- Wornow, Narayan, Opsahl-Ong, McIntyre, Shah and RĂ©, "Automating the Enterprise with Foundation Models", arXiv 2405.03710, 3 May 2024: arxiv.org
- Soric, Gracianne, Manolescu and Senellart, "Benchmarking Table Extraction from Heterogeneous Scientific PDF Documents", arXiv 2511.16134, 20 November 2025: arxiv.org
- Zhang et al., "ParseBench: A Document Parsing Benchmark for AI Agents", arXiv 2604.08538, April 2026: arxiv.org
Quarri is an AI-native data platform for the timber supply chain. It connects buying, production, sales and inventory for forest management, sawmill, wood products and pulp, paper and packaging operators.