Today we're publishing the first part of a running Q&A, written by the people who actually build Harold rather than by a marketing team.
The reason is straightforward. If you go looking for information about processing supplier documents, you find three separate conversations that no longer make sense as separate things.
Three industries that used to be different
OCR was one industry. It read characters off a page, and it needed rigid, consistent layouts to do it — you told it where on the document to look, and if the supplier changed their template, you changed yours.
EDI was another. It was your database talking to their database, and it was mostly trust. Consultants could put checks in, but it required a strong enough business relationship that both sides would build for each other. If you didn't have the buying power to demand it, you didn't get it.
Document automation was a third — a category name that mostly meant workflow software wrapped around the first two.
They have collapsed into one thing. Once you can add intelligence on top of character recognition, the document stops needing to be predictable, and the relationship stops needing to be formal. You don't need buying power to ask a supplier to send their invoice to a different email address. That single change does what EDI promised, without the project.
Why the existing material doesn't help
Most of what's published about this is written to support a sales process that hasn't changed since about 2013: big setup fee, consultancy, long contract, data locked in so leaving is painful.
So you get accuracy claims with no test conditions attached. You get "fully automated" with no discussion of what still needs a human. You get case studies with percentages and no arithmetic. None of it tells you the thing you actually want to know, which is: will this work on my invoices, and where will it fail?
What we're doing instead
We're answering questions in public, in plain terms, and we're including the limits.
Part 1 covers thirteen questions: what actually separates OCR from document automation from EDI, how many invoices a month justify automating, why "99% accurate" is a sales figure rather than an engineering one, whether you still need to check every invoice — the honest answer there is yes — what happens when the extraction gets something wrong, and why food and construction invoices are the hardest documents we handle.
Some of those answers are not flattering to our own product category. That's the point. A tool that quietly writes bad data into your ledger is worse than no tool at all, and pretending otherwise is how this industry lost people's trust in the first place.
Send us questions
Part 2 will cover pricing and what a document really costs to process, PO matching, credit notes and remittances, what happens to your data if a small software company disappears, and who shouldn't buy Harold.
If there's something you want answered, email hello@useharold.com and we'll put it in.