Recognizing data is good. Providing the right data in its entirety is better.
PEDIF Data Refinement
PEDIF can do more than recognize and structure data from PDF documents. The extracted information can also be enriched and matched against existing master or reference data.
This turns the information contained in a document into exactly the data your ERP, EDI, or downstream system needs for further processing.
PDF in. Data recognized. Data refined. Structured processing out.








































What does Refinement mean in PEDIF?
Business documents do not always contain all the information a receiving company needs for processing.
And even if the required information is present, it may not match the recipient's own master data.
This is where PEDIF Refinement comes in.
PEDIF can:
1. enrich data
Missing information is added based on defined rules or available data sources.
2. match data
Incoming information is compared with the receiving company's master and reference data and translated into the values required by its systems.
1. Automatically enrich data
When the PDF does not contain everything your system needs
A customer does not necessarily include every piece of information on an order that your ERP requires to create a sales order.
PEDIF can add this information during processing.
Typical examples include:
- internal customer number
- GLN
- delivery address
- debtor number
- location or plant
- sales organization
- additional customer- or process-specific information
Fixed enrichment
Certain information can be permanently linked to a customer, document type, or process.
Example:
An order from Müller GmbH is recognized.
The order does not contain the internal customer number.
PEDIF knows from the configured rules:
Müller GmbH → Customer number 4711
Customer number 4711 is automatically added to the structured dataset.
Dynamic enrichment via lookup
Information can also be determined dynamically from a reference or master data source.
PEDIF uses an existing value as a search criterion and identifies the corresponding target value.
Example:
The order contains a GLN.
PEDIF uses the GLN for a lookup and automatically determines:
GLN → Customer → Customer number → Correct delivery address
This means the required data does not have to be fully present on the PDF.
2. Match data with master data
Because your customer does not always use your master data
This is a common challenge, especially with incoming orders:
A customer orders products using their own master data.
The supplier, however, needs to process the order in its ERP using its own master data.
Both companies are referring to the same product, but may use different:
- item numbers
- product descriptions
- units
- packaging units
- product references
- customer material numbers
A practical example
The customer orders using their own item number
The PDF order may contain:
Customer item number: 12345-A Description: Filter size 200 Quantity: 20 units
The supplier's ERP, however, knows this product under a different item number and its own master data.
PEDIF first recognizes the information from the order.
It then matches it against the stored product or reference data.
For example:
Customer item 12345-A
becomes:
Supplier item 987654
PEDIF can therefore map the customer's information to the correct master data used by the supplier.
The order can then be transferred using the product data required by the supplier's ERP system.

From PDF to the correct ERP dataset
The process can be simplified into four steps.
1. Recognize
PEDIF identifies the relevant information in the incoming document. For example: customer, order number, item number, quantity, delivery address, or delivery date.
2. Enrich
Missing information is added using fixed rules or dynamic lookups. For example: GLN → Customer number, or Customer → Default delivery address.
3. Match
Incoming values are matched against your own master and reference data. For example: Customer item number → Internal item number.
4. Transfer
The refined data is transferred in structured form to the defined downstream process, such as ERP, EDI, XML, JSON, CSV, or API.
Why is pure data recognition not always enough?
Reading a PDF correctly initially answers only one question:
“What is written on the document?”
For automated downstream processing, a second question is often more important:
“What information does my own system need?”
This is exactly the gap that Refinement closes.
Without Refinement
PDF says: Item 12345-A ERP expects: Item 987654 An employee has to know the mapping and correct it manually.
With Refinement
PDF says: Item 12345-A PEDIF recognizes: Item 12345-A PEDIF matches: 12345-A → 987654 ERP receives: Item 987654
Your rules define the result
Refinement does not mean guessing the meaning of data.
Enrichment and matching are based on defined rules and data sources.
You can specify:
Typical Refinement use cases
Customer Lookup
A recognized customer, GLN, or other reference is mapped to the internal customer number. Result: The ERP receives the correct debtor.
Address Lookup
A customer, location, or GLN reference is mapped to a stored delivery address. Result: The order is assigned to the correct delivery location.
Product Lookup
The customer's item number is mapped to the supplier's item number. Result: The order is processed using the supplier's own product master data.
Data enrichment
A value required for the process is not present on the PDF but can be determined unambiguously from the context. Result: PEDIF adds the information automatically.
Especially valuable for incoming customer orders
Customer orders often bring together two different master data environments.
The buyer works with their own catalog.
The supplier works with its ERP and its own product master data.
An order can therefore be completely correct from the customer's perspective and still contain information that cannot be transferred directly into the supplier's ERP.
PEDIF acts as a translation layer between these two worlds:
This allows the sender to continue using their familiar ordering process, while the recipient receives data in the form required by its own systems.
From PDF data to process-ready data
PEDIF should not only be able to answer what is written on a document.
The goal is to prepare the information so that it can actually be used in the next process step.
Recognize
What is written on the document?
Enrich
Which required information is missing?
Match
What is the corresponding value in my master data?
Transfer
What data does my target system need?
Frequently Asked Questions
Next step
Do you have your own mapping or master data rules?
Show us a typical document process together with the relevant master or reference data. Together, we can evaluate which information PEDIF can recognize, enrich, and match, and how this can be turned into a structured transfer process for your target system.
Technical process assessment within a defined scope; no blanket guarantee for every PDF and no legal advice.
Contact Us
Questions, need help choosing the right setup?
Book a Meeting
Prefer to talk directly? Pick a time that works for you.