OCR Data Capture With Background Security Features.
Configure split-stream OCR capture so background security features stay visible
This article is for support, QA, implementation, and field engineering teams configuring OCR data capture for documents that also use background security features.
Applies To- Builds that include the split OCR capture stream change.
- Validated with the IPP Class Driver.
- Validated with
pdfocr8OCR capture and no white background cleanup options enabled.
When OCR was enabled with pdfocr8, data capture could work, but the OCR output could include a full-page white layer. If a template placed security features behind the document content, that white layer could hide the background features.
The split-stream change fixes this by using:
pdfwritefor the final printable document stream.pdfocr8for a separate OCR capture stream.
The user-facing configuration can still specify pdfocr8. In the split-stream build, the application internally keeps the printable stream on pdfwrite when a raster OCR engine is used, then sends the OCR stream to data capture.
Use the following baseline settings for the validated IPP Class Driver workflow:
{
"Engine": "pdfocr8",
"TryOCR": true,
"OCRMode": "Always",
"OCRLanguage": "eng",
"Segment1": " -sColorConversionStrategy=LeaveColorUnchanged -dALLOWPSTRANSPARENCY ",
"Segment2": "",
"Segment3": "",
"CaptureEngine": "",
"CaptureSegment1": "",
"CaptureSegment2": "",
"CaptureSegment3": ""
}
If the OCR language data is not found automatically, also configure the OCR data path:
{
"OCRDataPath": "C:\\Program Files\\Troy Group\\Troy Sentry Print Client\\Print Processor Service\\Resources\\tessdata"
}
Template Settings
For this workflow, leave white background cleanup options off.
Do not enable:
- Remove white vector fill backgrounds, method 1.
- Remove white vector fill backgrounds, method 2.
- Remove OCR white page background.
Known note: Remove white vector fill backgrounds (method 1) has been observed to throw an error and prevent printing in this scenario. Leave it off unless engineering specifically asks for a controlled test.
Expected BehaviorWith the recommended settings:
- Data capture should populate successfully.
- Portrait documents should print in portrait orientation.
- Background security features should remain behind the page content and remain visible.
- Security features should not need to be moved to the foreground just to appear.
- White background cleanup should not be required.
- Keep
Engineset topdfocr8when OCR data capture is required. - Keep
TryOCRset totrue. - Use
OCRModeset toAlwaysfor documents that require OCR capture. - Use
OCRLanguageset toengunless another language is required and installed. - Keep
CaptureEngineblank for the validated configuration. Blank allows the application to automatically use the OCR engine for the capture stream while keeping the print stream onpdfwrite. - Keep the
Segment1value shown above so Ghostscript preserves color behavior and PostScript transparency handling.
Before calling the configuration complete, verify:
- The workstation or server is running the split-stream build.
- The printer is using the IPP Class Driver.
- The OCR settings match the recommended baseline.
- White background cleanup options are disabled in the template.
- A test portrait document prints successfully.
- Data capture fields populate correctly.
- Background security features are visible without moving them to the foreground.
If data capture does not populate
- Confirm
Engineispdfocr8. - Confirm
TryOCRistrue. - Confirm
OCRModeisAlways. - Confirm
OCRLanguageiseng. - Confirm the
tessdatafiles are installed andOCRDataPathis set if needed. - Confirm the environment is using the split-stream build, not an older build.
If background features are hidden
- Confirm white background cleanup options are off.
- Confirm the template feature is actually configured as a background feature.
- Confirm the installed build contains the split-stream change.
- Confirm the job is being sent through the IPP Class Driver path used during validation.
If printing stops after enabling cleanup
- Turn cleanup back off.
- Retest with the recommended OCR settings.
- Capture the print processor logs and the exact template setting that caused the failure.
If OCR output looks like junk data
- Confirm the source job is being sent through the IPP Class Driver.
- Confirm
OCRModeisAlways. - Confirm
pdfocr8is available in the installed Ghostscript/runtime path. - Confirm the document is not being captured from the printable stream only.