Multi-language extraction
Arabic (native) and English out of the box, with configurable extractors for other RTL and LTR languages. Print, typed, and handwriting supported.
- Arabic
- English
- Handwriting
Arabic and English extraction from IDs, invoices, contracts and handwriting: structured JSON in S3, feeding Bedrock knowledge bases.
Smart OCR is Phoenix's productized document-AI pipeline, live on AWS Marketplace. It reads Arabic and English documents, including handwriting, and delivers structured JSON in Amazon S3 that downstream systems, RAG knowledge bases, and dashboards can consume directly.
Deployed to customer AWS accounts, so the documents never leave the customer's data boundary. Sold as a product, tuned per customer on top of the same core pipeline.
Arabic (native) and English out of the box, with configurable extractors for other RTL and LTR languages. Print, typed, and handwriting supported.
IDs and passports, invoices and receipts, contracts, KYC packs, medical documents, government forms, tabular data: extractor templates per doc-type.
Every extracted document becomes a normalised JSON record in Amazon S3 with confidence scores per field: feed BI, integrate to SAP or CRM, index into Bedrock.
Extraction pipeline plugs directly into Amazon Bedrock knowledge bases: the same document corpus becomes a RAG source for enterprise copilots.
01
Doc types, languages, volume, target downstream systems: priced against a per-page tier.
02
Product installed into the customer's own AWS account: data stays inside their boundary.
03
Extractor templates configured per doc type; sample corpus benchmarked to accuracy targets.
04
Volume ramp, integrations wired to SAP / CRM / knowledge base, ongoing accuracy monitoring.
Every engagement lands specific artefacts, not slides.
Smart OCR is a productized pipeline tuned for Arabic, English and handwriting, with per-document-type extractor templates (IDs, invoices, contracts, KYC), confidence scoring, and JSON-in-S3 output ready for downstream systems. It's built on AWS services (including Textract where it fits) but delivered as a product, not a toolkit.
No. Smart OCR deploys into the customer's own AWS account. Documents are extracted, JSON output lands in the customer's own S3 buckets, and everything stays inside the customer's data boundary.
National IDs and passports, invoices and receipts, contracts, KYC packs, medical documents, government forms and tabular data. New document types are added by configuring an extractor template: usually days, not weeks.
Every extracted field carries a confidence score. Per-doc-type dashboards report accuracy and low-confidence rates. Human-in-the-loop review can be wired in for high-value docs, and low-confidence samples feed back into extractor tuning.
Yes, that's the intended shape. Extracted JSON is stored in S3 in an already-normalised form, so plugging it into Amazon Bedrock Knowledge Bases is a straight indexing step. Same corpus becomes the ground truth for enterprise copilots.
Whether you're scoping an SAP move to cloud, restarting a stalled programme, or just trying to figure out where data and AI fit, start with a conversation.
Talk to us about your project
Start a secure conversation
Enter your details and verify your email to chat with Phoenix Assistant.
Verify your email
We sent a 6-digit code to .
Or type your own message below