Skip to main content
PHOENIX CONSULTING & DEVELOPMENT LIMITED logo
Available on AWS Marketplace Phoenix Smart OCR

Smart OCR

Document AI for enterprise workflows, in Arabic and English.

Phoenix Smart OCR is a document-AI platform built on AWS that extracts structured data from the paper and PDF flow that enterprises still run on: IDs, invoices, forms and even handwriting, in both Arabic and English. Extracted fields land as structured JSON in Amazon S3, ready to feed Bedrock knowledge bases and downstream SAP workflows.

What it does

Feature set

Bilingual extraction

Arabic and English extraction from IDs, invoices, forms, receipts and handwriting, with confidence scores per field.

Structured JSON output

Every document parsed into a structured JSON payload landed in Amazon S3 for immediate downstream consumption.

Bedrock-ready

Extracted content feeds Amazon Bedrock knowledge bases so GenAI copilots can ground themselves in enterprise documents.

SAP-workflow friendly

Output shape mapped for direct hand-off into SAP business processes: invoice posting, HR onboarding, master-data updates.

Use cases

Where it shows up

  • Automated invoice capture and posting into SAP FI/CO
  • KYC and customer onboarding across financial services and telco
  • HR document capture: IDs, contracts and forms feeding SuccessFactors
  • Regulatory and compliance archives: searchable Arabic and English document lakes

Built on: AWS · Amazon Bedrock · Amazon S3 · SAP integration

FAQ

Frequently asked

How does Smart OCR compare to Amazon Textract or open-source OCR?

Smart OCR is a productized pipeline tuned for Arabic, English and handwriting, with per-document-type extractor templates (IDs, invoices, contracts, KYC), confidence scoring, and JSON-in-S3 output ready for downstream systems. Built on AWS services including Amazon Textract where it fits, but delivered as a product, not a toolkit.

Does the document data leave our AWS account?

No. Smart OCR deploys into the customer's own AWS account via AWS Marketplace. Documents are extracted, JSON output lands in the customer's own S3 buckets, and everything stays inside the customer's data boundary.

How is pricing structured on AWS Marketplace?

Metered against volume: pages, documents or contracted throughput tiers, depending on the plan chosen at subscription. Billing goes through the customer's AWS account with no separate procurement. Contact us for volume tier details and enterprise commitments.

How long does deployment take?

Typically 2 to 4 weeks from subscription to production: infrastructure provisioning inside your AWS account, extractor-template tuning per document type, integration with your downstream systems (SAP, CRM, Bedrock knowledge base as applicable), and a benchmarked accuracy validation on a sample corpus.

Which document types are supported out of the box?

National IDs and passports, invoices and receipts, contracts, KYC packs, medical documents, government forms and tabular data. New document types are added by configuring an extractor template, usually days, not weeks.

The earliest conversations are usually the most useful.

Whether you are scoping document AI for SAP workflows or grounding a Bedrock knowledge base in enterprise documents, start with a conversation.