What It Is
Text Recognition is the official OCR skill in the Skill Library. Users upload an ID, invoice, table, or product image in your app; the service extracts text and, where available, structured fields for form filling and downstream workflows. OCR is synchronous: one request processes one image and returns a text summary plus type-specific structured details. It does not use the asynchronous video task flow.Supported Types
- General and handwritten text
- ID card and bank card
- Vehicle license and driver license
- Passport and business license
- VAT invoice and mixed receipts
- License plate and table recognition
Core Capabilities
How to Enable
1
Click Use
Open Text Recognition in the Skill Library, or tell superun that your app needs ID, invoice, or image text recognition.
2
Let superun prepare dependencies
superun enables OCR, project file storage, and the required cloud runtime, then installs the OCR credential in the project’s server environment.
3
Add the app workflow
Upload one image in the app, submit recognition from a server-side function, and display the synchronous result.
OCR is a capability for the app you build, not the assistant’s general image-understanding tool. The browser must never receive or call with the OCR credential directly.
Image and Result Requirements
- Process one image per call. Common formats include JPG, PNG, BMP, GIF, TIFF, and WebP.
- Upload local images to project storage first, then call OCR from the server with an accessible HTTP(S) URL. Local filesystem paths and browser
blob:URLs do not work. - Provider URL parameters are limited to 2048 bytes. Use a shorter storage or CDN URL when a signed URL is too long.
- Results contain recognized text and type-specific structured details. Do not assume every category returns the same fields.
- Keep the OCR credential in server-side environment variables only; never place it in frontend code, database rows, logs, or API responses.
Billing
- A successful recognition is billed once; a failed provider call is not billed as a success.
- Usage appears under the Text Recognition consumption category.
- Unit pricing follows the current product bill and is not hard-coded in this guide.
Privacy and Compliance
- IDs and invoices may contain sensitive personal information.
- The OCR processing path does not save the original image or result and does not log the image URL or recognized content.
- If your app uploads the source image into project storage, retention of that uploaded file follows your app’s own storage and deletion policy.
- Confirm the subject’s authorization before processing someone else’s ID or other sensitive document.
FAQ
Am I charged when recognition fails?
Am I charged when recognition fails?
No. Billing is reported after successful recognition; failed recognition is not charged as a successful call.
Does the platform keep my uploaded ID image?
Does the platform keep my uploaded ID image?
The OCR processing path does not persist the image or result. An image that your app uploaded to project storage remains subject to your app’s own storage and deletion rules.
Why can’t OCR read a browser-local image URL?
Why can’t OCR read a browser-local image URL?
Local paths and
blob: URLs only exist in the current browser. Upload the image to project storage, then call OCR from a server-side function with an accessible URL.
