Skip to main content

What It Is

Text Recognition is the official OCR skill in the Skill Library. Users upload an ID, invoice, table, or product image in your app; the service extracts text and, where available, structured fields for form filling and downstream workflows. OCR is synchronous: one request processes one image and returns a text summary plus type-specific structured details. It does not use the asynchronous video task flow.

Supported Types

  • General and handwritten text
  • ID card and bank card
  • Vehicle license and driver license
  • Passport and business license
  • VAT invoice and mixed receipts
  • License plate and table recognition

Core Capabilities


How to Enable

1

Click Use

Open Text Recognition in the Skill Library, or tell superun that your app needs ID, invoice, or image text recognition.
2

Let superun prepare dependencies

superun enables OCR, project file storage, and the required cloud runtime, then installs the OCR credential in the project’s server environment.
3

Add the app workflow

Upload one image in the app, submit recognition from a server-side function, and display the synchronous result.
OCR is a capability for the app you build, not the assistant’s general image-understanding tool. The browser must never receive or call with the OCR credential directly.

Image and Result Requirements

  • Process one image per call. Common formats include JPG, PNG, BMP, GIF, TIFF, and WebP.
  • Upload local images to project storage first, then call OCR from the server with an accessible HTTP(S) URL. Local filesystem paths and browser blob: URLs do not work.
  • Provider URL parameters are limited to 2048 bytes. Use a shorter storage or CDN URL when a signed URL is too long.
  • Results contain recognized text and type-specific structured details. Do not assume every category returns the same fields.
  • Keep the OCR credential in server-side environment variables only; never place it in frontend code, database rows, logs, or API responses.

Billing

  • A successful recognition is billed once; a failed provider call is not billed as a success.
  • Usage appears under the Text Recognition consumption category.
  • Unit pricing follows the current product bill and is not hard-coded in this guide.

Privacy and Compliance

  • IDs and invoices may contain sensitive personal information.
  • The OCR processing path does not save the original image or result and does not log the image URL or recognized content.
  • If your app uploads the source image into project storage, retention of that uploaded file follows your app’s own storage and deletion policy.
  • Confirm the subject’s authorization before processing someone else’s ID or other sensitive document.

FAQ

No. Billing is reported after successful recognition; failed recognition is not charged as a successful call.
The OCR processing path does not persist the image or result. An image that your app uploaded to project storage remains subject to your app’s own storage and deletion rules.
Local paths and blob: URLs only exist in the current browser. Upload the image to project storage, then call OCR from a server-side function with an accessible URL.