OCR.space provides a free and paid OCR API for extracting text from images and PDFs, returning results in JSON format.
OCR.space provides a free and paid OCR API for extracting text from images and PDFs, returning results in JSON format. On Nagent, OCR.space is exposed as a fully-configurable ai document extraction integration that any agent can call — 3 actions, and API key authentication. No code is required to wire OCR.space into your workflow — connect it once via the External Integrations panel and reuse it across every agent you build.
Agent builders use OCR.space to automate the kinds of tasks ai document extraction teams previously handled manually. Concrete examples — each one is a single agent step in Nagent — include:
Every action and trigger is paired with a structured input/output schema (visible in the sections below), so when you wire OCR.space into Helix — our agentic agent builder — the editor knows exactly what each step expects and produces. Configure once, deploy anywhere across your Nagent agents.
Every operation an agent can call against OCR.space, with input parameters and output schema. Drop these into any step of an agent built in Helix.
OCRSPACE_GET_CONVERSIONSRetrieve OCR API conversion statistics and usage data (PRO accounts only). Returns the number of conversions for Engine1, Engine2, and total conversions. Data is updated once daily and shows conversions from start of month to end of yesterday. Free API keys will return 0 conversions.
Input parameters
Start date option for conversion statistics. Case-sensitive.
Output
Data from the action execution
Error if any occurred during the execution of the action
Whether or not the action execution was successful or not
OCRSPACE_OCR_PARSE_IMAGE_POSTExtract text from images and PDF documents using OCR (Optical Character Recognition). Supports 27 languages, table recognition, orientation detection, and word-level coordinate extraction. Provide exactly one of `file`, `url`, or `base64Image`; providing multiple or none triggers E301/OCRExitCode 99. Input can be provided as file upload, public URL, or base64-encoded data URI. Response is nested JSON; extract text from `ParsedResults\[*\].ParsedText`. Returns extracted text with optional overlay coordinates and searchable PDF generation. For poor-quality scans, enable both `detectOrientation` and `scale` and ensure `language` matches the document.
Input parameters
Publicly accessible URL of the image or PDF file. IMPORTANT: Must also set 'filetype' parameter (e.g., 'PNG', 'JPG', 'PDF') when using this option.
Binary content of image or PDF file for upload (multipart/form-data). Supports JPG, PNG, GIF, PDF, BMP, TIF formats.
If True, automatically upscales low-resolution images before OCR processing to improve text recognition accuracy.
If True, optimizes OCR for table/structured data recognition. Recommended for receipts, invoices, and tabular documents.
File format specification. REQUIRED when using 'url' or 'base64Image' parameters. Valid values: 'PDF', 'GIF', 'PNG', 'JPG', 'TIF', 'BMP'. Not needed for 'file' parameter.
OCR language code: ara=Arabic, bul=Bulgarian, chs=Chinese Simplified, cht=Chinese Traditional, hrv=Croatian, cze=Czech, dan=Danish, dut=Dutch, eng=English, fin=Finnish, fre=French, ger=German, gre=Greek, hun=Hungarian, kor=Korean, ita=Italian, jpn=Japanese, pol=Polish, por=Portuguese, rus=Russian, slv=Slovenian, spa=Spanish, swe=Swedish, tha=Thai, tur=Turkish, ukr=Ukrainian, vnm=Vietnamese
OCR processing engine selection. 1=Standard engine (default, reliable), 2=Experimental engine (may have better accuracy for some documents)
Base64-encoded image as a data URI string. Format: 'data:image/\[format\];base64,\[encoded-data\]'. IMPORTANT: Must also set 'filetype' parameter when using this option.
If True, automatically detects and corrects text orientation (rotation). Returns detected orientation angle in TextOrientation field.
If True, returns word-level bounding box coordinates (Left, Top, Height, Width) for each detected word. Useful for document layout analysis.
If True, generates a searchable PDF with an invisible text layer overlay. Returns PDF URL in SearchablePDFURL field.
If True (and isCreateSearchablePdf=True), hides the text layer in the generated searchable PDF. The text is still searchable but not visible.
Output
Data from the action execution
Error if any occurred during the execution of the action
Whether or not the action execution was successful or not
OCRSPACE_PARSE_IMAGE_URLExtract text from images via URL using simplified GET endpoint. Only supports URL-based submissions - no file uploads or base64 encoding. Faster and simpler than POST endpoint for basic use cases.
Input parameters
Publicly accessible URL of the image or PDF file to OCR. Must be a valid HTTP/HTTPS URL.
OCR language code: ara=Arabic, bul=Bulgarian, chs=Chinese Simplified, cht=Chinese Traditional, hrv=Croatian, cze=Czech, dan=Danish, dut=Dutch, eng=English, fin=Finnish, fre=French, ger=German, gre=Greek, hun=Hungarian, kor=Korean, ita=Italian, jpn=Japanese, pol=Polish, por=Portuguese, rus=Russian, slv=Slovenian, spa=Spanish, swe=Swedish, tha=Thai, tur=Turkish, ukr=Ukrainian, vnm=Vietnamese
If True, returns word-level bounding box coordinates (Left, Top, Height, Width) for each detected word. Useful for document layout analysis.
Output
Data from the action execution
Error if any occurred during the execution of the action
Whether or not the action execution was successful or not
No publicly available marketplace agent is found using this tool yet. There are 66 agents privately built on Nagent that already use OCR.space.
Build on Nagent
Connect OCR.space to any Nagent agent in minutes — no API key management, no boilerplate. Just configure and deploy.
The five questions agent builders ask before adopting a new integration.
Open the External Integrations panel inside Nagent (app.nagent.ai/externalIntegration), find OCR.space, and click "Connect Now." You'll authenticate with an API key — Nagent handles credential storage and refresh automatically. Once connected, OCR.space is available to any agent in your workspace.
No. Nagent provides no-code integration for every tool. Once OCR.space is connected, you configure its 3 actions directly in the agent builder UI — no API calls, no boilerplate, no schema management.
Helix — Nagent's agentic agent builder — lets you drop OCR.space steps into any workflow visually. Pick an action (e.g., one of those listed above), fill in the inputs (Helix knows the required vs. optional schema for each parameter), and connect it to upstream/downstream steps. Triggers run as the entry point of an agent, so when a OCR.space event fires, the agent kicks off automatically.
Every OCR.space action and trigger ships with a fully-typed schema — input parameters with name, type, required flag, and description, plus the output payload shape. The schemas are documented in the sections above. Helix uses these schemas to validate your configuration at build time and to type-check the data flowing between steps.
Yes. While OCR.space ships with 3 pre-built ai document extraction actions, you can layer custom logic around them inside Helix — pre/post-processing steps, conditional branches, retries, or stitching OCR.space together with other connected tools. For deeper customization, talk to our team about Nagent's Agentic AI Lab — forward-deployed engineers who build OCR.space-based workflows tailored to your business.