Refractr: schema-driven extraction with field confidence and alpha logging
Refractr extracts structured fields from raw text and documents through an API. You supply a schema with each call and receive extracted values with field-level confidence scores. The service is labeled as an alpha release.
The homepage shows examples for invoices, bills of lading and support emails. A developer can ask for the fields an application needs, such as invoice number, vendor and total, instead of receiving an unrestricted prose answer.
Define the output for each request
Refractr says schemas do not need registration or a compilation step before use. Its examples specify typed placeholders for strings, numbers and dates, including nested line-item arrays. The same requested output shape can therefore be used across documents with different layouts.
The publisher lists PDF, DOCX and ODT input alongside text and several image formats. Scans and photos use its OCR option. Choose fields according to the downstream task: a delivery-note workflow may need the recipient and pallet count, while invoice processing also needs values that reconcile with totals.
For typed placeholders, the FAQ says a field returns the declared type or null when information is missing. Keep required-field checks in your own application. A response can match the requested shape while containing a value that needs correction.
Use confidence to route review
The homepage illustrates an invoice where a missing purchase-order number returns null and a lower-confidence total goes to a person for review. That makes confidence useful as a routing signal, rather than a substitute for checking the document.
Pick review thresholds using examples from your own workload. Compare the returned fields against the original files, including unfamiliar layouts and ambiguous labels. Check how your system handles nulls before automatically creating accounting records or sending customer responses.
The publisher reports an 82.0 field-level accuracy index and 0.49-second median document latency in its September 28, 2026 extraction benchmark. These are measurements reported by Refractr, not independent Franklin tests. Its FAQ also distinguishes a guarantee about response shape from model-based extraction accuracy.
Treat correct JSON as an interface requirement. It does not establish that the extracted date, total or identity is correct, and confidence scores need to be evaluated for the errors that matter in your application.
Check alpha data handling and credit use
During alpha, Refractr logs requests and responses to debug failures and improve extraction quality. Review that policy before submitting confidential documents or personal information. The publisher says logs are not sold or shared, but the logging itself still matters when deciding what you may upload.
The pricing section lists one credit per successful extraction and two credits for scans or photos using OCR. Failed extractions are free, and the alpha includes a daily free allowance. Confirm current terms before budgeting a production workload.
Start with representative non-sensitive documents and a small integration. Verify field correctness, missing-value handling and review routing before expanding the volume; the API's useful output is data your application can validate and use.