Meet Fare AI. Pages in. JSON out.

Fare AI is PearFare's own family of extraction models, purpose-built and trained on millions of documents. Send scraped HTML or PDFs, or just the URLs and we collect the pages free. Schema-validated JSON is back in your portal or S3 bucket typically within 3 hours from intake, up to a quarter-million pages a day.

input: listing_4482.htmlraw
output: listing_4482.jsonrec 4,482 / 128,000
SCHEMA VALID
example batch: 96,400 pages in 5h 47m
20M+
pages of monthly capacity
100%
of delivered records validate
3hr
typical batch turnaround
72hr
typical, samples to a production extractor
Models

Frontier models, free at your fingertips.

Fare 4.1 Lite and Fare 4 are the heart of PearFare. Every extractor we run is a dedicated tune of one of them on your own documents, and the free pilot puts them to work on 10,000 of your pages before you spend a dollar. Your pages tune your extractor and nothing else.

Fare 4.1 Lite
light and fast, built for volume

Lite is the workhorse. It tunes on your samples in two to three business days, runs high-volume extraction at speed, and every free pilot report is scored on its output.

inputHTML, PDF, CSV, TXT
outputschema-locked JSON, validated twice
disciplineabsent fields come back null
reportper-field precision and recall
Fare 4
deep capacity for hard layouts

Fare 4 is the higher-capacity model in the pair. It drafts your schema from a plain field list, checks every upload for schema fit before extraction starts, and takes the hard layouts.

draftingfield list in, approved schema out
fit checkevery upload, before extraction
escalationhard layouts and edge cases
tierincluded with paid engagements
Process

How you get started.

Create your account and start your pilot in about a minute. You handle steps one and three, we do the rest. Tuning typically takes two to three business days; the first week covers samples, your report review, and your first production batches.

Day 0
01You

Submit samples

Upload around 200 example pages, plus a plain list of the fields you want back. We draft the formal JSON schema from that list, and you approve it before tuning starts.

pages.zippermits.pdfcrawl.csv
+2-3 days
02Us

Model tuning

A dedicated extractor is tuned and benchmarked on your pages. You see the accuracy report before you send a production batch.

precision98.4
recall97.9
Anytime
03You

Submit batches

Drag batches into the portal, or send a URL list and we collect the pages for you, free. Larger pipelines submit through the API or S3.

⇧ batch.zipqueued
Within 12h
04Us

Delivery

Typical turnaround is under 3 hours; the contract commits to 12, backed by automatic credits. Validated records are delivered to your portal, webhook, or S3 bucket.

records_0716.json, 128,000 rows, valid
Pricing

Flat pricing, quoted per workload.

Self-serve plans are paused for now. After your free pilot we quote a flat rate per thousand pages for your volume, and past five million pages a month pricing is set to your workload.

Pricing
Quoted
per 1,000 pages, fixed in writing for your term
  • Run the free pilot on 10,000 of your pages and review the scored report
  • We quote a flat per-thousand-page rate for your monthly volume
  • The quote states included volume, the 12-hour delivery commitment, and the term
  • PO and invoice billing available
Every paid engagement includesfree page collectionschema-locked outputper-field accuracy reportvalidity warranted in the contractautomatic SLA credits14-day purgeencrypted at restmutual confidentiality in the contractno training for anyone else
Coverage

Built around your schema.

These four run at volume today. Open a card to see a finished record.

view →

Real-estate listings

address str
price_usd int
beds, baths num
mls_id str
sample record{ "address": "1847 Birchwood Ln", "price_usd": 412000, "beds": 4, "baths": 2.5, "sqft": 2140, "hoa_monthly_usd": 85, "mls_id": "55-20417" }
view →

Leads & companies

name, title str
company str
phone e164?
industry enum
sample record{ "name": "Maya Chen", "title": "VP Operations", "company": "Halstead Logistics", "phone": "+14145550188", "industry": "logistics" }
view →

Product catalogs

sku str
title, brand str
price money
in_stock bool
sample record{ "sku": "NG-88213", "title": "TrailPro 65L Backpack", "brand": "Northgate", "price": 189.99, "in_stock": true }
view →

Records & filings

entity str
filing_type enum
date iso8601
filing_id str
amount_usd money?
sample record{ "entity": "Ridgeline Builders LLC", "filing_type": "building_permit", "filing_id": "2231-B", "date": "2026-06-22", "amount_usd": 84500 }
Guarantees

Three things we put in writing.

100% valid JSON

Records are generated so they can only take shapes your schema allows, then checked again before delivery. Valid means every record conforms to your schema; accuracy is scored separately, field by field, in your report. The contract warrants that every delivered record validates, and anything that does not comes back flagged and unbilled.

Automatic SLA credits

We measure the share of pages delivered inside the 12-hour window each billing month. If it slips below 99%, a 25% credit lands on your next invoice automatically; below 95%, 50%. No request needed, and failed pages are never billed.

No training for anyone else

Your pages tune your dedicated extractor and nothing else. They are processed under the contract's mutual confidentiality terms on PearFare-owned hardware and purged 14 days after delivery.

Questions

Common questions.

How is accuracy measured?

Every pilot, and every production batch after it, ships with a field-level precision and recall report scored on your own pages. Precision is the share of extracted values that are correct; recall is the share of values present on the page that we captured. Tuned extractors typically score 97 to 99% field accuracy on stable layouts.

Where do the accuracy numbers come from?

During tuning we hand-label a random sample of your pages and keep it out of tuning. Every report scores the extractor's output against that labeled set, field by field. You never need to supply labeled data.

What happens when a page fails?

It gets a repair pass. If it still fails validation, it comes back flagged and unbilled.

Do I need a JSON schema?

No. Send a plain list of the fields you want back with your samples. We draft the JSON schema, you approve it, and the approved schema is what every delivered record validates against; once you approve it, it is treated as the schema you supply under the terms. Every engagement includes one tuned schema; additional schemas are priced on request.

What inputs do you take, and what counts as a page?

Scraped HTML, PDFs with selectable text, CSV, and TXT, zipped into batches. One HTML document is one page; for PDFs, each PDF page counts as one page, so a 12-page permit file bills as 12 pages. For scanned or photographed documents, email pilot@trypearfare.com and we will confirm fit before your pilot.

Can you collect the pages for us?

Yes, at no extra charge; we call it Managed Collection. Send a list of sources or URLs and we fetch the pages before extraction, at no extra charge. We collect publicly accessible pages only, never behind logins or access controls, at controlled request rates, and we may decline a source. The 12-hour delivery clock starts when collection completes and the batch enters the queue. Details are in the terms of service.

What does the free pilot include?

Everything a paid batch gets: tuning of your extractor on your samples, up to 10,000 pages processed, validated JSON, and the scored accuracy report. No payment details are required. If the report misses your bar, we re-tune on additional samples before you commit. The pilot carries no delivery time commitment; paid engagements do.

How is pricing set right now?

Self-serve plans are paused while we finalize pricing. After your free pilot, we quote a flat per-thousand-page rate for your workload in writing, and that rate holds for the term of your agreement.

What happens to batches larger than 250,000 pages?

They stage across consecutive days in portions of up to 250,000 pages, and the 12-hour delivery window applies to each portion from its scheduled submission. The portal shows the staging schedule at intake. Sustained volume above the cap is quoted with reserved capacity.

Where does client data live?

On bare-metal GPU nodes PearFare owns outright, at a single private site in Wisconsin. No public cloud runs any processing. Batches you upload through the portal, and finished results waiting for download, are staged in encrypted Cloudflare storage at the edge, and staged copies are removed within 14 days. Processing copies on PearFare hardware are purged 14 days after delivery, or immediately on delivery confirmation with the zero-retention option, requested from your portal settings.

How is the service secured?

API traffic terminates at the Cloudflare edge with TLS 1.3, DDoS mitigation, and a managed web application firewall before reaching PearFare nodes. Drives are encrypted at rest, and administrative access is limited to SSH key pairs, with password login disabled at the operating-system level.

Run your free pilot: 10,000 pages, scored for accuracy.

Submit samples today. Tuning typically takes two to three business days, and your validated JSON with its accuracy report follows right after. Free, scored on your own pages, no payment details required.

Create your account

Account setup takes about a minute, no sales call. Email your samples and field list to pilot@trypearfare.com and they are attached to your workspace. When the report meets your bar, we quote your workload and you keep submitting.