# Verify a document

Last updated: 2026-10-10

`POST /v1/verify/document` takes an answer or report as written, finds every quote in it with its link by fixed rules, and checks each pair on the live page, as [Verify citations](/docs/verify-citations) does: `exact`, `fuzzy` or `none`, score, offsets in your document and on the page, and the verdict on each page, under one signed receipt. Price: US$0.10 per part of up to 10 pairs.

Use it before sending a finished answer or report that quotes web pages. Call the free preview first: it lists the pairs, what was skipped and why, the number of parts and the exact price.

For a part of 1 to 9 pairs, one [verify/quote](/docs/verify-quote) call per pair costs less; the free preview lists each pair and its part.

## Request

```json
{
  "document": "WHO says \"Air pollution kills an estimated seven million people every year\" ([WHO](https://www.who.int/health-topics/air-pollution)).",
  "format": "markdown",
  "match": "fuzzy",
  "part": 0
}
```

| Field | Required | Meaning |
|---|---|---|
| `document` | yes | The answer or report as text, 1 character to 60,000 bytes UTF-8 |
| `format` | no | `markdown` (default) or `text` (only bare URLs and `[n]` references are links) |
| `match` | no | `fuzzy` (default), `exact` or `passages`, as verify/citations; it also chooses what is found (below) |
| `case_sensitive` | no | default `false`; must be `false` with `passages` |
| `threshold` | no | 0.5 to 1, default 0.9 (fuzzy only) |
| `min_words` | no | 1 to 20, default 4: a quoted span with fewer words is skipped as `too_short` |
| `part_size` | no | 1 to 10, default 10: pairs per part. A free trial call needs 3 or fewer |
| `part` | no | 0-based part to check, default 0. Ignored by the preview |

Unknown fields are refused with 400. Every URL in the part is checked before payment, as on verify/citations: a blocked or robots-disallowed URL refuses the call with 400 or 403, `details.field` naming the pair (for example `pairs[3].url`), and nothing is charged.

## Preview

`POST /v1/verify/document/preview` is free. Send the same body; nothing is fetched or charged.

```json
{
  "document_sha256": "9f2c…",
  "extractor": "doc-v1",
  "pairs_total": 1,
  "part_size": 10,
  "parts": 1,
  "price": { "usd_per_part": "0.10", "sats_per_part": 100, "usd_total": "0.10", "sats_total": 100 },
  "pairs": [
    { "pair": 0, "part": 0, "quote": "Air pollution kills an estimated seven million people every year", "url": "https://www.who.int/health-topics/air-pollution", "doc_start": 10, "doc_end": 74, "source": "inline_link" }
  ],
  "skipped": [],
  "skipped_truncated": 0,
  "skipped_counts": {},
  "limit_reached": false
}
```

- `doc_start`/`doc_end`: code-point offsets of the quoted words in `document` as sent. A pair found more than once (same quote, same URL) is one pair with `doc_spans`, checked and charged once. `doc_spans` holds the first 20 places; `doc_spans_truncated` counts the rest.
- `source`: `inline_link`, `reference`, `footnote`, `numbered`, `bare_url`, `blockquote` or `sentence` (passages).
- `skipped[]`: `{ doc_start, doc_end, reason, text }`; reasons `too_short`, `too_long`, `no_url`, `bad_url`, `beyond_limit`. It lists the first 100; `skipped_truncated` counts the rest and `skipped_counts` gives every skip by reason.
- The preview contacts no URL, so a dead or blocked URL shows up only on the paid call.

The preview takes from the same per-address limit as the paid routes and works on the edge server too.

## Response

The verify/citations response for this part, plus the document fields:

```json
{
  "document_sha256": "9f2c…",
  "extractor": "doc-v1",
  "part": 0, "parts": 1, "pairs_total": 1,
  "results": [
    { "pair": 0, "quote": "Air pollution kills…", "url": "https://www.who.int/…", "doc_start": 10, "doc_end": 74, "page": 0, "final_url": "…", "http_status": 200, "content_kind": "real", "match": "fuzzy", "score": 0.96, "start": 412, "end": 478, "context": "…", "occurrences": 1 }
  ],
  "pages": [ "as verify/citations" ],
  "summary": { "exact": 0, "fuzzy": 1, "none": 0, "failed": 0 },
  "receipt": "eyJ…",
  "tier": "paid"
}
```

- `doc_start`/`doc_end` are offsets in your document; `start`/`end` are offsets on the page, as in [Verify a quote](/docs/verify-quote). Read `content_kind` before relying on `match`.
- The receipt (kind `document`) signs `document_sha256`, the extractor, the options, the part and, per pair, its page, the SHA-256 of the quote, its document offsets (with `doc_spans` and `doc_spans_truncated` for a repeated pair) and its verdict. It does not contain the quote text. Whoever holds the document can tie every part's receipt to it. Check it with [`POST /v1/receipt/verify`](/docs/receipt-verify).

## How pairs are found

The rules are fixed and named by `extractor` (`doc-v1`) in every answer and receipt; a change gets a new name.

- Code blocks and inline code are ignored. With `format` `markdown`, links are inline links, reference links, footnotes, `[n]` references with a numbered list and bare URLs; with `text`, only bare URLs and `[n]` references.
- A quote is text between paired `"…"`, `“…”`, `«…»` or `„…“` inside one paragraph, or a blockquote (its last line starting with `—`, `--`, `-` or `Source:` is the attribution and its link the source). Single quotes are not used.
- A quote's link is the first link after it before the next quote or the paragraph end; else the last link before it in the same paragraph; else the blockquote's attribution; else it is skipped as `no_url`.
- With `match` `passages`, every sentence that carries a link is a claim and gets its best-matching passages as evidence, as in verify/citations.
- Then `bad_url`, `too_short` (`min_words`, not for passages), `too_long` (over 1,000 characters), duplicates and the 100-pair limit (`beyond_limit`) are applied, and pairs are numbered in document order. Part `p` holds pairs `p × part_size` to `p × part_size + part_size − 1`.

### Fixes to doc-v1

- 2026-10-10: in an indented code block (4 spaces or a tab, after a blank line or heading), every line is now code. Before, only its first line was skipped, so a quote and link on a later line of the block could be taken as a pair. Nothing else changed.

## Parts and price

Each paid call checks one part of up to 10 pairs, at US$0.10: the same price as a verify/citations batch.

| Document | Calls |
|---|---|
| 4 quotes | 1 |
| 10 quotes | 1 |
| 23 quotes | 3 |

## Limits

| Limit | Value |
|---|---|
| `document` | 60,000 bytes UTF-8 (400 over it) |
| Pairs per document | 100 (10 parts of 10); the rest are skipped as `beyond_limit` |
| Pairs per part | `part_size`, at most 10 |
| Places listed per repeated pair | 20 in `doc_spans`; the rest counted in `doc_spans_truncated` |
| Skipped items listed | 100 in `skipped`; the rest counted in `skipped_truncated` |
| Quote or claim | 1,000 characters |
| Free-trial part | 3 pairs; a larger part gets the normal 402 offer (`details.reason` `trial_too_large`) |

Matching shares its slots, queue and failure limit with verify/citations: 6 uncharged failed calls a minute per caller address across both routes, then 429 `rate_limited` (`details.reason` `uncharged_failures`); a full match queue gives 503 `over_capacity` with `Retry-After`. See [Limits](/docs/limits).

## Charging

A part is charged when it is checked, including pages that are 404s, bot walls or timeouts, and a pair whose fuzzy work ran out (it carries an `error`). Never charged: any 400, 422 `no_pairs` (no quote with a link was found), 422 `host_not_found` (no URL in the part resolves), 422 `match_too_costly` (every pair in the part failed), 429, 503, and the edge server's 501 `not_on_edge`. See [Payments](/docs/payments).

## Errors

See [Errors](/docs/errors). `part` past the last part gets 400 with `details.parts`.
