# Verify citations

Last updated: 2026-10-10

`POST /v1/verify/citations` checks 1 to 10 citations, each a quote and the URL it is attributed to, in one call. For each one you get the same answer as [Verify a quote](/docs/verify-quote) (`exact`, `fuzzy` or `none`, score, offsets, context) plus the verdict on its page, all under one signed receipt. Price: US$0.10 per batch.

Use it on the sources at the end of an answer before the user sees it: a made-up quote, a misquote, or a link to a dead or blocked page is caught in one paid call.

For 1 to 9 pairs, one [verify/quote](/docs/verify-quote) call per pair costs less, without the Crossref data.

## Request

```json
{
  "citations": [
    { "url": "https://www.iana.org/help/example-domains", "quote": "These domains may be used as illustrative examples in documents" },
    { "url": "https://www.iana.org/domains/reserved", "quote": "IANA-managed Reserved Domains" }
  ],
  "match": "fuzzy"
}
```

| Field | Required | Meaning |
|---|---|---|
| `citations` | yes | 1 to 10 objects `{ "url", "quote" }`. `url` is http or https, at most 2,048 characters; `quote` is at most 1,000 characters after normalisation |
| `match` | no | for every citation: `fuzzy` (default) also accepts small differences; `exact` needs the same text after normalisation; `passages` treats each quote as a paraphrased claim (below) |
| `case_sensitive` | no | default `false`; must be `false` with `passages` |
| `threshold` | no | fuzzy similarity needed, 0.5 to 1, default 0.9 |

Several citations may share a URL: each distinct URL is fetched once and every quote on it is matched against the same copy. Every URL is checked before payment; if one is malformed, blocked or disallowed by robots.txt, or any quote is empty or too long (or, with `passages`, has no word to compare), the whole request is refused with 400 or 403 and not charged. The error's `details.field` names the citation, for example `citations[2].quote`.

## Response

From the free sample (`GET /v1/sample/citations`), shortened:

```json
{
  "results": [
    {
      "url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
      "page": 0,
      "final_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
      "http_status": 200,
      "content_kind": "real",
      "match": "fuzzy",
      "score": 0.963,
      "start": 211,
      "end": 238,
      "context": "...the agency said that emissions fell 12 % in 2025, after two years of slow growth...",
      "occurrences": 1,
      "advice": "The page loaded and looks like real content."
    },
    {
      "url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/missing",
      "page": 1,
      "final_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/missing",
      "http_status": 404,
      "content_kind": "http_error",
      "match": "none",
      "score": null,
      "start": null,
      "end": null,
      "context": null,
      "occurrences": 0
    }
  ],
  "pages": [
    {
      "requested_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
      "final_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
      "http_status": 200,
      "content_kind": "real",
      "title": "Regional climate report 2025 (demo page)",
      "content_sha256": "0e429607…",
      "normalised_sha256": "086a29af…",
      "injection_flags": []
    }
  ],
  "summary": { "exact": 0, "fuzzy": 1, "none": 2, "failed": 0, "retracted": 0, "doi_unchecked": 0 },
  "receipt": "eyJhbGciOiJFZERTQSIs…",
  "tier": "paid"
}
```

- `results` has one entry per citation, in request order. `page` is its index in `pages`, which has one entry per distinct URL with the page verdict, hashes and network evidence.
- `match`, `score`, `start`, `end`, `context` and `occurrences` mean what they mean in [Verify a quote](/docs/verify-quote). Offsets are into the normalised text of that citation's page; `normalised_sha256` on the page lets anyone holding the text check them.
- Read `content_kind` before relying on `match`. A `none` on a page that is not `real` (a 404, bot wall or JavaScript shell) means we could not read the page, not that the quote is absent from it.
- `summary` counts the verdicts; `failed` counts citations that carry an `error` instead of a verdict (below). `retracted` counts citations on a paper Crossref marks retracted, and `doi_unchecked` citations whose page's DOI lookup was skipped, so their retraction status is unknown (see Retracted papers; both absent when DOI lookups are off).
- The receipt signs, per page, the URL, status, `content_kind`, hashes and evidence, and per citation its page index, the SHA-256 of the quote and the verdict. It does not contain the quote text. Check it with [`POST /v1/receipt/verify`](/docs/receipt-verify).

## Passages: paraphrased claims

With `"match": "passages"`, each citation's `quote` is a claim. Each result is then the one described in [Verify a quote: passages](/docs/verify-quote): up to 3 passages from the page that share most of the claim's words, each with `start`, `end`, `score`, `text`, `missing` and `flags`, plus `scorer`.

- `match` is `passages` or `none`, and `summary` adds a `passages` count.
- The response has a top-level `note`: the passages are evidence for you to judge, not a verdict.
- The receipt signs each citation's passages (`start`, `end`, `score`) and the `scorer`.
- The work is linear in page size, so passages share no work limit and never fail with `match_too_costly`.

The price is the same.

## PDF

A cited PDF is read and matched as in [Verify a quote: PDF](/docs/verify-quote). Each match and passage on it has `pdf_page` and `pdf_page_end`, and the receipt signs them. Its entry in `pages` has `pdf` (`page_count`, `pages_read`, `truncated`). If one PDF in the batch cannot be queued for reading, the whole batch gets 503 `over_capacity` (`details.reason` `pdf_queue_full`) and is not charged.

## Retracted papers (DOI)

A cited page with a DOI gets `crossref` in its `pages` entry: the work's `title`, `authors` (up to 10, with `authors_total`), `year`, `container_title` (the journal) and `type` from Crossref's public metadata, and `retracted`: `true` when Crossref lists a retraction, withdrawal or removal notice for it (Crossref carries the Retraction Watch database). `updates` lists every notice that updates the work (`type` such as `retraction`, `correction` or `expression_of_concern`, notice `doi`, `source`, `date`), and `update_to` is filled when the cited DOI is itself such a notice. `summary.retracted` counts the citations on retracted works. A citation that matches its quote on a retracted paper still needs a warning before you use it.

The DOI is taken from a `doi.org` link, else the page's own `citation_doi` (or similar) meta tag, else a DOI in the URL path; `doi_source` says which (`url` or `page`). `status` is `found`, `not_found` (not a Crossref DOI; DataCite and other agencies are not checked) or `skipped` with a `reason` (`rate_limited`, `busy`, `timeout` or `unavailable`) when Crossref is slow or we are at our share of its public limits. A skipped lookup never fails the call or changes the charge; `summary.doi_unchecked` counts the citations whose page's DOI lookup was skipped, so `retracted`: 0 with `doi_unchecked` above 0 means not every DOI was checked. Lookups are made one at a time across all callers (at most 120 a minute for paid calls; free trial calls have their own budget of 6 a minute, so they never use the paid one) and answers are cached for a day. The receipt signs each page's `doi`, `status` and `retracted`.

## Partial failures

A quote that would cost too much to match fuzzily on its page gets `"error": { "code": "match_too_costly", ... }` and null match fields; the rest of the batch is still checked and the batch is charged. The fuzzy searches of one batch also share a work limit (twice what one `verify/quote` call may use), spent in page order: a citation reached after it is used up gets the same error with a message saying so, while an exact match is still found. If every citation in the batch fails this way, the call gets 422 `match_too_costly`, not charged, and `details.failed` lists only which citations failed. Use `"match": "exact"` or shorter quotes.

Each caller IP address gets 6 uncharged failed batches a minute (a bucket that refills one every 10 seconds). A batch holds a place in that bucket while it runs and gives it back unless every citation failed, so a caller with 6 batches still running also waits. Over the limit, `verify/citations` answers 429 `rate_limited` with `details.reason` `uncharged_failures` and a `Retry-After` header, before any page is fetched, and nothing is charged. This normally comes before the 402; if your own parallel batches take the last place meanwhile, it comes after your payment is verified and before it is settled. Charged batches never use the bucket up.

A few batches are matched at once across all callers and a short queue waits behind them. When that queue is full, `verify/citations` answers 503 `over_capacity` with `details.reason` `match_queue_full` and a `Retry-After` header (a few seconds), and nothing is charged: before the 402 if the queue is already full, or after your payment is verified and before it is settled if it filled while your pages were fetched. Retry the same request after `Retry-After`.

## Dead domains

A URL whose host name does not resolve is reported as a result: `content_kind` `http_error`, `http_status` null, `match` `none`, and the batch is charged. If no URL in the batch resolves, the call gets 422 `host_not_found` and is not charged.

## Charging

If a page turns out to be disallowed by robots.txt or rate limited after payment (for example after a redirect), the whole batch answers with an error and is not charged. Results about the pages themselves, including 404s, bot walls and timeouts, are charged. See [Payments](/docs/payments).

A free-trial call covers at most 3 citations. A larger batch is not run as a trial: it gets the normal 402 offer, with `details.reason` `trial_too_large`, and is not charged.
