# Fetch a page

Last updated: 2026-10-10

`POST /v1/fetch` fetches one URL and tells you what came back: a `content_kind` verdict, HTTP status, final URL after redirects, title, hashes, injection flags and, if you ask, up to 20,000 characters of clean text. Price: US$0.002 per call.

Use it when you need to know whether a URL gave you real content, a bot challenge, a JavaScript shell, a paywall or an error, with a receipt that records it.

## Request

```json
{ "url": "https://www.iana.org/help/example-domains", "return_text": true }
```

| Field | Required | Meaning |
|---|---|---|
| `url` | yes | http or https URL, at most 2,048 characters |
| `return_text` | no | also return up to 20,000 characters of clean text; default `false` |

## Response

From the free sample (`GET /v1/sample/fetch`), shortened:

```json
{
  "page": {
    "requested_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
    "final_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
    "http_status": 200,
    "content_type": "text/html",
    "content_kind": "real",
    "signals": ["status_200"],
    "title": "Regional climate report 2025 (demo page)",
    "redirects": [],
    "content_sha256": "0e429607…",
    "raw_sha256": "179b9ac4…",
    "retrieved_at": "2026-10-09T14:19:34.909Z",
    "bytes": 847,
    "truncated": false
  },
  "text": "Regional climate report 2025\n\nThis is a demo page…",
  "text_chars": 636,
  "text_truncated": false,
  "injection_flags": [],
  "advice": "The page loaded and looks like real content.",
  "receipt": "eyJhbGciOiJFZERTQSIs…",
  "tier": "paid",
  "rail": "x402",
  "network": "eip155:84532"
}
```

- `content_kind` and `signals` are explained in [How it works](/docs/how-it-works). `advice` is one plain sentence about the verdict.
- `content_sha256` hashes the full extracted text, not the 20,000-character copy. `raw_sha256` hashes the bytes as received. `content_sha256` is null when no body was read.
- `injection_flags` lists prompt-injection patterns found in the text, such as `instruction_override` or `hidden_instructions`. They are flags, not a verdict: an empty list means none of our patterns matched.
- `text` is null unless you set `return_text`.

## PDF

An `application/pdf` page is read too, within the same 2 MB and 10-second limits. Only its text layer is read, with no OCR, so a scanned PDF comes back as `empty` with the signal `pdf_no_text_layer`.

- `signals` includes `pdf`. `title` is the PDF's own title, if it has one.
- `text` has the pages in order, separated by a form feed (`\f`).
- Only text inside each page's box is read, as a reader sees it. Text drawn outside the page is not in `text`.
- `page.pdf` gives `page_count`, `pages_read` and `truncated`. Very long PDFs stop at a page or character cap, with the signal `pdf_truncated`. The receipt signs these as well.
- A PDF we could not read is `unsupported_type` with a signal saying why: `pdf_encrypted`, `pdf_invalid`, `pdf_over_budget`, `pdf_timeout` or `pdf_out_of_memory`. It is charged like any other page we cannot read.
- When too many PDFs are waiting to be read, the call gets 503 `over_capacity` (`details.reason` `pdf_queue_full`) with `Retry-After`. It is not charged.

Only the full service reads PDFs: an edge server answers 501 `not_on_edge` (`details.reason` `pdf`) for one, not charged. Limits are in [Limits](/docs/limits).

## Charging

A bad URL, a blocked address or a host name that does not resolve is refused before payment and not charged. Results about the page itself, including 404s, bot walls and timeouts, are charged: that is the answer you paid for. See [Payments](/docs/payments).
