Fetch a page
Last updated: 2026-10-10
POST /v1/fetch fetches one URL and tells you what came back: a content_kind verdict, HTTP status, final URL after redirects, title, hashes, injection flags and, if you ask, up to 20,000 characters of clean text. Price: US$0.002 per call.
Use it when you need to know whether a URL gave you real content, a bot challenge, a JavaScript shell, a paywall or an error, with a receipt that records it.
Request
{ "url": "https://www.iana.org/help/example-domains", "return_text": true }
| Field | Required | Meaning |
|---|---|---|
url | yes | http or https URL, at most 2,048 characters |
return_text | no | also return up to 20,000 characters of clean text; default false |
Response
From the free sample (GET /v1/sample/fetch), shortened:
{
"page": {
"requested_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
"final_url": "https://vehcdj664efetfrsolne5umanq.srv.us/v1/sample/page",
"http_status": 200,
"content_type": "text/html",
"content_kind": "real",
"signals": ["status_200"],
"title": "Regional climate report 2025 (demo page)",
"redirects": [],
"content_sha256": "0e429607…",
"raw_sha256": "179b9ac4…",
"retrieved_at": "2026-10-09T14:19:34.909Z",
"bytes": 847,
"truncated": false
},
"text": "Regional climate report 2025\n\nThis is a demo page…",
"text_chars": 636,
"text_truncated": false,
"injection_flags": [],
"advice": "The page loaded and looks like real content.",
"receipt": "eyJhbGciOiJFZERTQSIs…",
"tier": "paid",
"rail": "x402",
"network": "eip155:84532"
}
content_kindandsignalsare explained in How it works.adviceis one plain sentence about the verdict.content_sha256hashes the full extracted text, not the 20,000-character copy.raw_sha256hashes the bytes as received.content_sha256is null when no body was read.injection_flagslists prompt-injection patterns found in the text, such asinstruction_overrideorhidden_instructions. They are flags, not a verdict: an empty list means none of our patterns matched.textis null unless you setreturn_text.
An application/pdf page is read too, within the same 2 MB and 10-second limits. Only its text layer is read, with no OCR, so a scanned PDF comes back as empty with the signal pdf_no_text_layer.
signalsincludespdf.titleis the PDF's own title, if it has one.texthas the pages in order, separated by a form feed (\f).- Only text inside each page's box is read, as a reader sees it. Text drawn outside the page is not in
text. page.pdfgivespage_count,pages_readandtruncated. Very long PDFs stop at a page or character cap, with the signalpdf_truncated. The receipt signs these as well.- A PDF we could not read is
unsupported_typewith a signal saying why:pdf_encrypted,pdf_invalid,pdf_over_budget,pdf_timeoutorpdf_out_of_memory. It is charged like any other page we cannot read. - When too many PDFs are waiting to be read, the call gets 503
over_capacity(details.reasonpdf_queue_full) withRetry-After. It is not charged.
Only the full service reads PDFs: an edge server answers 501 not_on_edge (details.reason pdf) for one, not charged. Limits are in Limits.
Charging
A bad URL, a blocked address or a host name that does not resolve is refused before payment and not charged. Results about the page itself, including 404s, bot walls and timeouts, are charged: that is the answer you paid for. See Payments.