How it works
Last updated: 2026-10-10
How AttestPage fetches a page, what each content_kind verdict means, and what a signed receipt records about the fetch.
The fetcher
When a route reads a page, our fetcher requests the URL you sent:
- GET only, with no cookies, logins or headers passed through from you. User agent
QuoteproofBot/1.0 (+https://vehcdj664efetfrsolne5umanq.srv.us/bot). - http and https on ports 80 and 443. Private, loopback and other internal addresses are refused; the address is resolved once and the connection is pinned to it, and each redirect is checked again.
- Up to 5 redirects, 5 seconds to connect, 10 seconds in all, 2 MB of body.
- HTML, XHTML, plain text and JSON. PDFs too, for
fetch,verify/quote,verify/quotesandverify/citations: the text layer only (no OCR, so a scanned page has no text), and each match gives its page number. No JavaScript is run. - robots.txt and our opt-out list are honoured; see Our fetcher.
- No proxies and no attempt to get past CAPTCHAs or bot checks. A challenge page is reported as
bot_wall.
Pages are not stored. The text is extracted, hashed, matched, and dropped when the response is sent.
content_kind
Every page result has one content_kind, from automated rules over the status, headers and text. They can be wrong; signals lists the clues behind each verdict.
| content_kind | Meaning |
|---|---|
real | The page loaded and looks like real content. |
bot_wall | The site served a bot challenge to our fetcher. Do not treat this as the page's content. |
js_shell | The page needs JavaScript to show its content; the text we read is mostly empty scaffolding. |
paywall | The page appears to be behind a paywall or login; the text may be a teaser only. |
empty | The page loaded but has little or no readable text. |
off_site | The URL redirected to a different site; the content is from that site. In check/links, off_site_error: true marks one whose site answers 4xx/5xx. |
http_error | Error status, or a connection problem (signals such as outcome_dns_error, outcome_connect_error). |
unsupported_type | Not HTML, text, JSON or a readable PDF, so not read. A PDF we could not read has the signal pdf and one of pdf_encrypted (needs a password), pdf_invalid (damaged or not a PDF), pdf_over_budget (needs more memory than we allow), pdf_timeout or pdf_out_of_memory. |
pdf | check/links only: a live PDF link (status 2xx, application/pdf, same site), not read. In check/links an unread link with a 4xx/5xx status is http_error and one redirected to another site is off_site (with off_site_error: true if that site answers 4xx/5xx). |
robots_disallowed | robots.txt or an opt-out disallows our fetcher. |
timeout | The site did not answer in time. |
too_large | Over the 2 MB cap, so not read in full. |
injection_flags
Patterns often used for prompt injection, found in the page text: instruction_override, role_reassignment, prompt_markers, prompt_extraction, addresses_ai, command_execution, exfiltration_request, plus hidden_instructions when a match is in text a browser does not show. They are flags, not a verdict. An empty list means none of our patterns matched.
Receipts
Every paid answer carries a receipt: a compact JWS signed with Ed25519 (alg EdDSA, typ attestpage-evidence+jws). The key id is https://still-rapids-9yt7.here.now/.well-known/jwks.json#ap-<thumbprint>, where https://still-rapids-9yt7.here.now is the fixed issuer address (it stays the same when the service address changes); the public keys are at /.well-known/jwks.json and /.well-known/did.json.
Payload fields (version 1):
| Field | Meaning |
|---|---|
v | schema version, 1 |
iss | https://still-rapids-9yt7.here.now |
iat | signing time, Unix seconds |
kind | quote, quotes, citations, document, fetch, links, packages or attest |
request_sha256 | SHA-256 of your request body as canonical JSON |
quote_sha256 | SHA-256 of the quote (quote receipts) |
url, final_url, http_status, content_kind | what was requested and what came back |
content_sha256 | SHA-256 of the full extracted text |
raw_sha256 | SHA-256 of the body bytes as received |
retrieved_at | when the page was fetched |
server_ip | the vetted IP address our fetcher connected to for the final response (the site's address, not yours) |
tls | the site's certificate on the final response over HTTPS: cert_sha256 (SHA-256 of the DER certificate), issuer, subject (each up to 1,024 bytes), valid_from, valid_to; null over plain HTTP |
headers | these response headers when the site sent them: date, last-modified, etag, content-type, content-length (each up to 512 bytes) |
result | the route's result (match, signals, per-link results, per-page and per-citation results, or attestation) |
tier | paid, trial or sample |
payment | paid calls: rail, network, asset, amount, payer, payment_id (L402: payer null, payment_hash and preimage); otherwise null |
server_ip, tls and headers are on fetch and quote receipts and on each entry of a links receipt and each page of a citations receipt. They are null when no response came back, and always null on an edge server, which cannot pin the connection to an IP or read the certificate. Those byte caps count the value as UTF-8 once JSON-encoded (a " or \ counts 2, a CJK character 3); a longer value is cut on a whole character and ends with …. Receipts issued before these fields existed do not have them and still verify.
A receipt records what our fetcher saw at that time. Pages change, and a page can show different content to different visitors. It does not show that anything the page says is correct.
Check receipts with POST /v1/receipt/verify or offline.