# How it works

Last updated: 2026-10-10

How AttestPage fetches a page, what each `content_kind` verdict means, and what a signed receipt records about the fetch.

## The fetcher

When a route reads a page, our fetcher requests the URL you sent:

- GET only, with no cookies, logins or headers passed through from you. User agent `QuoteproofBot/1.0 (+https://vehcdj664efetfrsolne5umanq.srv.us/bot)`.
- http and https on ports 80 and 443. Private, loopback and other internal addresses are refused; the address is resolved once and the connection is pinned to it, and each redirect is checked again.
- Up to 5 redirects, 5 seconds to connect, 10 seconds in all, 2 MB of body.
- HTML, XHTML, plain text and JSON. PDFs too, for `fetch`, `verify/quote`, `verify/quotes` and `verify/citations`: the text layer only (no OCR, so a scanned page has no text), and each match gives its page number. No JavaScript is run.
- robots.txt and our opt-out list are honoured; see [Our fetcher](/bot).
- No proxies and no attempt to get past CAPTCHAs or bot checks. A challenge page is reported as `bot_wall`.

Pages are not stored. The text is extracted, hashed, matched, and dropped when the response is sent.

## content_kind

Every page result has one `content_kind`, from automated rules over the status, headers and text. They can be wrong; `signals` lists the clues behind each verdict.

| content_kind | Meaning |
|---|---|
| `real` | The page loaded and looks like real content. |
| `bot_wall` | The site served a bot challenge to our fetcher. Do not treat this as the page's content. |
| `js_shell` | The page needs JavaScript to show its content; the text we read is mostly empty scaffolding. |
| `paywall` | The page appears to be behind a paywall or login; the text may be a teaser only. |
| `empty` | The page loaded but has little or no readable text. |
| `off_site` | The URL redirected to a different site; the content is from that site. In check/links, `off_site_error: true` marks one whose site answers 4xx/5xx. |
| `http_error` | Error status, or a connection problem (signals such as `outcome_dns_error`, `outcome_connect_error`). |
| `unsupported_type` | Not HTML, text, JSON or a readable PDF, so not read. A PDF we could not read has the signal `pdf` and one of `pdf_encrypted` (needs a password), `pdf_invalid` (damaged or not a PDF), `pdf_over_budget` (needs more memory than we allow), `pdf_timeout` or `pdf_out_of_memory`. |
| `pdf` | check/links only: a live PDF link (status 2xx, `application/pdf`, same site), not read. In check/links an unread link with a 4xx/5xx status is `http_error` and one redirected to another site is `off_site` (with `off_site_error: true` if that site answers 4xx/5xx). |
| `robots_disallowed` | robots.txt or an opt-out disallows our fetcher. |
| `timeout` | The site did not answer in time. |
| `too_large` | Over the 2 MB cap, so not read in full. |

## injection_flags

Patterns often used for prompt injection, found in the page text: `instruction_override`, `role_reassignment`, `prompt_markers`, `prompt_extraction`, `addresses_ai`, `command_execution`, `exfiltration_request`, plus `hidden_instructions` when a match is in text a browser does not show. They are flags, not a verdict. An empty list means none of our patterns matched.

## Receipts

Every paid answer carries a `receipt`: a compact JWS signed with Ed25519 (`alg` EdDSA, `typ` `attestpage-evidence+jws`). The key id is `https://still-rapids-9yt7.here.now/.well-known/jwks.json#ap-<thumbprint>`, where `https://still-rapids-9yt7.here.now` is the fixed issuer address (it stays the same when the service address changes); the public keys are at [/.well-known/jwks.json](https://vehcdj664efetfrsolne5umanq.srv.us/.well-known/jwks.json) and [/.well-known/did.json](https://vehcdj664efetfrsolne5umanq.srv.us/.well-known/did.json).

Payload fields (version 1):

| Field | Meaning |
|---|---|
| `v` | schema version, 1 |
| `iss` | `https://still-rapids-9yt7.here.now` |
| `iat` | signing time, Unix seconds |
| `kind` | `quote`, `quotes`, `citations`, `document`, `fetch`, `links`, `packages` or `attest` |
| `request_sha256` | SHA-256 of your request body as canonical JSON |
| `quote_sha256` | SHA-256 of the quote (quote receipts) |
| `url`, `final_url`, `http_status`, `content_kind` | what was requested and what came back |
| `content_sha256` | SHA-256 of the full extracted text |
| `raw_sha256` | SHA-256 of the body bytes as received |
| `retrieved_at` | when the page was fetched |
| `server_ip` | the vetted IP address our fetcher connected to for the final response (the site's address, not yours) |
| `tls` | the site's certificate on the final response over HTTPS: `cert_sha256` (SHA-256 of the DER certificate), `issuer`, `subject` (each up to 1,024 bytes), `valid_from`, `valid_to`; null over plain HTTP |
| `headers` | these response headers when the site sent them: `date`, `last-modified`, `etag`, `content-type`, `content-length` (each up to 512 bytes) |
| `result` | the route's result (match, signals, per-link results, per-page and per-citation results, or attestation) |
| `tier` | `paid`, `trial` or `sample` |
| `payment` | paid calls: rail, network, asset, amount, payer, payment_id (L402: payer null, payment_hash and preimage); otherwise null |

`server_ip`, `tls` and `headers` are on fetch and quote receipts and on each entry of a links receipt and each page of a citations receipt. They are null when no response came back, and always null on an edge server, which cannot pin the connection to an IP or read the certificate. Those byte caps count the value as UTF-8 once JSON-encoded (a `"` or `\` counts 2, a CJK character 3); a longer value is cut on a whole character and ends with `…`. Receipts issued before these fields existed do not have them and still verify.

A receipt records what our fetcher saw at that time. Pages change, and a page can show different content to different visitors. It does not show that anything the page says is correct.

Check receipts with [POST /v1/receipt/verify](/docs/receipt-verify) or [offline](/docs/verify-offline).
