THE PAGEFLOCK API / VERSION 1

A clear path
from request to result.

Search the web. Read a public page. Find published contacts. Every operation has one request shape, a visible cost and a response you can build with. Create an account and choose a paid plan to start. Copy the key when it appears; we only store its hash.

  1. Create a key.

    Create your account, complete email confirmation when required, then create a named key in API keys. Save its secret when it appears; it cannot be retrieved later. Choose an active paid subscription before creating your first key.

  2. Set two environment variables.

    Set WAGGLE_BASE_URL to the website origin you are using, without a trailing slash, and WAGGLE_API_KEY to your secret. Run these examples from your server or terminal.

  3. Inspect the result.

    A successful page extraction returns HTTP 200, request_id, data, credits_used: 1 and elapsed_ms. Keep the request ID; it connects your response to Activity.

curl --fail-with-body --include --max-time 45 \
  -X POST "$WAGGLE_BASE_URL/api/pageflock/v1/scrape" \
  -H "Authorization: Bearer $WAGGLE_API_KEY" \
  -H 'Content-Type: application/json' \
  --data '{
  "url": "https://example.com"
}'

This edits a code example. It does not send a request or use credits.

Run a request in your workspace

Python examples use the standard library. JavaScript examples target Node.js 20+. Go, Ruby and Java examples use their standard HTTP libraries. cURL examples include response headers and fail with the error body intact. The client timeout is 45 seconds; a lost response does not prove the request was never completed.

Send Authorization: Bearer YOUR_API_KEY and Content-Type: application/json with JSON POST requests. The base path is /api/pageflock/v1. Keys share their account’s credit balance, rate limits and concurrency allowance.

Existing API keys and the /api/waggle/v1 alias continue to work. Responses expose X-PageFlock-* headers alongside their X-Waggle-* compatibility names.

HeaderRequiredValue
AuthorizationYesBearer followed by your pageflock API secret.
Content-TypeFor JSON POST bodiesapplication/json
Idempotency-KeyOptional on the seven synchronous operationsA stable identifier for one logical request. See Retry protection.
Keep keys on your server.

Do not include secrets in browser bundles, query strings, saved configurations or public repositories. In API keys, rotate immediately or choose a 24-hour handover while you update your application. Revoked or expired keys return 401; requests already accepted can finish.

Follow the key handover guide

02 / ENDPOINT REFERENCE

Choose the job you need.

The seven operations below are synchronous. They return their result in the same HTTP response. Saving or configuring a request in the workbench does not run it.

POST/api/pageflock/v1/scrape1 credit / success

Readable text, links and metadata from a public page.

Static HTML costs 1 credit. Optional JavaScript rendering costs 5 where available. Login, private and blocked sources are unsupported.

Request body

FieldType / defaultWhat it means
urlstring · required—A complete public http:// or https:// URL, 8–2,048 characters. Standard web ports only; no embedded credentials or control characters.
mode"manual" | "auto""manual"Auto mode tries bounded static extraction first and escalates to JavaScript rendering only when the result looks like a client-rendered shell. It settles at the selected path cost; use max_cost=1 to guarantee a one-credit static request.
max_costinteger | nullnullAuto-mode ceiling: 1 or 5 credits. The five-credit ceiling permits rendering where the deployment has passed its browser-isolation check. Unused reserved credits are refunded when static extraction wins.
min_quality_scoreinteger | nullnullOptional 0–100 response-quality floor. When the final result scores below this value, the request returns 422 and uses zero credits. The score measures response shape, not factual accuracy.
extractobject{}Optional named fields for static HTML: {"heading":{"selector":"h1"},"links":{"selector":"a","attribute":"href","multiple":true}}. Up to 10 fields. Supports a single compound CSS selector combining a tag, class, ID or attribute presence/equality, such as a.product[href]. Traversal, selector lists, pseudo-classes and escapes are not supported. Included in the one-credit request.
renderbooleanfalseRun JavaScript in a network-isolated Chromium browser. Costs 5 credits on success; requires rendering availability on the deployment.
wait_msinteger500After-load wait in milliseconds, 0–3,000. Nondefault values require render=true.
wait_forstring | nullnullOptional CSS selector, 1–200 characters, to await in rendered mode. Total rendering deadline is 30 seconds.
screenshotbooleanfalseWhen render=true, include a bounded full-page PNG as base64 in data.screenshot. Screenshots are limited to 4 MiB and are unavailable when the deployment renderer is unavailable.
return_htmlbooleanfalseFor static extraction, include the bounded raw HTML source in data.html (up to 2 MB).
actionsarray[]Rendered browser actions: up to 20 click, fill or scroll steps. Click/fill use one simple CSS selector; fill values and scroll amounts are bounded. Arbitrary JavaScript, downloads, authentication and private pages are not supported.

Make the request

curl --fail-with-body --include --max-time 45 \
  -X POST "$WAGGLE_BASE_URL/api/pageflock/v1/scrape" \
  -H "Authorization: Bearer $WAGGLE_API_KEY" \
  -H 'Content-Type: application/json' \
  --data '{
  "url": "https://example.com"
}'

Illustrative JSON with shortened content. Fields in the table below are inside data; source content and timing will vary.

View an example HTTP 200 response
{
  "request_id": "00000000-0000-4000-8000-000000000001",
  "data": {
    "url": "https://example.com/",
    "title": "Example Domain",
    "text": "Example Domain\nThis domain is for use in illustrative examples in documents.",
    "links": [
      {
        "text": "Learn more",
        "url": "https://iana.org/domains/example"
      }
    ],
    "metadata": {
      "description": null,
      "language": "en",
      "quality": {
        "score": 58,
        "grade": "usable",
        "mode": "static"
      }
    },
    "evidence": {
      "source_url": "https://example.com/",
      "retrieved_at": "2026-09-17T00:00:00+00:00",
      "content_sha256": "…"
    },
    "status_code": 200
  },
  "credits_used": 1,
  "elapsed_ms": 840
}
FieldTypeWhat it means
extracted / extractionobjectWhen extract is provided, named values contain text or raw attributes; missing matches are null (single) or [] (multiple). Up to 50 matches per field, 2,000 characters per value and 50,000 characters total, allocated in alphabetical field-name order for deterministic retries. extraction.truncated_fields names shortened fields. Attributes are not resolved or fetched. Static mode only.
urlstringFinal source URL after checked redirects.
titlestring | nullText from the page title, or null when absent.
textstringReadable page text, up to 100,000 characters. Script, style, noscript and template elements are removed.
linksarrayUp to 500 links inspected from the page; only HTTP/HTTPS links are returned.
links[].text / urlstringAnchor text up to 300 characters and an absolute URL up to 2,048 characters. A returned link is not a promise that a later fetch is allowed.
metadata.description / languagestring | nullThe description meta tag and HTML lang attribute, when present.
status_codeintegerThe successful HTTP status returned by the source.
metadata.qualityobjectExplainable 0–100 response-shape signal with a strong, usable or thin grade, readable-character/link counts and pass/fail checks. It describes content shape; it does not verify truth or freshness.
evidenceobjectSource URL, retrieval timestamp and SHA-256 fingerprint of the HTML pageflock received. Keep this beside exported data for traceability.
screenshotobject | absentWhen requested with render=true, contains a full-page image/png encoded as base64. The image is bounded to 4 MiB.

Behavior to build around

Static mode costs 1 credit. Set render=true for JavaScript extraction at 5 credits, or mode=auto to try static first and settle at 1 or 5 credits based on the selected path. Rendering is only attempted where the deployment has passed its browser-isolation self-test; an unavailable renderer leaves auto mode on static HTML and charges one credit.

Rendering uses a fresh cookie-free browser with a 30-second deadline, 40 broker requests and bounded input/output. Public HTTP GET requests pass source, robots and DNS checks; blocked subresources are counted in metadata.rendering. An opt-in full-page PNG screenshot is available with screenshot=true. Bounded click, fill and scroll actions are available through actions; custom cookies, proxy configuration, logins, downloads and arbitrary JavaScript are not supported. Collections still use static extraction.

Source refusals, non-success source responses and fetch failures return 422 with zero credits. Set min_quality_score when you also want a thin final response rejected with zero credits. The quality floor measures response shape and does not verify factual accuracy.

POST/api/pageflock/v1/contacts1 credit / success

Published emails, phone contacts and social links, with sources.

One page per request, up to 50 contacts per category. Published addresses are not mailbox-verified. Suggested pages are never fetched automatically.

Request body

FieldType / defaultWhat it means
urlstring · required—A complete public http:// or https:// URL, 8–2,048 characters. Standard web ports only; no embedded credentials or control characters.

Make the request

curl --fail-with-body --include --max-time 45 \
  -X POST "$WAGGLE_BASE_URL/api/pageflock/v1/contacts" \
  -H "Authorization: Bearer $WAGGLE_API_KEY" \
  -H 'Content-Type: application/json' \
  --data '{
  "url": "https://www.w3.org/Consortium/contact"
}'

Illustrative JSON with shortened content. Fields in the table below are inside data; source content and timing will vary.

View an example HTTP 200 response
{
  "request_id": "00000000-0000-4000-8000-000000000001",
  "data": {
    "url": "https://example.com/contact",
    "checked_at": "2026-09-15T10:00:00+00:00",
    "pages_checked": 1,
    "emails": [
      {
        "value": "hello@example.com",
        "found_in": "mailto",
        "verification": "not_checked",
        "source_url": "https://example.com/contact"
      }
    ],
    "phones": [],
    "socials": [],
    "contact_pages": [],
    "found": true,
    "truncated": false,
    "scope": "Published contacts from this page's static HTML and organization markup. Mailboxes, phone ownership and social profiles are not verified. Linked pages are not fetched; each follow-up costs a separate request."
  },
  "credits_used": 1,
  "elapsed_ms": 840
}
FieldTypeWhat it means
url / checked_atstringFinal fetched source URL and UTC extraction timestamp.
pages_checkedintegerAlways 1. Suggested contact pages are not followed automatically.
emails[]objectvalue, found_in, verification and source_url. verification is always not_checked. Deduplicated email addresses; up to 50.
phones[]objectvalue, normalized, found_in, verification and source_url. Only explicit telephone links or supported organization markup; up to 50.
socials[]objectplatform, url, found_in and source_url; up to 50 public links found in the page. Social profiles are not fetched or ownership-verified.
contact_pages[]objecturl, label, fetched: false and source_url for suggested contact pages; up to 50. Each follow-up is a separate request.
foundbooleanTrue when at least one email, phone or social link was found. Suggested pages alone do not make this true.
truncatedbooleanTrue when an extraction limit was reached. It does not count additional contacts or guarantee complete coverage when false.
scopestringA reminder of the extraction and verification boundaries.

Behavior to build around

A published address is source evidence, not proof that a mailbox accepts mail. pageflock does not guess addresses or perform SMTP verification through this endpoint.

Keep source_url and checked_at beside every contact in your own application. found_in identifies the evidence channel, such as mailto, page_text, tel or supported structured markup.

A successfully fetched page with no contacts returns empty arrays and found: false. It still uses one credit. Missing JavaScript-only contact widgets are not rendered.

POST/api/pageflock/v1/products1 credit / success

Product details from published data on pages that allow direct access.

Reads published Product JSON-LD, or a selected Amazon, Walmart or Shopee product page. Some marketplaces reject direct requests; blocked or unavailable pages return an error without a credit charge. Missing values stay empty.

Request body

FieldType / defaultWhat it means
urlstring · required—A complete public http:// or https:// URL, 8–2,048 characters. Standard web ports only; no embedded credentials or control characters.
marketplace"amazon" | "walmart" | "shopee" | nullnullChoose a dedicated public product-page parser. The URL must belong to the selected marketplace. Omit for generic Product JSON-LD.

Make the request

curl --fail-with-body --include --max-time 45 \
  -X POST "$WAGGLE_BASE_URL/api/pageflock/v1/products" \
  -H "Authorization: Bearer $WAGGLE_API_KEY" \
  -H 'Content-Type: application/json' \
  --data '{
  "url": "https://scrapeme.live/shop/Bulbasaur/"
}'

Illustrative JSON with shortened content. Fields in the table below are inside data; source content and timing will vary.

View an example HTTP 200 response
{
  "request_id": "00000000-0000-4000-8000-000000000001",
  "data": {
    "url": "https://example.com/products/notebook",
    "products": [
      {
        "name": "Field notebook",
        "description": null,
        "sku": "NOTE-01",
        "brand": "Example",
        "offers": [
          {
            "price": "12.00",
            "currency": "USD",
            "availability": "https://schema.org/InStock",
            "url": "https://example.com/products/notebook"
          }
        ]
      }
    ],
    "source": "published_json_ld",
    "scope": "Up to 20 published products and 20 offers each. Missing fields stay null; prices are not independently verified."
  },
  "credits_used": 1,
  "elapsed_ms": 840
}
FieldTypeWhat it means
urlstringFinal fetched product-page URL.
productsarrayUp to 20 published Product JSON-LD records in generic mode, or one identified product in marketplace mode. Each product has a nonempty name.
products[].namestringPublished name, up to 1,000 characters.
products[].description / sku / brandstring | nullPublished values when supported; otherwise null. Limits are 10,000 / 200 / 500 characters respectively.
products[].offersarrayUp to 20 published offers per product. No independent price or stock verification.
offers[].price / currency / availability / urlstring | nullPrice is returned as a string; currency and availability preserve the source values. A supported offer URL must be HTTPS. Do not infer missing price, currency or availability.
source / scopestringsource is published_json_ld; scope describes the extraction boundary.

Behavior to build around

Marketplace mode parses supported public HTML and embedded product records, adding product_id, images, rating, review_count and per-field provenance. Full public product URLs are required. Marketplace search, seller history and live inventory guarantees are not included.

A page without a supported named product, a mismatched product ID, a blocked page or an app shell returns 422 and uses zero credits. Missing fields remain null. Source layouts can change. Walmart has limited live acceptance; Amazon and Shopee require further live coverage before broad reliability claims.

POST/api/pageflock/v1/youtube1 credit / success

Public video details and exact engagement when published.

Public, embeddable videos only. Extended metadata, captions and related videos are returned when available and permitted. Missing engagement counts are labelled, not estimated.

Request body

FieldType / defaultWhat it means
urlstring · required—HTTPS YouTube video URL, 8–2,048 characters, with an 11-character video ID. Accepts watch?v=, youtu.be, shorts/, embed/ and live/ URLs on supported YouTube hosts. No custom port.

Make the request

curl --fail-with-body --include --max-time 45 \
  -X POST "$WAGGLE_BASE_URL/api/pageflock/v1/youtube" \
  -H "Authorization: Bearer $WAGGLE_API_KEY" \
  -H 'Content-Type: application/json' \
  --data '{
  "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw"
}'

Illustrative JSON with shortened content. Fields in the table below are inside data; source content and timing will vary.

View an example HTTP 200 response
{
  "request_id": "00000000-0000-4000-8000-000000000001",
  "data": {
    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "video_id": "jNQXAC9IVRw",
    "title": "Me at the zoo",
    "channel": {
      "name": "jawed",
      "url": "https://www.youtube.com/@jawed"
    },
    "thumbnail_url": "https://i.ytimg.com/vi/jNQXAC9IVRw/hqdefault.jpg",
    "source": "youtube_oembed",
    "scope": "Public embed metadata: title, channel and thumbnail. Extended public details were unavailable for this video."
  },
  "credits_used": 1,
  "elapsed_ms": 840
}
FieldTypeWhat it means
url / video_idstringCanonical https://www.youtube.com/watch?v= URL and video identifier.
titlestringTitle returned by public embed metadata, up to 2,000 characters.
channel.name / channel.urlstring | nullChannel name and supported HTTPS YouTube channel link, when supplied.
thumbnail_urlstring | nullSupported HTTPS i.ytimg.com thumbnail URL, when supplied. The image itself is not downloaded.
metricsobjectExact views, likes, dislikes and comments when the public source supplies them. Engagement percentage is calculated only when views, likes and comments are all available; otherwise counts remain null.
transcript / related_videos / other_channel_videosobject / arraysOptional permitted public caption text and videos named by the source. Transcript status and reason explain when no public caption track was available.
source / scopestringsource identifies the public data source; scope describes which metadata was available.

Behavior to build around

pageflock starts with public oEmbed metadata. It may add available public watch-page metadata, exact engagement counts, related videos and a permitted public caption track; missing fields remain null or explicitly unavailable.

Private, removed or non-embeddable videos return 422 and use zero credits. pageflock does not use a logged-in watch page or download media.

GET/api/pageflock/v1/usageNo acquisition credits

Use the same Bearer key. This read-only request returns your account’s current balance and effective limits. Read these values instead of hard-coding the allowances of a pricing tier.

curl --fail-with-body --include --max-time 45 \
  -X GET "$WAGGLE_BASE_URL/api/pageflock/v1/usage" \
  -H "Authorization: Bearer $WAGGLE_API_KEY"
FieldTypeWhat it means
credits_remainingintegerCurrent account balance. In-flight operations can hold reserved credits until they finish or are refunded.
planstringEffective plan identifier from the current paid subscription or retained legacy entitlement. A pending checkout does not change it.
limits.requests_per_minuteintegerShared admission limit for metered operations across your keys and dashboard.
limits.concurrent_requestsintegerMaximum simultaneous accepted operations, including collection work.
limits.pages_per_collection / active_collectionsintegerMaximum pages in one collection and simultaneous active collections.
overviewobjectSeven UTC days of request, completion, failure, running and credit totals, plus success_rate and average_ms. A rate or average is null when there is no applicable completed work.
capabilitiesobjectOperation availability, credit costs and scope. An available operation still depends on a valid request and an eligible source; it is not an uptime guarantee.

A batch processes a list of URLs. A crawl follows eligible links from one exact origin. Collections use static page extraction and continue after you close the browser. Search queries and dedicated contact/product/video output are separate synchronous operations.

POST/api/pageflock/v1/collections202 Accepted
FieldType / defaultWhat it means
namestring"Untitled collection"Collection label, 1–100 characters.
mode"batch" | "crawl""batch"Batch processes the supplied URLs. Crawl discovers links from exactly one seed URL within the same origin.
urlsstring[] · required—1–25 public URLs, further restricted by your plan. A crawl requires exactly one URL. Batch URLs are deduplicated.
max_pagesinteger5Crawl page ceiling, 1–25 and within your plan. Batch size is determined by the deduplicated URL list.
client_idstringGenerated1–64 characters. Supply a stable identifier to retry the same collection submission. Reusing it returns the existing account-owned collection; do not reuse it for changed work.
curl --fail-with-body --include --max-time 45 \
  -X POST "$WAGGLE_BASE_URL/api/pageflock/v1/collections" \
  -H "Authorization: Bearer $WAGGLE_API_KEY" \
  -H 'Content-Type: application/json' \
  --data '{
  "name": "My research",
  "mode": "crawl",
  "urls": [
    "https://example.com"
  ],
  "max_pages": 5,
  "client_id": "my-unique-run-001"
}'

The response contains an id, status, counts, credits_used and expires_at. Keep its ID and poll every four seconds or slower. The collection body uses client_id for duplicate submissions; the synchronous Idempotency-Key header does not apply here.

MethodPath after /api/pageflock/v1Behavior
GET/collectionsLatest 50 account-owned collections, worker_available and retention_days.
GET/collections/{id}Progress and pages with URL, status, error and per-page request_id.
GET/collections/{id}/resultsPaginated JSON: page starts at 1; page_size defaults to 100 (maximum 100). Includes pagination totals. Add download=true for a streamed full JSON export.
GET/collections/{id}/results?format=csvDownload URL, title, text and status_code columns with formula-safe cells.
POST/collections/{id}/cancelStop pending pages; an already-running page may finish and be charged on success.
DELETE/collections/{id}Permanently delete a terminal collection and its results. Cancel active work and wait first.
StateMeaning
queuedAccepted and waiting for a worker.
runningThe worker is processing or discovering pages.
completedEligible work completed. A crawl can finish below max_pages when it finds no more links.
partialSome pages completed and some failed.
failedNo successful completed dataset; inspect page errors.
cancelledCancellation stopped further processing; successful pages remain available until expiry.

Each successful page costs one credit. Failed or interrupted page reservations are refunded; pending cancelled pages are not charged. There are at most 50 saved collections per account. Results expire seven days after collection creation, so download them before expires_at.

Collections and retention guide

All seven synchronous operations use the success envelope below. Errors use {"error":{"message":"…"}}. Do not expect a success-shaped data object when the HTTP status is an error.

FieldTypeWhat it means
request_idstringpageflock's ID for the metered operation. Keep it with the response and use it to locate the request in Activity or ask for help.
dataobjectEndpoint-specific fields documented below. Example JSON illustrates the schema; it is not a live result or a completeness guarantee.
credits_usedintegerThe original operation's credit charge: 1 or 5 on success. On an idempotent replay this body field stays unchanged; use the response header to see that the replay adds zero credits.
elapsed_msintegerDuration of the original operation in milliseconds. Source response times vary; this is not an SLA.
HeaderHow to use it
X-PageFlock-Credits-UsedCredits charged by this HTTP call, when present. An idempotent replay uses 0 even though its unchanged JSON body records the original charge.
X-PageFlock-Request-IdMetered request ID when one has been created. Early authentication, validation or admission failures can have no ID. The success body also carries request_id.
Retry-AfterSeconds to wait when a rate/concurrency limit or pending duplicate response supplies it.
X-PageFlock-Idempotency-Replayedtrue when a retained result is replayed without another acquisition.
X-PageFlock-Idempotency-Statuscreated, replayed, in_progress, conflict, unavailable or capacity. Inspect it alongside the HTTP status and error message.
X-PageFlock-Idempotency-Expires-AtUTC expiry of the retry receipt, 24 hours after the original claim.
Cache-Controlno-store; keep your own response copy if your application needs it.

Find the original operation by request ID in Activity. History stores request metadata and inputs; it does not ordinarily retain the whole synchronous response. Saved requests store configurations and never run automatically.

For the seven synchronous POST operations, send an optional Idempotency-Key before the first request. Generate a UUID for each new logical operation. Keep that key with its exact request inputs and reuse it if the connection drops.

Same key, same operation, same validated inputs.

The key is scoped to your account across API keys and the dashboard. It must be 16–128 ASCII letters, digits or _ . : -. Payload defaults are normalized before comparison. A different operation or payload needs a new key.

SituationWhat happensNext step
Same request already completedWithin 24 hours, the retained status, body and original request ID are returned. No second acquisition charge.Read X-PageFlock-Idempotency-Replayed: true and X-PageFlock-Credits-Used: 0. The body’s credits_used remains the original charge.
Same request still running409 with Retry-After: 2.Wait, then retry the same key and unchanged inputs.
Key reused with different work409 conflict; the changed operation is not executed.Compare the saved inputs. Use a new key only for an intentional new operation.
Result cannot be retainedThe first response is delivered, but a later replay receives a 410 unavailable receipt.Check Activity and your own response store. This key never triggers an automatic refetch.
Request interrupted in the serviceRecovery records a terminal interrupted outcome and refunds the original reservation once.Review its status. A new acquisition requires a deliberate new key.
Original operation failed after its claimThe original error status and message are retained and replayed too.Reusing the key retrieves that outcome. Use a new key only when deliberately attempting the operation again.
Receipt capacity reached429 with Retry-After: 3600 and status capacity. Existing retained keys can still be replayed.Wait for receipts to expire; do not remove retry protection just to force another acquisition.
Older than 24 hoursAfter 24 hours from the original claim, an expired key may execute as new work.Do not retry an old job automatically. Check stored results and create new work deliberately.

Opting in allows pageflock to retain the request and its response for up to 24 hours: up to 2 MiB per response, 16 MiB of retained response data per account, and 1,000 unexpired receipts per account. The raw idempotency key is stored as a hash. Account closure removes these receipts and results.

Without this header, each successful submission is a new billable operation. A client timeout is ambiguous: inspect Activity before submitting again. Replays still count toward your account request-rate limit, but do not consume provider or destination acquisition budgets. Even with retry protection, use a bounded retry count and honor Retry-After; changing the key on every retry defeats duplicate protection.

Use GET usage for your effective limits. Rate is how many operations can begin in a time window; concurrency is how many can be running together. They are separate controls shared across account keys, the dashboard and collection workers.

BoundaryCurrent behavior
Account rate10–240 metered requests per minute, depending on the effective plan.
Account concurrency1–12 simultaneous operations, depending on the effective plan.
Destination rate30 requests per minute per destination domain, shared across accounts for URL operations. Search also has deployment-wide budgets.
Request body8 KiB maximum for public JSON request bodies.
Page fetch30-second execution timeout; source response size is bounded (1 MB by default).
Static outputUp to 100,000 text characters and 500 inspected links per page.
Search outputUp to 10 normalized results per call; fewer results do not reduce the charge.
Collections10–1,000 pages and 1–20 active collections, depending on plan; at most 50 saved collections.
OutcomeCharge
Static extraction, contacts, products or YouTube metadata succeeds1 credit
JavaScript page rendering succeeds (where available)5 credits
Web, news or video search succeeds5 credits, including empty results
Contact page succeeds with no emails, phones or social links1 credit; found is false
Product page has no supported named product data0 credits; 422 response
Invalid input, blocked source, failed page or provider failure0 credits; a reservation made for the failed operation is refunded
Retained idempotent replay0 additional credits
Configure, save, or reopen a request0 credits; explicit execution is separate

Estimate a workload by mixing one-credit requests and five-credit searches. Count calls, not result rows. Paid plans renew monthly and add credits after each successful payment. Unused credits roll over while your account remains open. Cancel before the next renewal from Plans & billing. The estimator selects one monthly plan; larger workloads need an Enterprise quote. Existing purchased credits are retained without automatically subscribing you.

First inspect the HTTP status, error.message, and any request ID or retry header. An error message can describe the failed field or source policy; do not retry every 422 or rotate keys to work around a limit.

{
  "error": {
    "message": "Check these fields: limit"
  }
}
StatusMeaningUseful next step
401Missing, expired, revoked or invalid authentication.Check the Bearer header and key handover deadline. Use an active key for this account.
402Insufficient available credits for the operation.Read usage. A search or page extraction requires one available credit.
403Account confirmation or allowed-origin checks were not satisfied.Confirm your email when required. Call from your server and do not forward a foreign browser Origin header.
404The requested account-owned resource was not found.Check the collection ID and account; expired or deleted resources may no longer exist.
409An operation conflicts with account/collection state, or an idempotency key is pending or mismatched.Read the message and retry header. Wait for pending work, resolve the state or correct the inputs.
410An idempotency receipt exists but its response is unavailable.Check Activity and your stored output. Do not silently replace the key and execute again.
413The JSON request body exceeds 8 KiB.Reduce the body; keep collection URLs within the documented limits.
422Invalid fields, unsupported URL, source refusal or an extraction failure.Correct the field, inspect the page’s static source or choose an eligible public source. Failed URL extraction costs zero.
429A rate, concurrency, deployment budget or retry-receipt capacity was reached.Honor Retry-After. Reduce concurrency or request frequency; all account keys share limits.
502Search provider execution failed.The search reservation is refunded. A retained error replays with the same key; deliberately attempting fresh work requires a new key.
503A required service, search provider configuration or collection worker is unavailable.Check the status page. A transport failure can be ambiguous: preserve the idempotency key and inspect Activity before new work.

pageflock checks source decisions, network destinations, robots.txt for page extraction, and each redirect destination. It refuses local/private networks, credentials in URLs, unsupported ports and blocked sources. It does not bypass logins, paywalls or CAPTCHAs, or fetch LinkedIn profile pages. A search-result URL or a social link is not authorization to scrape that page.

Contacts carry source_url and found_in; keep their evidence and extraction timestamp with downstream records. Published contact information is not verified mailbox, phone or social ownership.

DataRetention and access
Synchronous request metadata30 days, including operation, URL/query, saved inputs, status, charge, timing and request ID. Full response bodies are not ordinarily retained.
Opt-in idempotency resultsUp to 24 hours for retry protection, within documented storage limits.
Collection resultsSeven days from collection creation; download JSON or CSV before expiry.
Saved request configurationsAccount-owned inputs retained until removed or the account is closed. They contain no saved API secret and do not execute on open.
Account exportAccount details, key metadata, saved configurations, purchases, collections, bounded recent request metadata and retained retry receipts/results. Key secrets, password hashes and session cookies are excluded.

Settings provides account export and closure. Exporting does not extend a collection’s expiry or include every synchronous response. Keep your own copy of important results and choose an appropriate retention policy for contact information in your application.

READY FOR YOUR FIRST REAL REQUEST?

Take it into the workbench.

Choose an operation, inspect the cost, run deliberately and keep the result.

Open your workspace Browse practical guides
YOUR PRIVACY

Choose what works for you. You can reopen these settings from the footer at any time.