Choose how pages are gathered
A batch processes a list of URLs you provide. A crawl starts from one page and follows permitted links on the same exact website origin. Choose a maximum page count within your account’s collection limit.
Plans allow 10–1,000 pages per collection. Multiple collections can progress together; each collection processes one page at a time. Collection workers share your account’s rate and concurrency limits with individual requests. Each source is checked before it is fetched. The worker continues when you leave the browser, so monitor the collection’s progress instead of resubmitting it.
Watch each page’s outcome
A collection can contain successful and failed pages. Each successful page uses one credit; failed or blocked pages do not. Inspect the page states when the overall collection finishes instead of assuming every page succeeded.
Cancel when you want to stop queued work. A page that is already running may still complete and use its credit. When retrying a collection submission after a connection problem, use the same client_id to avoid creating another collection.
Download before seven days
Open the collection and download JSON or CSV while the content is available. The collection displays its result expiry. Saved page content is kept for seven days; request metadata has a separate 30-day retention window.
Your account export contains account and collection metadata, not the collected page text. If your workflow needs a permanent dataset, store the collection download in your own system and choose an appropriate retention policy there.