FIELD NOTES / THE BASICS

Search, extract or crawl? Start with the right tool.

Three different jobs. One useful workflow. A guide to choosing your starting point.

← All field notes
01

Start with what you know.

If you have a question but no URLs, start with search. If you already know the page you want, start with extraction. If you have a list of pages or want to follow links within one site, create a collection. Choosing the right starting point keeps your workflow understandable and avoids unnecessary requests.

In pageflock, web search returns result titles, links and snippets — and, where Google surfaces them, a knowledge panel, a direct answer, related questions and a local business pack with phone, address and website — through an independently configured Serper connection. A result is a lead to investigate; the snippet is not the full page, and its wording can differ from the current source.

02

Read a page, then decide what to keep.

The pageflock Extraction Engine accepts a public URL and returns its title, readable text, links and metadata. It reads static HTML, with optional JavaScript rendering when a page needs it; login-only pages are outside this workflow. Dedicated extractors go further on a known page: contacts, product data with images and ratings, YouTube metadata with the description and stats, and a company profile that reads a whole company's public pages in one call.

Keep the source URL beside anything you save. Check the actual text before treating it as input to a report or model. For a search-led workflow, inspect the results first and extract only the relevant URLs.

03

Move to a collection when one page is not enough.

Use a batch when you already have your URLs. Use a same-origin crawl when the useful pages are connected within one website. A collection runs in the background and records each page's result. Your browser does not need to stay open.

Download the completed results as JSON or CSV within seven days. A successful search costs five credits, and each successful extracted page costs one. One search followed by three successful extractions therefore uses eight credits. Failed or blocked extraction is refunded.

Choose your API ↗
YOUR PRIVACY

Choose what works for you. You can reopen these settings from the footer at any time.