Give the collection a purpose.
Write down the question your dataset should help answer. A tightly scoped collection is easier to review than a folder of unrelated pages. For example, gather a project's public documentation to support a knowledge base, with the source URL attached to each extracted page.
Try a single page in the extractor first. Check that the relevant content exists in the returned static HTML. A blank or incomplete result is a reason to inspect the source, not to send a much larger batch.
Gather a bounded set of pages.
Open Collections in your workspace. Choose a batch for a known URL list, or a crawl to follow links within one website. Pick a page limit that fits your plan and start the job. Watch the per-page outcomes rather than assuming the entire collection succeeded.
Source rules, robots checks and network restrictions still apply to every page. A failed or blocked page does not become usable simply because it was included in a collection. You can cancel work that no longer serves your purpose.
Export and review before using it.
Download JSON when you want structured fields in code, or CSV for a tabular review. Keep the original export so you can compare later transformations against what was actually collected. Remove duplicate URLs and inspect unusually short responses.
Collection results remain available for seven days after the retention rules apply; export promptly. Store your own copy under a retention policy appropriate to your project. pageflock collects source material; it does not verify every claim in that material or decide whether it is suitable for your use.