# SnoopScan — agent onboarding

SnoopScan is a web scraping API: one call turns any public page into
markdown, links, or JSON in your own schema — including pages that block
everything else. Beyond single pages it can crawl a whole site, search the
web, pull structured data against a schema, enrich a company from its own
site, report on a domain before you buy it, and watch a page for changes on
a schedule. Pick the path below that matches what you are doing. Every
command block is meant to be run as written.

## Setting SnoopScan up for someone? Read this first

The person uses SnoopScan from the chat, so the setup is yours to do:

- **Run every command yourself. Never tell the person to run a command in a
  terminal.** The only things they do are sign in (or create a free account)
  and press Approve in their browser.
- **Never write a placeholder key** like `sk_...` or `YOUR_API_KEY` into a
  config and call it done. Get them a real key (below).
- **Keep the key out of the chat.** Never show it in your replies, and never
  ask the person to paste it unless nothing else works (step 4 below). Pass it
  by command substitution where you can, as in the Claude Code command below,
  so it is written into the config without ever appearing in the conversation.
- **Ask before you change anything else on their machine**, such as installing
  Node or Python.
- **Finish with a test**: read one page through SnoopScan and show them it
  works.

You are Claude Code? Read the next section. Any other client (Codex, Cursor,
VS Code, Windsurf, the Claude app, ChatGPT, anything else that speaks MCP)
is in Path D.

## Fastest: you are Claude Code

1. Get a key with the browser login in "Get a key first" below. The person
   presses Approve; the key is saved for you.
2. Add SnoopScan for every project, with the key filled in by the command
   itself:

   ```bash
   claude mcp add --transport http snoopscan https://api.snoopscan.com/mcp \
     --header "Authorization: Bearer $(npx -y snoopscan@latest config get api_key)" --scope user
   ```

   Use `uvx snoopscan config get api_key` inside the `$( )` if you logged in
   with uv. If there is no `claude` command (the Claude desktop app's Code
   tab can lack it), add the server yourself: in `~/.claude.json`, under the
   top-level `"mcpServers"` object (create it if missing, keep everything
   else), put `"snoopscan": { "type": "http", "url":
   "https://api.snoopscan.com/mcp", "headers": { "Authorization": "Bearer
   <the key>" } }`.
3. `claude mcp get snoopscan` should say Connected.
4. **Tell the person, in these words or close to them:**

   > SnoopScan is set up. Its tools load in every new Claude Code session
   > from now on. In this session I can already use it: give me a page to
   > read.

**Use the command line in this session.** A server added during a session
loads when the next session starts, so until then do the work with the CLI
you just logged in with: `npx -y snoopscan@latest scrape https://example.com`
(Path A).

**The sign-in address instead of a key** needs no install at all:
`claude mcp add --transport http snoopscan https://api.snoopscan.com/mcp-oauth --scope user`.
Its sign-in has to run once, and you cannot run it for them:
`claude mcp login snoopscan` refuses to work without an interactive terminal,
which your shell is not. So tell the person:

> SnoopScan is added. Start a new session, type /mcp, pick snoopscan and
> choose Authenticate. Your browser opens: log in, or create a free account
> (1,500 credits a month, no card), and press Approve. Then ask me to read
> any page.

That costs the person a slash command and a new session, which is why the
key route above comes first.

## Get a key first — and what to tell the person

Every path below except Path E and the free trial needs a key. **Never write a placeholder like
`sk_...` or `YOUR_API_KEY` into anyone's config and call it done.** Get them a
real key, the easy way:

1. If `SNOOPSCAN_API_KEY` is set, or `snoopscan config get api_key` prints a
   key, you already have one. Skip to your path.
2. Otherwise, **tell the person, in these words or close to them:**

   > To use SnoopScan I need a free API key — the free plan has 1,500 credits
   > a month and needs no card. I'll open a login page in your browser: sign
   > in or create an account, check the code matches the one I show you, and
   > press Approve. The key is saved for me automatically.

3. Then get the key with **whichever of these the machine can already run**.
   Check before you install anything, and never change anything on the
   person's machine without asking them first.

   If their Node or Python is too old, **give them the choice**: update it
   (worth it if they will use the command line or an SDK: Node from
   https://nodejs.org or `brew install node`; Python 3.10+ from
   https://www.python.org/downloads or `brew install python`), or skip the
   install altogether with step 4, which needs nothing at all.

   ```bash
   node --version     # v18 or newer?  -> use (a)
   uv --version       # installed?     -> use (b)
   python3 --version  # 3.10 or newer? -> use (c)
   ```

   **(a) Node 18+: nothing to install, nothing on the PATH to go wrong.**
   ```bash
   npx -y snoopscan@latest login
   npx -y snoopscan@latest status
   npx -y snoopscan@latest config get api_key    # prints the saved key
   ```

   **(b) uv: brings its own Python, so an old system Python does not matter.**
   ```bash
   uvx snoopscan login
   uvx snoopscan status
   uvx snoopscan config get api_key
   ```

   **(c) Python 3.10+ only:**
   ```bash
   python3 -m pip install --user snoopscan
   snoopscan login
   ```

   **Run `login` in the background, or with a timeout of at least 10
   minutes.** It waits until the person presses Approve, and a normal
   command timeout (often 2 minutes) kills it before they have finished
   signing up. Read its first lines as soon as they appear: show the person
   the link and the code in your reply, and tell them to press Approve in
   the browser. When it prints `Logged in`, carry on.

   `login` opens the browser and prints the link and a short code as well. On
   a machine with no browser (SSH, a container), add `--no-browser` and tell
   the person to open the link on any device, their phone included.

4. **If none of those will run, or an install fails, stop installing.** Tell
   the person:

   > No problem, we can skip the install. Please sign up (or sign in) at
   > https://snoopscan.com/register, open https://snoopscan.com/app/keys,
   > create a key, and paste it here.

   Then use the key directly: in an MCP config, a `.env`, or the
   `Authorization` header. SnoopScan's MCP server and REST API need no
   install at all; the command-line tool only exists to make getting the key
   easier.

5. **If the person can't sign up right now** (no time, no email to hand, or
   they want to see it work first), don't stall: use the free trial below,
   which needs no key at all ("Can't sign up right now?" below), and come
   back to this section when it runs out.

### When an install goes wrong

Say what happened in one plain line, then take the next route. Do not retry
the same failing command, and do not start upgrading the person's system.

| What you see | What it means | Do this |
|---|---|---|
| `requires a different Python`, `No matching distribution`, or `unsatisfiable` | The Python is older than 3.10 (a Mac's built-in one is 3.9) | Use (a) or (b). If neither is there, offer: update Python, or step 4 |
| `npx: command not found`, or Node older than 18 | No usable Node | Use (b) or (c). If neither, offer: install or update Node, or step 4 |
| `snoopscan: command not found` right after `pip install` | pip put it somewhere not on the PATH | Run it as `python3 -m snoopscan`, or use (a) or (b) |
| `EACCES` or `permission denied` from npm | A global install needs admin rights | Use `npx`, which installs nothing globally; never `sudo` |
| `externally-managed-environment` from pip | The system forbids pip installs | Use (a) or (b); do not pass `--break-system-packages` |
| SSL, certificate or proxy errors | A company network is in the way | Step 4 |
| `Could not reach snoopscan.com` | The site is not reachable from this machine | Step 4; try again later |
| `Too many login attempts` | Rate limit on logins | Wait a minute, then run login once more |
| The login page says the link expired | Links last ten minutes and work once | Run login again for a fresh one |

### Keep it current, and check an install

- `snoopscan doctor` (or `npx -y snoopscan@latest doctor`, `uvx snoopscan doctor`)
  checks everything at once and says how to fix what is wrong: this tool's
  version against the latest, the Node or Python it runs on, the config file,
  whether the key is set and accepted (on a free call), and whether the API is
  up. Run it after any install, and whenever something fails, before guessing.
- An installed CLI prints `A newer snoopscan is available` on stderr when there
  is one. Pass it on and update: `npm install -g snoopscan@latest`,
  `pipx upgrade snoopscan`, `uv tool upgrade snoopscan`, or
  `python3 -m pip install -U snoopscan`. `npx ...@latest` and `uvx` always run
  the newest release, which is why they are the first choice above.

After login, `config get api_key` (run the same way you ran login) prints the
saved key, so you can put the real value into an MCP config or `.env` yourself
instead of asking the person to paste it.

Every request carries the key as `Authorization: Bearer <key>`. A request
without one is refused with a message that says how to get one; pass that on
rather than guessing. A 401 with a key that looks right almost always means it
was copied with a trailing newline — strip whitespace.

## Can't sign up right now? The free trial, no key

SnoopScan answers a few calls with **no key at all**, so you can do the job
now and sign up later:

- **MCP:** add `https://api.snoopscan.com/mcp` with **no** `Authorization`
  header. `scrape`, `search` and `parse` work on the trial; every other tool
  answers with the sign-up steps. In Claude Code:

  ```bash
  claude mcp add --transport http snoopscan https://api.snoopscan.com/mcp --scope user
  ```

  (The server loads in the next session. In this one, use REST below.)
- **REST:** `POST /v1/scrape`, `/v1/search` and `/v1/parse` with **no**
  `Authorization` header:

  ```bash
  curl -X POST https://api.snoopscan.com/v1/scrape \
    -H "Content-Type: application/json" \
    -d '{"url": "https://example.com/pricing", "formats": ["markdown"]}'
  ```

  `/v1/parse` takes a file upload (up to 10 MB) as it does with a key.

What the trial is, so you can plan around it:

- **20 free successful calls a day per visitor** (per address), reset at
  midnight UTC. No card, no account.
- **Plain pages only**: a page that needs a browser needs a key. Formats are `markdown`,
  `html`, `rawHtml` and `links`; screenshots, JSON extraction, page actions
  and a chosen country need a key. Search returns up to 10 results, with page
  content for the first 3.
- **Every answer shows the calls left**: a `trial` object on REST (`remaining`
  of `limit`, plus `X-RateLimit-Remaining`), a closing line on MCP ("Free
  trial: 19 of 20 free calls left today"). Pass it on when it gets low.
- **A blocked page doesn't use up the allowance.**
- **Never send a placeholder or empty `Authorization` header.** Any header at
  all, even a wrong one, is treated as a key and refused.

**Go back to "Get a key first" when** an answer says the page needs a
browser ("This page needs a browser"), or the calls run out (a
429 `RATE_LIMITED`, or "This address has used all 20 free SnoopScan calls for
today"). Tell the person:

> That page needs SnoopScan's browser, which comes with a free account
> (1,500 credits a month, no card). I'll open a login page: sign in or create
> an account and press Approve, and I'll carry on from there.

(Say "You've used today's free calls" instead of the first sentence when that
is the reason.) Then swap the no-key server for the keyed one: remove it
(`claude mcp remove snoopscan --scope user`) and add it again with the key as
in "Fastest" above, or add the header to the config you wrote.

## Choose your path

- **Shell access, want results now** → Path A
- **Writing application code (Python/JS)** → Path B
- **No install, one-off REST call** → Path C
- **Want this as a callable tool in Claude Code, Codex, Cursor, the Claude app, ChatGPT, etc.** → "Fastest" above, or Path D
- **No key yet, just checking what this does** → Path E
- **Can't sign up right now, but the job needs doing** → the free trial above

## Path A — CLI

Needs Node 18+ or Python 3.10+; pick the install from "Get a key first"
above. For a lasting install rather than `npx`/`uvx` each time:

```bash
npm install -g snoopscan           # or: uv tool install snoopscan, or: pipx install snoopscan
snoopscan login                    # browser sign-in; saves the key (see above)
snoopscan status                   # checks the key and that the API is reachable
```

Every command shares `-o/--output` (write to a file instead of stdout),
`--json` (full response, not just the content), `--pretty`, `--tier`
(`auto` by default; `browser` for a page that only works in a real
browser), `--timeout` (ms), `--max-age` (accept a cache hit this many ms
old), and `--formats` (comma-separated: `markdown,html,rawHtml,links`).

```bash
snoopscan scrape https://example.com/pricing --formats markdown,links

snoopscan crawl https://example.com --limit 50 --wait > site.jsonl
snoopscan crawl-status <job_id>

snoopscan map https://example.com                         # every URL on the site, no page bodies

snoopscan search "site:example.com pricing" --scrape --limit 10

snoopscan extract https://example.com/product/1 https://example.com/product/2 \
  --schema schema.json --prompt "price and stock status"

snoopscan parse ./quarterly-report.pdf                    # local file to markdown — PDF, DOCX, XLSX, HTML

snoopscan products https://shop.example.com                # a store's whole product catalog
snoopscan posts https://blog.example.com                   # every post on a blog or forum

snoopscan company https://acme.com                          # firmographics + contacts from a company's own site
snoopscan company https://acme.com --no-contacts            # firmographics only, cheaper

snoopscan domain example.com                                 # registration + DNS + backlinks, all three by default
snoopscan domain example.com --no-backlinks

snoopscan monitor create --name "Pricing" --urls https://example.com/pricing \
  --interval 60 --goal "Alert when the price or plan names change"
snoopscan monitor list
snoopscan monitor run <monitor_id>
```

`company` enriches one company from its own site: name, phone, address,
LinkedIn, headcount, industry, contact emails, social links and contact form.
To find businesses by trade and town, use the `findLeads` MCP tool or
`POST /v1/leads`.

Full command list: `scrape`, `crawl`, `crawl-status`, `map`, `search`,
`extract`, `parse`, `products`, `posts`, `company`, `domain`, `monitor`,
`config`, `status`. The domain buyer's report and the bulk domain screen have
no command yet: use the MCP tools (`domainReport`, `domainSnapshots`,
`domainScreen`) or REST (`/v1/domain/report`, `/v1/domain/screen`). The JS CLI mirrors every command and flag exactly:
`npm install -g snoopscan`.

## Path B — SDKs

Put the real key in the project's environment first:
`echo "SNOOPSCAN_API_KEY=$(snoopscan config get api_key)" >> .env` (and make sure
`.env` is git-ignored).

```python
import os
from snoopscan import SnoopScan
snoop = SnoopScan(api_key=os.environ["SNOOPSCAN_API_KEY"])  # the real key, never a placeholder
page = snoop.scrape("https://example.com/pricing")
print(page.markdown)

pages = snoop.crawl_and_wait("https://example.com", limit=50)   # blocks until the crawl finishes
for page in pages:
    print(page.url, len(page.markdown or ""))
```

```javascript
import { SnoopScan } from 'snoopscan';
const snoop = new SnoopScan({ apiKey: process.env.SNOOPSCAN_API_KEY! }); // the real key
const page = await snoop.scrape('https://example.com/pricing');
console.log(page.markdown);
```

The SDKs use the same names and fields as the CLI and the REST API:
`scrape`, `crawl`, `map`, `search`, `extract`, `company`, `domain` and
`monitor` work the same way in all three.

## Path C — REST

```bash
curl -X POST https://api.snoopscan.com/v1/scrape \
  -H "Authorization: Bearer $SNOOPSCAN_API_KEY" -H "Content-Type: application/json" \
  -d '{"url": "https://example.com/pricing", "formats": ["markdown", "links"], "onlyMainContent": true}'
```

The main fields for each endpoint:

- **`/v1/scrape`** — `formats` (`markdown`, `html`, `rawHtml`, `links`,
  `media`, `summary`, `screenshot`, or a `json` object with your schema —
  ask for several at once), `onlyMainContent` (default `true`, drops nav/
  footer/cookie banners), `includeTags`/`excludeTags` (CSS selectors),
  `waitFor`, `timeout`, `maxAge` (accept a cached hit this old).
- **`/v1/crawl`** — `url`, `limit` (default 100, up to 100,000),
  `maxDepth` (default 3), `maxConcurrency` (default 5), `includePaths`/
  `excludePaths`, `allowExternalLinks`, `ignoreSitemap`,
  `deduplicateSimilarURLs` (default `true`). Async: returns a job id,
  poll `/v1/crawl/{id}`.
- **`/v1/map`** — `url`, `search` (filter the URL list by a term),
  `limit` (default 5,000, up to 30,000), `includeSubdomains`,
  `ignoreSitemap`. No page bodies are fetched — this is discovery only.
- **`/v1/search`** — `query`, `limit` (default 10, up to 50), `sources`,
  `location`, and an optional `scrapeOptions` to fetch each result too
  (same shape as `/v1/scrape`'s body).
- **`/v1/extract`** — `urls` (1–100), `schema` (your JSON schema),
  `prompt`, and an optional `scrapeOptions`. Pass your own `model`
  (provider and key) and model calls are billed to your key, at no credit
  cost.
- **`/v1/company`** — `url` (a domain or any URL on it — `acme.com` and
  `https://acme.com/about` enrich the same company), and `contacts`
  (default `true`). Returns company details and the contacts the site
  publishes.
- **`/v1/domain`** — `domain`, plus three booleans that default to
  `true`: `registration`, `dns`, `backlinks`.
- **`/v1/domain/report`** — `domain`, `depth` (`quick`, `buyer` by
  default, or `deep`), and optionally `snapshots` and `verifyLinks`. A
  buyer's report before purchasing a domain: registration, DNS, today's site,
  archive history, backlinks and whether linking pages still link, with
  findings and concerns tied to their evidence. Archived snapshots it could
  not read in time come back later from `/v1/domain/history/snapshots`.
  Hosted API only.
- **`/v1/domain/screen`** — `domains` (up to 500) and `dns` (default
  `true`). Which names can be registered, from each registry's own record
  and DNS: cut a long list to a shortlist, then report on the few that
  matter. Hosted API only.
- **`/v1/monitor`** — create with `name`, `urls`, `intervalMinutes`
  (default 60, minimum 5), `webhook`, and `goal` (a note of what you are
  watching for). `GET /v1/monitor/{id}/checks` for history.

Every endpoint above, plus `/v1/parse`, `/v1/products` and `/v1/posts`,
is documented in full at https://snoopscan.com/docs.

## Path D — MCP

SnoopScan runs an MCP server at `https://api.snoopscan.com/mcp` (streamable
HTTP, `Authorization: Bearer <key>`), and a sign-in address with no key at
`https://api.snoopscan.com/mcp-oauth`. Every tool the API has is a tool there:
`scrape`, `fetchMore`, `crawl`, `crawlStatus`, `crawlPages`, `map`, `search`,
`parse` (a document at a URL, such as a PDF, as markdown), `extract`, `listProducts`, `listPosts`, `checkChanges`, `domain`, `company`,
`findContacts`, `people`, `hiring`, `findLeads`, `leadsStatus`, and the
domain buyer's tools `domainReport`, `domainSnapshots` and `domainScreen`.
The domain buyer's tools, the product and post listings and the company,
contact, people, hiring and lead tools run on the hosted API only.

Where a client below takes a key, get it first ("Get a key first" above) and
write it into the config without repeating it in the chat. A server added
during a session loads in the next one; use Path A meanwhile.

**Claude Code** — see "Fastest: you are Claude Code" above.

**Codex** (the Codex CLI, the IDE extension and the Codex app share
`~/.codex/config.toml`) signs in by itself, so no key is needed:

```bash
codex mcp add snoopscan --url https://api.snoopscan.com/mcp-oauth
codex mcp login snoopscan     # in the background, with a long timeout
```

`login` prints a link: put it in your reply and tell the person to sign in or
create a free account, then press Approve. `codex mcp list` should show
snoopscan. With a key instead: `url = "https://api.snoopscan.com/mcp"` and
`http_headers = { Authorization = "Bearer <the key>" }` under
`[mcp_servers.snoopscan]`.

**Cursor** — merge into `~/.cursor/mcp.json` (every project) or
`.cursor/mcp.json` (this one):

```json
{ "mcpServers": { "snoopscan": { "url": "https://api.snoopscan.com/mcp", "headers": { "Authorization": "Bearer <the key>" } } } }
```

Or `"url": "https://api.snoopscan.com/mcp-oauth"` with no headers, and Cursor
asks the person to sign in when it connects.

**VS Code and GitHub Copilot** — merge into `.vscode/mcp.json`:

```json
{ "servers": { "snoopscan": { "type": "http", "url": "https://api.snoopscan.com/mcp", "headers": { "Authorization": "Bearer <the key>" } } } }
```

For every workspace, `code --add-mcp` takes the same server as one JSON
object with `"name": "snoopscan"` added. The `-oauth` address with no headers
signs in instead.

**Windsurf** — under `"mcpServers"` in `~/.codeium/windsurf/mcp_config.json`
(or `~/.config/devin/mcp_config.json` on newer releases; use whichever
exists):

```json
{ "mcpServers": { "snoopscan": { "serverUrl": "https://api.snoopscan.com/mcp", "headers": { "Authorization": "Bearer <the key>" } } } }
```

Then tell the person to press refresh in Cascade's MCP panel if it is not
listed yet.

**If the person uses the Claude app** (desktop, claude.ai or the phone), there
is no command and nothing to run. Connectors there take a web address, and
this one signs in, so there is no key to copy. Tell them:

"In Claude, open Settings, then Connectors. Click + and choose Add custom
connector. Name it SnoopScan, paste `https://api.snoopscan.com/mcp-oauth`,
and click Add. Then click Connect, log in or create a free account, and press
Approve."

On a Team or Enterprise plan, an owner adds the connector for the organization
first.

**If the person uses ChatGPT**, it connects through developer mode (Plus, Pro,
Business, Enterprise and Education, on the web). Tell them:

"In ChatGPT, open Settings, then Security and login, and turn on Developer
mode. Then go to chatgpt.com/plugins, press +, and create an app named
SnoopScan with the server address `https://api.snoopscan.com/mcp-oauth` and
OAuth sign-in. Sign in or create a free account, and press Approve. In a
chat, pick it from Developer mode in the + menu."

**Anything else** that speaks MCP over streamable HTTP takes
`https://api.snoopscan.com/mcp` with the header `Authorization: Bearer <the
key>`, or `https://api.snoopscan.com/mcp-oauth` with no header if it signs
in. A client that only runs local servers can bridge with
`npx -y mcp-remote https://api.snoopscan.com/mcp --header "Authorization: Bearer <the key>"`.
Setup pages for forty-odd clients: https://snoopscan.com/integrations.

If `/mcp` was added without a key, it still connects: `scrape`, `search` and
`parse` run on the free trial (see "Can't sign up right now?" above), and
every other tool answers with the setup steps — follow them.

## Path E — no key yet, just checking what this even does

`https://snoopscan.com` has a free playground on the homepage: type a URL,
pick read, browse, map, crawl or search, and see the result. No signup. For
an agent that needs the answers itself, use the free trial above over MCP or
REST.

## If something looks wrong

- A 401 almost always means the key was copied with a trailing newline —
  the whitespace survives shell substitution more often than not.
- `snoopscan status` (or `GET /health`) tells you whether the key is bad or
  the API is down. Check it before assuming either.
- Branch on the error's `code`, not its message:
  `INVALID_REQUEST`, `UNAUTHORIZED`, `FORBIDDEN_SCOPE`, `RATE_LIMITED`,
  `TIMEOUT`, `BLOCKED`, `FETCH_FAILED`, `TARGET_ERROR`,
  `EXTRACTION_FAILED`, `ROBOTS_DENIED`, `JOB_NOT_FOUND`,
  `INSUFFICIENT_CREDITS`, `PROXY_UNAVAILABLE`, `SEARCH_UNAVAILABLE`,
  `PLACES_UNAVAILABLE`, `SERP_UNAVAILABLE`, `COMPANY_UNAVAILABLE`, `PLATFORMS_UNAVAILABLE`,
  `ENGINE_REFUSED`, `INTERNAL`.
- `BLOCKED` or `TARGET_ERROR` on a page that loads in a normal browser:
  retry with `--tier browser` (CLI) or `"tier": "browser"` in the body
  before calling the page unreachable.
- `PLACES_UNAVAILABLE`, `COMPANY_UNAVAILABLE` or `PLATFORMS_UNAVAILABLE`:
  the endpoint is only on the hosted API, `https://api.snoopscan.com`, not
  on a self-hosted copy. `PROXY_UNAVAILABLE`: set `proxy` to `auto`.
