Raspbytes

Documentation · Scraper API guide

How to prepare custom fields for the Scraper API

Turn page content into predictable structured records with CSS, XPath, or JSON selectors, then extend the request to repeated records, pagination, detail pages, and JavaScript-rendered websites.

Selectors

CSS, XPath, or JSON

Structured

Single or repeated records

Bounded

Pagination and detail pages

Custom fields describe the exact values the Scraper API should return, such as product names, prices, article dates, ratings, or profile details. Use this guide when a managed parser does not match the page or when your application needs its own response shape.

When to use custom fields

Use custom fields when you know the structure of the page and want a predictable result tailored to your application. For example, this HTML:

Source HTML
<h1 class="product-name">Wireless Headphones</h1>
<span class="price">£79.99</span>

can become:

Structured result
{
  "product_name": "Wireless Headphones",
  "price": "£79.99"
}

Choose a managed parser when one already covers the target and fields you need. Choose custom fields for a custom layout, additional values, or a response contract you control. A request uses either a managed parser or custom fields, not both.

Start with the output you want

Define the result before inspecting the page. A product record might look like this:

Desired result
{
  "name": "Wireless Headphones",
  "price": "£79.99",
  "availability": "In stock",
  "image_url": "https://example.com/images/headphones.jpg"
}

This gives you four selectors to prepare. Use stable, descriptive field names. A field name must start with a letter or underscore and may then contain letters, numbers, underscores, or hyphens.

Define custom fields

Put custom fields under extractor.fields. Each field supports the following contract:

sourceSelector type: css, xpath, or json.
selectorExpression used to locate the value.
attributeOptional HTML attribute to return instead of visible text.
multipleReturn every matching value as an array instead of the first match.
Single-page custom fields
{
  "url": "https://example.com/products/wireless-headphones",
  "extractor": {
    "fields": {
      "name": {
        "source": "css",
        "selector": "h1.product-name"
      },
      "price": {
        "source": "css",
        "selector": ".product-price"
      },
      "image_url": {
        "source": "css",
        "selector": ".product-gallery img",
        "attribute": "src"
      }
    }
  }
}

When attribute is omitted, the field returns text. When it is present, the field returns that attribute. Set multiple totrue to return every match as an array.

Choose reliable selectors

Prefer semantic IDs, stable class names, data-* attributes, and meaningful HTML relationships. For example,[data-product-price] is normally more durable than a long positional selector such asmain > div:nth-child(2) > div:nth-child(4) > span.

Use your browser developer tools to inspect the value, then confirm that the selector identifies only the intended element. Avoid generated class names and deeply nested positional paths because small layout changes can break them.

Return all product features
{
  "features": {
    "source": "css",
    "selector": ".product-features li",
    "multiple": true
  }
}

Extract repeated records

Listing pages usually contain several records with the same structure. Use record_selector to identify each item container; field selectors then run relative to the matching item.

Repeated product records
{
  "url": "https://example.com/products",
  "extractor": {
    "record_selector": {
      "source": "css",
      "selector": ".product-card",
      "limit": 50
    },
    "fields": {
      "name": { "source": "css", "selector": ".product-name" },
      "price": { "source": "css", "selector": ".product-price" },
      "product_url": {
        "source": "css",
        "selector": "a.product-link",
        "attribute": "href"
      }
    }
  }
}

Relative evaluation prevents a name from one card being paired with a price or URL from another. The record limit can be between 1 and 100.

Use XPath or JSON when appropriate

CSS selectors are usually simplest. XPath is useful when selection depends on document hierarchy or more complex relationships. When an XPath field runs inside a repeated record, begin it with .// so it remains relative to that record.

Record-relative XPath
{
  "record_selector": {
    "source": "xpath",
    "selector": "//article[contains(@class, 'product-card')]"
  },
  "fields": {
    "name": { "source": "xpath", "selector": ".//h2" },
    "price": {
      "source": "xpath",
      "selector": ".//span[contains(@class, 'price')]"
    }
  }
}

Use source: "json" for page-level structured JSON. JSON fields cannot be combined with record_selector; use CSS or XPath for repeated HTML records.

JSON fields
{
  "fields": {
    "product_name": {
      "source": "json",
      "selector": "product.name"
    },
    "price": {
      "source": "json",
      "selector": "product.offers.price"
    }
  }
}

Add pagination and detail pages

Pagination follows one same-host next-page link at a time. It requires a repeated-record selector, supports CSS or XPath, and accepts a page limit from 2 to 20, including the first page.

Listing pagination
{
  "record_selector": {
    "source": "css",
    "selector": ".product-card"
  },
  "fields": {
    "name": { "source": "css", "selector": ".product-name" },
    "price": { "source": "css", "selector": ".product-price" }
  },
  "pagination": {
    "source": "css",
    "selector": "a.next-page",
    "attribute": "href",
    "limit": 5
  }
}

Use follow when a listing contains summary cards but the full record lives on each detail page. The link selector runs relative to the listing record, while follow.fields runs from the detail-page root. A detail field replaces a listing field with the same name.

Follow product detail pages
{
  "record_selector": {
    "source": "css",
    "selector": ".product-card"
  },
  "fields": {
    "name": { "source": "css", "selector": ".product-name" },
    "product_url": {
      "source": "css",
      "selector": "a.product-link",
      "attribute": "href"
    }
  },
  "follow": {
    "source": "css",
    "selector": "a.product-link",
    "attribute": "href",
    "limit": 20,
    "fields": {
      "description": {
        "source": "css",
        "selector": ".product-description"
      },
      "sku": {
        "source": "css",
        "selector": "[data-product-sku]",
        "attribute": "data-product-sku"
      }
    }
  }
}

Handle JavaScript-rendered pages

Enable render_js when the required elements appear only after client-side rendering. The request starts in managed Chromium instead of trying standard HTTP retrieval first.

Browser-rendered extraction
{
  "url": "https://example.com/dynamic-products",
  "method": "GET",
  "render_js": true,
  "follow_redirects": true,
  "extractor": {
    "record_selector": {
      "source": "css",
      "selector": ".product-card"
    },
    "fields": {
      "name": { "source": "css", "selector": ".product-name" },
      "price": { "source": "css", "selector": ".product-price" }
    }
  }
}

Use it only when the target requires rendering. Standard retrieval is usually faster when the data is already present in the returned HTML.

Handle warnings and choose a response format

A missing field does not necessarily fail the request. A single-value field may return null, a multiple field may return an empty array, and the response can include a warning identifying the missing value or affected record.

Result with optional values
{
  "name": "Wireless Headphones",
  "price": "£79.99",
  "previous_price": null,
  "features": []
}

Choose response_format: "json" for software integrations orresponse_format: "markdown" for LLM, agent, and readable text workflows. The format changes presentation, not extraction behaviour.

A normal new request performs a fresh retrieval. An explicit idempotency key makes retries safe by replaying the original request outcome; omit it when every submission must collect the page again.

Preparation checklist

Define the exact output fields you need before choosing selectors.
Confirm every selector matches the intended value.
Prefer stable IDs, class names, and data attributes.
Specify the correct attribute for links and images, such as href or src.
Use record_selector when the page contains repeated records.
Start record-relative XPath fields with .//.
Set multiple to true for fields that can return several values.
Use a next-page selector and a sensible limit for pagination.
Place detail-page fields under follow.fields.
Enable render_js only when the content requires browser rendering.
Handle missing values, empty arrays, and warnings in your application.
Omit the idempotency key when every submission must perform a fresh retrieval.

Complete example

This request extracts products from a JavaScript-rendered listing, follows pagination, and visits each product page for additional fields.

Complete custom extraction request
{
  "url": "https://example.com/products",
  "method": "GET",
  "render_js": true,
  "response_format": "json",
  "extractor": {
    "record_selector": {
      "source": "css",
      "selector": ".product-card",
      "limit": 50
    },
    "fields": {
      "name": { "source": "css", "selector": ".product-name" },
      "price": { "source": "css", "selector": ".product-price" },
      "product_url": {
        "source": "css",
        "selector": "a.product-link",
        "attribute": "href"
      }
    },
    "pagination": {
      "source": "css",
      "selector": "a.next-page",
      "attribute": "href",
      "limit": 5
    },
    "follow": {
      "source": "css",
      "selector": "a.product-link",
      "attribute": "href",
      "limit": 25,
      "fields": {
        "description": {
          "source": "css",
          "selector": ".product-description"
        },
        "image_url": {
          "source": "css",
          "selector": ".product-gallery img",
          "attribute": "src"
        },
        "availability": {
          "source": "css",
          "selector": ".availability"
        }
      }
    }
  }
}

Start with a small field set and validate it against representative pages. Add repeated records, pagination, browser rendering, or detail-page following only when the source requires them.